The AI Paradox: Can Automated Carbon Footprinting Deliver on its Sustainability Promise?

By Jim Giles

In the race to meet ambitious net-zero targets, corporations are increasingly turning to a seductive new solution: Artificial Intelligence. As the pressure mounts to disclose Scope 3 emissions—those elusive indirect emissions occurring throughout a product’s value chain—sustainability professionals are finding themselves buried under a mountain of data. The promise of AI is clear: take a bill of materials or a simple product description, and let a Large Language Model (LLM) instantly calculate the Product Carbon Footprint (PCF).

However, a sobering new study from researchers at Watershed, a leading carbon accounting platform, suggests that while AI may offer speed, it often sacrifices accuracy. The study reveals a troubling "black box" phenomenon where impressive-sounding final emissions figures mask significant errors in the underlying logic, raising urgent questions about whether businesses are relying on high-tech guesswork to make critical environmental decisions.

The Burden of the Footprint: Why Automation is the Goal

To understand the appeal of AI-driven PCFs, one must first understand the sheer, agonizing complexity of traditional lifecycle assessments (LCAs). Measuring a product’s carbon footprint requires a granular breakdown of every raw material, chemical input, and manufacturing process involved. Analysts must hunt down energy data from disparate suppliers, verify material weights, and apply conversion factors to estimate emissions.

The process is notoriously labor-intensive and prone to human error. It is a task that often requires weeks or months of coordination across global supply chains. The result? A systemic failure in corporate transparency. According to research from PwC, a staggering 69 percent of companies have managed to create PCFs for less than a quarter of their entire product lineups. For companies with thousands of SKUs, the task is effectively impossible without a technological leap.

This is the gap that a new wave of sustainability-tech firms is rushing to fill. Companies like Makersite and Terrascope are marketing AI as the ultimate "time-saver." Makersite promises to "automate accurate LCAs across your entire product portfolio in seconds," while Terrascope suggests its models can achieve 70 percent accuracy without the need to pester suppliers for primary data. Watershed itself offers tools designed to bridge this gap, yet the company’s internal research team has turned a critical eye on the reliability of the very technology they are helping to advance.

Chronology of the Research: Putting Models to the Test

To test the capabilities of current generative AI, Krishna Rao and his colleagues at Watershed conducted an extensive benchmarking study. They selected 175 diverse products across sectors ranging from heavy chemicals and textiles to complex electronics. They then prompted four of the most advanced AI models—developed by Anthropic, DeepSeek, Google, and OpenAI—to generate carbon footprint estimates based on the provided specifications.

The results were, at first glance, promising. When asked for a final emissions estimate, the best-performing model arrived within "two multiples" of the expert-verified answer 77 percent of the time. In the world of broad corporate reporting, where estimates are often ballpark figures, this initial performance might lead a sustainability officer to conclude that the technology is ready for prime time.

However, the Watershed team did not stop at the final number. They insisted on seeing the "work." They asked the models to decompose the products into their constituent components and estimate the emissions associated with each step of the manufacturing process. This is where the narrative of AI success began to unravel. When the researchers audited the intermediate steps, the accuracy plummeted. In some cases, the model’s reliability dropped to as low as 37 percent.

"What surprised me most was the size of the gap," said Rao. The study highlights a classic problem in modern AI: the models are excellent at synthesizing vast amounts of general knowledge to produce a plausible-looking output, but they struggle with the specific, logical chains required for high-stakes scientific accounting.

Supporting Data: The Illusion of Accuracy

The discrepancy between the final number and the intermediate steps is not just a technical footnote; it is a fundamental flaw. For a manufacturer, the final PCF number is often just a reporting requirement. But the data behind the number is a strategic asset.

Why AI carbon footprinting tools are prone to misleading results

If a company intends to reduce its carbon impact, it must identify "hotspots"—specific components or vendors that contribute disproportionately to the total footprint. If the AI suggests that a specific piece of aluminum hardware is the primary driver of emissions when, in reality, the error lies in the model’s failure to account for the energy grid of the manufacturing region, the company may waste months of effort and capital on a "solution" that yields zero actual carbon reduction.

The Watershed benchmarking process was intentionally designed to be applied to specialist AI tools, including those used by Watershed’s own clients. By breaking down the task into components, the team proved that a model can arrive at a "correct" final answer for the wrong reasons. This is a dangerous outcome for corporate sustainability, where the goal is not just to report a number, but to change industrial processes.

Official Responses and the Industry Stance

The response from the sustainability tech sector has been one of cautious engagement. Proponents of AI-driven LCAs argue that the technology is in its infancy and that the "accuracy gap" is a result of data scarcity rather than a flaw in the AI itself.

Many firms are pivoting toward a "human-in-the-loop" approach, where AI acts as an assistant to the analyst rather than an autonomous oracle. In this model, the AI performs the heavy lifting of data aggregation, while a human expert verifies the intermediate logic. Makersite, for instance, emphasizes the connectivity of their data to real-world industrial databases, suggesting that the integration of proprietary datasets can compensate for the hallucinations and logical errors typical of off-the-shelf LLMs.

However, the Watershed findings serve as a necessary cooling-off period for the industry. The message is clear: automation cannot currently replace the rigor of traditional LCA methodology. Until models can demonstrate consistency in their step-by-step reasoning, they must be treated as productivity tools, not auditing tools.

Implications for Sustainability Reporting

The implications of this research are profound for the regulatory landscape. With the European Union’s Corporate Sustainability Reporting Directive (CSRD) and the SEC’s climate disclosure rules gaining traction, companies are under increasing pressure to provide audit-grade data.

If companies use AI to generate these reports, they must be prepared to defend the logic behind those numbers. The Watershed study suggests that relying on an AI "black box" could leave companies vulnerable to accusations of greenwashing. If a firm reports a reduction in emissions based on AI-generated data that turns out to be mathematically flawed, the reputational and legal risks are significant.

The Path Forward: Transparency and Verification

For sustainability professionals, the path forward requires a shift in mindset:

  1. Demand Interpretability: If you are using an AI tool for carbon accounting, demand to see the intermediate steps. If the tool cannot explain how it arrived at an estimate, it is not ready for compliance-level reporting.
  2. Benchmark Against Experts: Companies should conduct their own "sanity checks" by running a small sample of their product line through both traditional LCA consultants and AI tools to measure the variance.
  3. Prioritize Primary Data: AI is best used to fill in gaps where data is impossible to obtain, not to replace primary data that can be collected from suppliers.
  4. Educate Stakeholders: Boards and investors should be informed that AI in sustainability is a work in progress. Overselling the accuracy of automated footprints creates an expectation gap that will eventually lead to investor skepticism.

"They have to understand the intermediate steps," Rao reiterated. His message to the industry is one of cautious skepticism. AI is a powerful tool for navigating the complexity of modern supply chains, but it is not a substitute for the meticulous, often tedious work of carbon accounting. As the industry moves forward, the successful companies will be those that use AI to augment human intelligence, rather than those that seek to replace it with a faulty, automated shortcut.

In the final analysis, the climate crisis is a problem of physics and chemistry, not just data processing. While AI can help us organize our efforts, it cannot yet fathom the intricate, carbon-heavy realities of industrial manufacturing. For now, the most sustainable approach is to ensure that our technology remains under the watchful, critical eye of the humans responsible for our collective future.

Leave a Reply

Your email address will not be published. Required fields are marked *