Main Facts: The Promise and Peril of Automated Footprinting
As global regulatory pressures mount and the race toward Net Zero intensifies, multinational corporations are facing an unprecedented challenge: calculating the carbon footprint of every single item they produce. This process, known as a Product Carbon Footprint (PCF), is notoriously labor-intensive, often requiring months of manual data collection and expert analysis. Into this breach has stepped Generative Artificial Intelligence, promising to transform a "tedious drudgery" into a near-instantaneous digital output.
However, a landmark study from researchers at Watershed, a leading carbon accounting platform, has issued a stark warning. While AI models can produce "impressive-sounding" final emission estimates, they frequently fail at the granular level. The study, led by Krishna Rao and his colleagues, revealed that while the best-performing AI models could estimate a total footprint within a reasonable margin of error 77% of the time, their ability to accurately break down the components of those emissions—the "intermediate steps"—plummeted to as low as 37%.
The implications are significant for sustainability officers. If a manufacturer relies on a flawed AI decomposition of a product’s lifecycle, they may invest millions in supply chain shifts that target the wrong components, effectively chasing "ghost emissions" while the true environmental culprits remain unaddressed. As the industry moves from voluntary reporting to mandatory disclosures under frameworks like the EU’s Corporate Sustainability Reporting Directive (CSRD), the gap between AI’s perceived capability and its actual accuracy represents a looming compliance risk.
Chronology: From Manual LCAs to the AI Gold Rush
The evolution of carbon accounting has moved through three distinct phases over the last two decades, leading to the current inflection point.
The Era of Manual Life Cycle Assessments (2000s–2015)
In the early days of corporate sustainability, measuring a product’s impact was the domain of specialized consultants. Using Life Cycle Assessment (LCA) methodologies, experts would manually trace a product from "cradle to grave." This involved interviewing suppliers, auditing factory energy bills, and consulting massive, static databases like Ecoinvent. Because of the cost and time involved, companies typically only performed LCAs for their "hero products"—the top 1% of their portfolio.
The Rise of Carbon Accounting Software (2016–2022)
As "ESG" (Environmental, Social, and Governance) became a boardroom priority, the first wave of carbon accounting platforms emerged. These tools moved the data from spreadsheets to the cloud, allowing for better tracking of Scope 1 (direct) and Scope 2 (energy-related) emissions. However, Scope 3 (supply chain) emissions remained a "black box." The industry realized that manual LCAs could not scale to meet the demand of tracking thousands of SKUs.
The Generative AI Pivot (2023–Present)
The release of Large Language Models (LLMs) changed the trajectory of the industry. Startups and established players alike began integrating AI to "read" bills of materials (BOMs) and automatically map them to emission factors. In 2024, the market saw a surge in "automated PCF" solutions. Companies like Makersite, Terrascope, and Watershed itself began leveraging AI to bridge the data gap, promising to deliver insights in seconds that previously took months. This period has been characterized by high optimism, recently tempered by the critical findings of the Watershed research team.
Supporting Data: The Accuracy Gap in Black-Box Models
To test the efficacy of AI in this specialized field, the Watershed team designed a rigorous benchmarking process. They utilized 175 diverse products across sectors including chemicals, textiles, electronics, and industrial materials. The researchers pitted models from the world’s leading AI labs—Anthropic, DeepSeek, Google, and OpenAI—against expert-verified "ground truth" data.
The "Two Multiples" Metric
The study first looked at the "Macro Accuracy" of the models. The benchmark for success was whether the AI could come within "two multiples" of the expert answer (meaning if the true footprint was 10kg of CO2, the AI’s answer fell between 5kg and 20kg).
- Top Performance: The best models achieved this 77% of the time.
- The Interpretation: While a 23% failure rate is high for financial accounting, in the nascent world of carbon estimation, this was initially seen as a promising "ballpark" figure.
The "Decomposition" Failure
The study’s most critical finding emerged when the researchers asked the models to show their work—a process known as "Chain of Thought" (CoT) processing. They required the models to decompose products into their constituent materials and assign emissions to each part.
- Granular Accuracy: Success rates dropped to 37%.
- The "Hallucination" Problem: In many cases, the AI would correctly guess the total emissions for a product (perhaps by referencing similar products in its training data) but would attribute those emissions to the wrong materials. For example, it might overestimate the impact of a device’s plastic casing while drastically underestimating the impact of its integrated circuits.
The PwC Context
This data explains a persistent trend in the corporate world. A recent PwC study found that 69% of companies have created PCFs for less than a quarter of their product lineups. The "drudgery" of manual calculation is the primary barrier, but the Watershed data suggests that current AI "shortcuts" may be providing a false sense of progress rather than a reliable solution.

Official Responses: Industry Leaders Weigh In
The findings have sparked a debate among the "Carbon Tech" elite regarding the balance between speed and precision.
Krishna Rao, Watershed Researcher:
Rao, the lead author of the study, emphasized that the goal was not to dismiss AI, but to refine its application. "What surprised me most was the size of the gap [between final results and intermediate steps]," Rao stated. His message to the industry is one of cautious integration: "They [users] have to understand the intermediate steps. They cannot rely on the final results alone."
The Competitive Landscape:
While Watershed has taken a self-critical approach by benchmarking its own tools, other players in the space maintain that AI-driven speed is essential for climate action.
- Makersite continues to market its ability to "automate accurate LCAs across your entire product portfolio in seconds," emphasizing that for many companies, an 80% accurate view of 100% of their products is better than a 99% accurate view of only 1% of their products.
- Terrascope claims it can achieve 70% accuracy without requiring direct data from suppliers, leaning heavily on the predictive power of its proprietary models to fill in data gaps where primary information is missing.
Sustainability Professionals:
The general consensus among ESG analysts is a shift toward "Human-in-the-Loop" (HITL) systems. The consensus is that AI should be used as a "first-draft" engine, but a qualified human expert must still audit the "decomposition" phase to ensure the logic aligns with physical reality.
Implications: The High Stakes of "Hallucinated" Carbon Data
The Watershed study is more than an academic exercise; it has real-world consequences for the global economy and the environment.
1. Misallocation of Decarbonization Capital
If a global retailer uses AI to determine that their biggest emission source is packaging, they will spend millions on biodegradable alternatives. If the AI was wrong—and the true source was the logistics or the raw material extraction—the company’s carbon footprint will remain unchanged despite the investment. Flawed data leads to flawed strategy, which in turn leads to a failure to meet climate targets.
2. The Regulatory "Greenwashing" Trap
With the advent of the EU’s Green Claims Directive and the SEC’s climate disclosure rules, companies are legally liable for the environmental claims they make. If a company publishes a PCF generated by an AI that "hallucinated" the numbers, they could face massive fines and reputational damage. The "I didn’t know the AI was wrong" defense is unlikely to hold water with regulators.
3. The Need for "White-Box" AI
The study highlights the danger of "black-box" AI in science-based fields. For AI to be useful in carbon accounting, it must be "explainable." Future tools will likely need to provide citations for every emission factor used and allow users to override specific data points in the decomposition process.
4. A New Benchmarking Standard
Watershed’s study has effectively created a new "gold standard" for testing carbon-accounting AI. By releasing their benchmarking process, they are forcing the industry to move away from marketing claims of "instant accuracy" toward a more rigorous, transparent validation of how AI "thinks" about the physical world.
5. Supply Chain Relations
Finally, the reliance on AI to "guess" footprints without "troubling suppliers" (as some services claim) may actually damage supply chain transparency. Real decarbonization requires engagement with suppliers to change their energy mixes and processes. If AI is used to bypass this engagement, the primary incentive for suppliers to improve their own footprints—maintaining their status as a "low-carbon vendor"—is removed.
In conclusion, while AI remains a tantalizing solution to the bottleneck of carbon accounting, the Watershed study serves as a vital correction. The automation of sustainability is not a "set-and-forget" technology; it is a powerful tool that requires expert oversight, rigorous benchmarking, and a healthy dose of skepticism regarding the "impressive-sounding" numbers it produces.
