AI & Research

When AI Gets Sustainability Wrong

Oct 6, 2026 · 7 min read

Dr. Elham KheradmandCEO, Lucid Axon
In this article
  1. The number that was never disclosed
  2. Last year’s planet
  3. The spin machine
  4. Same claim, new author
  5. What good looks like
  6. The fair comparison

Hallucinated disclosures, stale data, and greenwashing amplified by models trained on corporate self-reporting.

AI can read a 200-page sustainability report in seconds. It can also invent a number that was never in it, quote a commitment the company dropped a year ago, and repeat the company’s own spin with more confidence than the company had.

These are three different failures with one root. A general-purpose model is fluent in text written by the companies it is asked to assess, so its answers sound more certain than the evidence underneath. The research on each failure is now good enough to be specific about it, and about what fixes it.

The number that was never disclosed

In a 2026 benchmark built on 94 real ESG reports, 15.6% of GPT-4o’s answers were fabricated or unsupported.[1] Another 34.8% were incomplete, and fewer than half were fully correct. The benchmark, ESG-Bench from the University of Sheffield, worded its questions to provoke errors, so read it as a stress test and not a field error rate.

The pattern behind the number matters more than the number. Asked for a metric the report does not contain, a model tends to produce one anyway. In sustainability data, “not disclosed” is often the most important answer, and it is the one a language model is least inclined to give.

Retrieval helps, but it does not make the problem go away. A GPT-4 pipeline extracting data from the reports of 166 Hong Kong-listed companies reached 76.9% accuracy,[2] which leaves roughly one data point in four wrong. It also could not read charts, where a good share of ESG data lives.

The clearest real-world warning comes from outside sustainability. In 2025, Deloitte Australia delivered a government report containing references to academic work that did not exist and an invented quote from a court judgment. Deloitte repaid A$97,587 of its fee.[11] An outside academic caught the errors, not the firm’s own review.

We have not found a verified case of an AI-fabricated sustainability disclosure. That may say more about how rarely AI use is disclosed than about how rarely it goes wrong.

Last year’s planet

Even a model that never hallucinates is working from old numbers. The European Central Bank assumes a one-year delay on corporate emissions data, and up to two years for sovereigns.[10] A portfolio’s “current” footprint describes an economy that has already moved on.

Much of that data was never reported at all. Third-party estimates often make up more than half of emissions observations, and they mostly reflect a company’s size and industry. One study found them at least 2.4 times less effective than reported figures for identifying the worst emitters.[6] On Scope 3, two major vendors’ estimates correlated at only about 0.14 to 0.16.[7]

The history moves too. Researchers documented widespread changes to past ESG scores at one large provider, enough to change the results of return tests.[9] And raters disagree with each other in the first place: correlations across six of them average 0.54.[8]

Then the rules changed. Between April 2025 and March 2026, the EU cut most companies out of its sustainability reporting directive, Canadian securities regulators paused their mandatory climate disclosure rule, and Canada narrowed its greenwashing law. The Net-Zero Banking Alliance, once nearly 150 banks, ceased operations in October 2025.[12]

A model trained on text from 2021 to 2024 will still describe those banks as members. It has no date attached to what it knows.

The spin machine

When generative AI writes sustainability disclosures, it greenwashes, and readers find the result more credible than the real thing. That is the finding of a 2026 experiment from the University of Auckland and Aalto University.[4]

Postgraduate students who had studied greenwashing used ChatGPT to draft sustainability statements for a fictional company. The drafts scored high on greenwashing, the students’ edits did little to fix them, and other students rated them more positive and credible than real corporate disclosures. It is one classroom study and needs replicating with professionals, but its direction is not a surprise.

The reason is what the models learned from. Most of what is written about a company’s sustainability is written by the company: reports, press releases, web pages. Research on climate disclosures has found much of that text to be cheap talk, with firms cherry-picking the risks that matter least.[5] Firms that talk the most cheaply also show higher emissions growth.

The adverse evidence lives elsewhere: in enforcement orders, court filings, NGO investigations and local-language news. It is sparse, scattered and harder to parse. A model that weighs sources by volume will hear the company first and loudest.

The same technology cuts the other way when pointed at the right question. The cheap-talk findings came from a language model trained to separate specific commitments from vague ones. A 2026 study used an LLM pipeline to turn ten years of reports from 600 European companies into dated, checkable numbers, validated against expert annotations.[3]

Same claim, new author

The claim belongs to whoever makes it, whichever tool drafted it. In April 2025, German prosecutors fined asset manager DWS €25 million for ESG marketing that did not match reality.[13] In 2024, the SEC penalized two advisers, one of them Toronto-based, for overstating their use of AI.[14]

A vendor selling AI for sustainability is exposed on both sides. Its claims about the AI and its claims about the sustainability are each statements of fact that must be substantiated.

In Canada, Bill C-15 received royal assent on March 26, 2026 and narrowed the Competition Act’s greenwashing provisions.[15] It removed the “internationally recognized methodology” test and private parties’ direct route to the Tribunal for claims about a business. Claims must still be adequately substantiated, the general ban on misleading representations still applies, and the Commissioner keeps full enforcement powers.

The amendment relaxed how a company substantiates an environmental claim. It did not remove the duty to substantiate one, and “the model wrote it” is not substantiation.

What good looks like

Each of these failures has a known fix, and the fixes are design choices. Six of them have evidence behind them.

  1. Cite the page. Every extracted figure should point to the page and to whether it came from text, a table or a chart. This is what lets a human check it.
  2. Let the model say “not disclosed.” In ESG-Bench, training models to abstain raised their accuracy on unanswerable questions to 98 to 99%. Abstention is a result, and often the finding.
  3. Validate against independent data. The European study checked its output against a benchmark dataset and expert annotations before drawing any conclusion.
  4. Date everything. Each data point needs its reporting period, its publication date and the date it was retrieved. Revisions should be versioned, never overwritten.
  5. Label what is reported, estimated and modelled. These are different kinds of evidence. Blending them hides the weakest numbers behind the strongest.
  6. Triangulate, then have an expert sign off. Company text should be checked against enforcement records, litigation, NGO findings and asset-level data. A domain expert makes the final call.

None of this is exotic. It is the discipline already expected of financial data, applied to a field that has gone without it.

The fair comparison

The honest benchmark for AI is current practice, and current practice is weak. Expert-built ESG ratings barely agree with each other, estimates stand in for missing disclosures, and histories get rewritten. A disciplined pipeline can already cover more companies, more consistently, than any analyst team.

So the question for anyone buying or building these tools is narrow and practical. Can it show its source? Can it say “not disclosed”? Does every number carry a date and a label? Is anything independent of the company in the evidence?

A system that answers yes to those is doing decision intelligence. One that cannot is repeating what companies say about themselves, faster and with more confidence than they said it. See how Lucid Axon works.

  1. [1]Sun et al., ESG-Bench: Benchmarking Long-Context ESG Reports for Hallucination Mitigation, AAAI 2026
  2. [2]Zou et al., ESGReveal, Journal of Cleaner Production, 2025
  3. [3]Forster et al., Assessing corporate sustainability with large language models: evidence from Europe, Nature Communications, 2026
  4. [4]Dimes, de Villiers, Denny & Leinonen, The Impact of Generative AI on the Perpetuation and Detection of Greenwashing in Sustainability Reports, Business Strategy and the Environment, 2026
  5. [5]Bingler, Kraus, Leippold & Webersinke, How cheap talk in climate disclosures relates to climate initiatives, corporate emissions, and reputation risk, Journal of Banking & Finance, 2024
  6. [6]Kalesnik, Wilkens & Zink, Do Corporate Carbon Emissions Data Enable Investors to Mitigate Climate Change?, Journal of Portfolio Management, 2022
  7. [7]Busch, Johnson & Pioch, Corporate carbon performance data: Quo vadis?, Journal of Industrial Ecology, 2022
  8. [8]Berg, Kölbel & Rigobon, Aggregate Confusion: The Divergence of ESG Ratings, Review of Finance, 2022
  9. [9]Berg, Fabisik & Sautner, Is History Repeating Itself? The (Un)predictable Past of ESG Ratings, ECGI working paper, 2021
  10. [10]European Central Bank, climate data FAQ
  11. [11]CFO Dive, Deloitte AI debacle seen as wake-up call for corporate finance, 2025
  12. [12]Ballotpedia News, Net Zero Banking Alliance ends operations after member exodus, October 2025
  13. [13]The D&O Diary, Deutsche Bank Asset Management Unit Pays €25m Greenwashing Fine, April 2025
  14. [14]U.S. Securities and Exchange Commission, press release 2024-36, March 2024
  15. [15]Osler, Further amendments to the environmental claims provisions of the Competition Act, 2026

Can your AI show its source?

Send us your methodology and five entities. We’ll run the assessment and show you the results.