This policy brief is based on Bank of Italy, Occasional Paper N. 1003. The views expressed are those of the authors and not necessarily those of the institutions the authors are affiliated with.
Abstract
Credible climate-related disclosure depends on reliable emissions data. Yet Scope 3 emissions data for euro area banks — relative to financed emissions from credit and investment portfolios — are affected by gaps, inconsistencies, and logical anomalies. This paper documents such anomalies and then evaluates whether three leading generative AI (GenAI) tools — Claude (Anthropic), ChatGPT (OpenAI), and Gemini (Google) — can address these shortcomings. GenAI-generated data are significantly correlated with estimates from professional providers and help identify anomalies. However, they also present quality and consistency issues and raise concerns about replicability and transparency. Going forward, the development of specialised language models and improved reporting standards could make GenAI a valuable complementary data source, but it cannot yet replace traditional reporting.
Climate-related disclosure has become a cornerstone of risk management for investors, policymakers, and financial institutions navigating the green transition, yet critical data gaps persist. This is particularly true for banks’ Scope 3 emissions, namely those related to their credit and investment portfolios, which remain poorly measured, sparsely reported, and riddled with inconsistencies. These emissions typically account for more than 95% of a bank’s total carbon footprint and are central to evaluating banks’ transition risks, and decarbonization performance. Despite their limitations, Scope 3 emissions are essential for such assessments, as Scope 1 and Scope 2 emissions relate only to banks’ internal operations and provide no information on financed activities. At the same time, they represent a particularly complex measurement challenge for banks compared to non-financial firms, since to compute them, banks must use not only value‑chain emissions but also those of borrowers and investee companies, ideally relying on borrower‑ and issuer‑level data, or otherwise on sectoral proxies. Given such challenges, financial institutions often rely on data provided by ESG data vendors when assessing borrower emissions. Our focus is on banks’ emissions of the euro area given their central role in the European financial intermediation and the particularly ambitious climate disclosure regulation developed in the region with the goal of fostering the channeling of funds to the climate transition.
Our analysis builds on new evidence showing that Scope 3 emissions of listed banks in the euro area display several unexpected and undesirable features, using data from well‑known professional providers such as LSEG Datastream, ISS, MSCI, and Bloomberg. First, Scope 3 emissions data are often missing even for large and medium-sized listed banks— although many are subject to transparency requirements— with data coverage ranging between 32% and 69% depending on the provider. Beyond coverage, the quality of available data raises serious concerns. Bank‑level emissions exhibit substantial heterogeneity and, in some cases, a low correlation across data providers, reflecting methodological differences for model‑based estimates, with some exceeding the average by more than 100%. Moreover, the data displays extreme volatility over time in the period 2018-2022 – during which data availability progressively increased, largely as a result of enhanced disclosure requirements: frequent and recurring annual variations exceeding 100% are common and likely driven not by changes in banks’ activities but by repeated methodological revisions by data providers. Further, banks’ carbon footprints — measured as the ratio of their Scope 3 emissions to total loans — are not systematically higher for those institutions most exposed to energy-intensive sectors such as mining, energy, or transport. Finally, and most strikingly, in several instances reported financed emissions (Scope 3) are lower than a bank’s own operational energy-related emissions (Scope 2), suggesting an implausible underestimation of portfolio-related Scope 3 emissions.
Figure 1. Banks’ scope 3 emissions: data coverage and carbon footprint as of 2022 by source

These unwarranted features, together with the low reliability of available bank‑level emissions data, hinder their use for research and policy purposes at the current stage. Against this background, we explore whether widely available generative AI (GenAI) tools, whose adoption is fast growing also for financial applications, can serve as a novel and valuable source of information. Thanks to their ability to access, extract, and elaborate information from multiple structured and unstructured sources — including reports, disclosures, and online databases — and to interact iteratively with users, GenAI tools may help recover fragmented climate‑related information and support data quality assessment. This paper provides first‑of‑its‑kind empirical evidence on whether such tools can help fill the documented data gaps and enhance banks’ Scope 3 emissions data, complementing insights from other domains discussed by Korinek (2023).
Figure 2. Heterogeneity of scope 3 emissions of individual banks across data sources

To assess whether GenAI could help bridge this data gap, the paper tests three leading tools — Claude (Anthropic), ChatGPT (OpenAI), and Gemini (Google) — on a common task: retrieving or estimating Scope 3 emissions for our sample of major euro area banks for the reference year 2022.
The tools were used in two modes. In retrieval mode, models were asked to recover emissions data, in these cases they were often extracted from banks’ sustainability reports, TCFD disclosures, and Climate Disclosure Project (CDP) questionnaires or other sources. In estimation mode, models were asked to provide their best estimate based on their training or broadly available data or complement the dataset where self-reported figures were unavailable.
While we standardized the prompt design via identical requests across tools, their architectures and functionalities led to remarkable differences in approaches and outcomes. Claude 3.5 and 3.7, without web search enabled, generated tables drawing on training data; Claude 3.7 proved capable of distinguishing between retrieved (banks’ self-reported) figures and estimated ones, enhancing transparency. ChatGPT required multiple iterations and struggled with long lists of banks processed simultaneously, highlighting computational constraints for data‑intensive tasks. Gemini 2.5 Pro conducted deep searches across more than 400 websites, retrieving information, for instance, directly from annual reports, websites and CDP questionnaires.
This diversity in approach is itself informative: the tools rely on different architectures in terms of multimodality and tokenization (i.e., the types of inputs they can process and the way information is internally represented), context window (i.e., the amount of information they can process in a single prompt), as well as alignment and training philosophies, which shape how models interpret and respond to prompts. This raises issues about the comparability of results across tools, even when models draw on similar or partially overlapping training datasets and web-based information (when enabled).
The analysis reveals a mixed picture: promising potential alongside built-in limitations.
On the positive side, GenAI-generated emissions figures are more often available and correlate significantly with some of the data supplied especially by some professional providers, such as Bloomberg and ISS. This pattern suggests that large language models are either trained on similar underlying datasets and/or apply estimation approaches broadly aligned with established estimation methodologies.
A notable advantage of GenAI over providers’ static databases is its interactive nature. When users flagged anomalies — such as instances where Scope 2 emissions appeared higher than Scope 3 — Claude and ChatGPT were often able to acknowledge the inconsistency with most realistic scenarios, where financed emissions exceed operational energy-related emissions, and then self-correct their estimates. This quick user-guided self-correction capability has no close analogue in traditional data providers.
However, significant limitations remain. Different tools produced different outputs, and even the same tool generated different results across models. Replicability is a core concern: identical queries submitted to different models within the same tool, or with bank lists of different lengths, produced different outputs — particularly for ChatGPT. This inconsistency makes it difficult to use GenAI data in a systematic, audit-proof framework. Transparency is equally problematic: in several instances, models labelled data as “self-reported” when the figures were likely to be estimates, or vice versa, blurring the boundary between retrieval and estimation modes. Finally, because large language models are constrained by the information set they are trained on or can access, they suffer from the same disclosure-related shortcomings: where corporate reporting is scarce or low quality, GenAI cannot reliably compensate for missing primary data.
Figure 3. GenAI-based scope 3 emissions: data coverage and carbon footprint for the 2022

Figure 4: Correlation between GenAI-based scope 3 emissions data and other data sources


The findings carry insightful implications for supervisors, regulators, and market participants working with emissions data.
GenAI tools are best fitted as a complementary source rather than a substitute for professional data providers or mandatory corporate disclosure. They can add value in specific use cases: cross-checking figures from traditional sources, flagging anomalies that warrant closer scrutiny, and partially filling gaps for institutions where public information exists but is scattered across multiple documents. Their interactive nature also makes them useful diagnostic tools for analysts.
At the same time, the replicability and transparency issues documented in this research prevent, at present , the use of GenAI‑generated emissions data as a primary input for investment decisions, prudential assessments, or regulatory compliance without substantial verification.
Two structural conditions would improve the reliability and usefulness of GenAI in this domain. First, the development of domain-specific language models fine-tuned on climate and financial data — rather than general-purpose models — could enhance both the accuracy and consistency of outputs. Second, and more fundamentally, progress depends on improved mandatory disclosure standards.
In the meantime, investors and supervisors should treat both professional providers’ and GenAI‑based data with appropriate caution, cross‑checking multiple sources and carefully examining methodological approaches and data quality. Going forward, as GenAI tools become widespread in the financial sector also for climate risk management, appropriate provisions would consider how AI-generated data could be integrated into regulatory frameworks, and how the use of AI in financial risk assessment might be governed and oversighted.
The Scope 3 emissions data currently available for euro area banks is insufficient for the purposes that regulators, investors, and financial institutions increasingly expect them to serve. Professional data providers offer partial coverage, volatile estimates, and in some cases logically implausible results. This paper shows that GenAI tools can contribute to addressing these gaps up to a point and are not a silver bullet. Inconsistent outputs, transparency limitations, and inherent disclosure-related data quality problems mean that AI-generated emissions data should be treated as one input among several, not as a reliable standalone source. Ultimately, the path to improved climate data runs through better corporate disclosure, not only better technology. Until the underlying reporting landscape improves, both human analysts and AI tools will continue to work with imperfect raw material.
Angelico, C. & Bernardini, E. (2026), Can GenAI fill banks’ emissions data gaps? Bank of Italy Occasional Paper No. 1003.
Korinek, A. (2023), ‘Generative ai for economic research: Use cases and implications for economists’, Journal of Economic Literature 61(4), 1281–1317.