This policy brief is based on the recent Guidance note 6 published by the BIS Irving Fisher Committee on Central Bank Statistics (IFC). The views expressed are their own and do not necessarily reflect the views of the BIS, the IFC or its members.
Abstract
The growing availability of information sources has offered central banks new opportunities to enhance their statistical function. By linking – or integrating – various data sets, they have been able to produce more granular, timely and diverse statistics in a cost-efficient way. These advancements have also enabled a better use of information available in society, such as administrative records, to improve statistical agility in responding to user needs. Yet integrating alternative data – often generated as a by-product of other processes – also raises challenges, including concerns over accuracy, representativeness and reliability. Central banks’ experience highlights both the strategic and operational importance of data integration for both users and producers of statistics. It also underscores the need for strengthening the global statistical infrastructure, through adequate data governance, management and public resources.
The digitalisation of economies has led to an unprecedented proliferation of data sources, offering central banks new opportunities to enhance their statistical function. By leveraging innovative techniques and tapping into novel data, they can produce more granular, timely and diverse statistics in a cost-efficient way. These advancements also enable a better use of the rich information available in today’s society, such as administrative data, to improve statistical agility in responding to user needs during unforeseen events and address the challenges posed by increasing non-response rates in surveys. However, this surge in alternative data supply raises challenges, including concerns over quality, representativeness and the risk of secondary data crowding out reliable official statistics.
The above opportunities and challenges have raised interest in the concept of data integration. This process, also known as linking or fusion, basically refers to combining multiple data sources. The main goal is to fill information gaps effectively and efficiently, by leveraging various existing data while minimising reporting burden.
This policy brief discusses the benefits, challenges and future directions of data integration in the context of central banking, emphasising the need to strengthen the global statistical infrastructure through adequate data governance, management and resources.
Data integration features three main benefits, namely enabling additional analytical insights, making better use of existing data and enhancing data accuracy.
First, it enables central banks to address diverse user demands for multidimensional information. Using a greater variety of data sources can better support policy areas such as climate change, financial stability and economic monitoring, for instance by linking micro and macro data, as well as diverse types like text, images and sensor data. Examples include assessing banks’ credit exposures by merging loan data with business registers or estimating aggregate consumption patterns through payments. Beyond enhancing analytical perspectives, data integration also provides more flexible approaches, allowing users to “zoom in” on specific events without losing macro perspectives and enabling compilers to create new aggregates quickly from granular inputs without initiating new data collections (Israël and Tissot (2021)).
Second, effective data integration can maximise the use of existing sources with three advantages. Firstly, it helps reduce reporting burden by identifying redundant data collections, such as overlapping statistical and supervisory reporting, as envisioned by the new Eurosystem Integrated Reporting Framework (IReF) project. Secondly, it can support filling information gaps cost-effectively, enabling users to meet data needs without new collections and allowing producers to enhance coverage using administrative, credit and business registers. The Covid-19 pandemic served as a good reminder for leveraging all the data available to maintain statistical production in an agile way despite sudden supply disruptions (de Beer and Tissot (2020)). Thirdly, data integration can improve timeliness, leveraging real-time sources like web-scraped data for nowcasting inflation, compiling consumer expenditure indicators and analysing market sentiment, ensuring responsiveness to new information demands.
Finally, central banks’ experience with data integration shows that it can enhance the accuracy and quality of statistics. By cross-referencing information, compilers can assess consistency, conduct plausibility checks and perform mirror analyses to ensure coherent and reliable measurement of statistical indicators. Leveraging secondary data also supports quality assurance, editing tasks and post-survey adjustments, helping compilers to address issues such as missing values or non-response in surveys (Scotti et al (2024)).
Combining multiple data sources does not come without challenges and costs. These can be grouped into three broad categories.
One important challenge has to do with the still fragmented information standards and limited adoption of common identifiers. The use of multiple standards – such as Statistical Data and Metadata eXchange (SDMX) for macroeconomic statistics, eXtensible Business Reporting Language (XBRL) used for a number of financial reporting exercises and other ISO standards (eg for payments and geospatial information) – may lead to fragmentation and higher costs for both data compilers and users, complicating maintenance and efficient information reuse (IFC (2025b)). Furthermore, while the adoption of global identifiers – eg Legal Entity Identifier (LEI) for financial entities – has significantly expanded over the past decades, their often voluntary adoption and competition with local or proprietary systems can ultimately undermine their effectiveness.
Implementing data integration into IT infrastructures may also raise a number of challenges, including the need to manage complex formats and large data sets, as well as to migrate from legacy systems to more modular, high-performance platforms. Such transitions often create a heterogeneous IT environment, requiring costly dual-system operations and potentially increasing security risks. Furthermore, while cloud services offer scalable and cost-effective solutions, they also raise concerns about data security, sovereignty and reliance on third-party providers. Maintaining robust IT security is also becoming increasingly difficult, for instance with advances in quantum computing that may threaten current cryptographic methods at some point (Auer et al (2025)). Additionally, a shortage of specialised staff, such as data engineers and scientists, is a persistent issue – particularly in low- and middle-income countries – hindering the modernisation of IT infrastructure.
Lastly, leveraging various information sources for statistical purposes raises several methodological, operational and ethical challenges. Secondary and alternative data sources, often generated as by-products of non-statistical activities, may lack alignment with established statistical quality frameworks, leading to potential biases and inconsistencies. They can also disrupt statistical continuity due to the inherent volatility of these data or because of the possibility of a sudden discontinuation in data providers. This makes it essential to ensure stable data-sharing agreements and contingency plans (MacFeely (2020)). Organisational silos further complicate integration processes by fragmenting data and hindering collaboration, resulting in inefficiencies and inconsistent information. Turning to ethical and legal considerations, these arise from the potential privacy and confidentiality risks associated with linking data, making it essential to implement robust rules, operational safeguards and advanced technologies to protect sensitive information. More fundamentally, they highlight the importance of carefully assessing the reliance on external data providers, not least to ensure that the independence and integrity of statistical compilation exercises are maintained.
Central banks’ experience shows that maximising data use and value through integration calls for making further progress in four areas: (i) data governance and management; (ii) adequate data access and sharing; (iii) data quality; and (iv) international cooperation.
First, sound data governance and management are fundamental to effective multisource statistics, as they provide the framework and tools needed to manage and access the various information assets both within and across organisations. For example, central banks’ efforts in data governance show that the establishment of a data stewardship function can help break down data silos and foster data reuse. It also represents a useful way to clarify responsibilities, streamline processes, for example to reduce redundancies, and optimise data access. Turning to data management systems, the implementation of well established standards such as SDMX helps ensure that data can be efficiently found, exchanged and unambiguously interpreted, enabling semantic interoperability and consistent metadata management.
Second, adequate access to and sharing of information also plays a central role in enabling data reuse. In practice, central banks have already taken several steps towards achieving this goal, for example by creating accessible data portals, establishing data centres – eg to provide secure, privacy-preserving environments for sharing both aggregated and granular data (IFC (2025c)) – and promoting structured exchange of information beyond organisational boundaries – ie with statistical offices, regulators and supervisors at the national and international level, such as the BIS International Data Hub (Bese Goksu and Tissot (2018)). Central banks also engage with other statistical bodies to exchange best practices, as seen in the recommendations of the Data Gaps Initiative (DGI) endorsed by the G20 for establishing internationally agreed principles on data sharing and access (Marini (2025)).
Third, effective data integration requires that diverse sources – particularly secondary and emerging alternative data – are incorporated into robust statistical and data quality frameworks, for instance through effective assurance and curation processes. This approach is in particular essential for managing the often-lower quality of novel data. Concretely, addressing quality challenges may require focussing on three key areas: first, documenting best practices for using non-official sources in statistics, such as publishing recommendations and guidance material informed by practical experience; advancing metadata quality frameworks can in addition play a central role, including to meet the needs of AI-ready data (IFC (2025d), Sirello et al (2025); Verhulst et al (2025)); lastly, promoting the development and adoption of data maturity frameworks, such as the one introduced by the G20 DGI, will help organisations benchmark and enhance their data processes.
Fourth, integrating the large variety of available data sources calls for strengthened collaboration with all the data stakeholders involved. Central banks can leverage their well established partnerships within the statistical system, notably with national statistical offices, regulators and other government bodies, as well as international organisations. Such collaborations can be instrumental in broadening the range of available sources and optimising data collection, particularly in relation to the financial system. Collaboration may also expand to other stakeholders in the data ecosystem, especially the private sector, academia and citizens, which increasingly generate vast amounts of data that could potentially be better used for statistical purposes (UNECE (2025)). However, collaboration with private stakeholders needs to be assessed against an adequate level of accessibility, equity and inclusiveness, particularly for countries with limited statistical capacity (Amutorine et al (2024)).

Amutorine, M, N Lawrence and J Montgomery (2024): “Increasing data sharing and use for social good: lessons from Africa’s data sharing practices during the Covid-19 response”, Data and Policy, vol 6, no e53, November.
Auer, R, D Dodson, A Dupont, M Haghighi, N Margaine, D Marsden, S McCarthy and A Valko (2025): “Quantum-readiness for the financial system: a roadmap”, BIS Papers, no 158, July.
de Beer, B and B Tissot (2020): “Implications of Covid-19 for official statistics: a central banking perspective”, IFC Working Papers, no 20, November.
Bese Goksu, E and B Tissot (2018): “Monitoring systemic institutions for the analysis of micro-macro linkages and network effects”, Journal of Mathematical and Statistical Science, April.
Irving Fisher Committee on Central Bank Statistics (IFC) (2025a): “Data integration in central banks: opportunities and challenges”, IFC Guidance Note, no 6, July.
———– (2025b): “SDMX adoption and use of open source tools”, IFC Report, no 17, February.
———– (2025c): “Data science in central banking: enhancing the access to and sharing of data”, IFC Bulletin, no 64, May.
———– (2025d): “Governance and implementation of artificial intelligence in central banks”, IFC Report, no 18, April.
Israël, J-M and B Tissot (2021): “Incorporating microdata into macro policy decision-making”, Journal of Digital Banking, vol 6, no 3.
MacFeely, S (2020): “In search of the data revolution: has the official statistics paradigm shifted?”, Statistical Journal of the IAOS, vol 36, no 4.
Marini, M (2025): “Self-assessment tool for data access and sharing maturity”, presentation at the DGI Global Conference, Cape Town, June.
Scotti, C, C Rondinelli, M Bottone, E Mattevi and A Neri (2024): “Central bank business surveys: version 2.0”, VoxEU, 9 December.
Sirello, O, B Bogdanova and M Erdem (2025): “Metadata in the age of AI: the role of official statistics in securing a virtuous cycle”, Statistical Journal of the IAOS, July.
Verhulst, S, A Zahuranec and H Chafetz (2025): “Moving toward the FAIR-R principles: advancing AI-ready data”, SSRN, no 5164337, March.
United Nations Economic Commission for Europe (UNECE) (2025): The future of national statistical offices: a call to action, May.