menu
close

Author(s):

Koji Takahashi | Seikei University
Joon Suk Park | Bank of Korea

Keywords:

ChatGPT , generative artificial agents , privacy paradox , Westin index , survey

JEL Codes:

M31 , C83 , C45 , D12 , L86

This policy brief is based on Generative AI for Surveys on Payment Apps: AI View on Privacy and Technology” BIS working papers No 1333. The views expressed in this note are those of the authors and do not necessarily reflect the official views of the Bank for International Settlements and Bank of Korea.

Abstract

This study evaluates the capacity of Generative AI (GenAI), specifically ChatGPT-4o, to simulate human survey responses regarding the adoption and perception of payment apps. The findings demonstrate that GenAI can effectively replicate the privacy calculus, a decision-making process where individuals compare the convenience of a service with potential data risks. The AI successfully identified that app users perceive significantly higher benefits and lower risks than non-users, mirroring real-world behavioral patterns without explicit prompting. However, the study also identifies a “diversity deficit,” as the AI-generated responses lack the natural variability found in human populations. Furthermore, the AI exhibits an inherent “privacy fundamentalist” bias, consistently overstating security concerns compared to actual Dutch survey participants. Consequently, while GenAI serves as a valuable tool for survey pre-testing and brainstorming, it remains a complement to, rather than a replacement for, traditional human-based market research.

Introduction

As businesses and policymakers increasingly seek rapid, cost-effective ways to gauge public opinion, the rise of Generative AI has sparked a debate over whether large language models can serve as digital twins for human survey participants. If an AI can reliably mirror the attitudes, biases, and decision-making processes of a specific population, it could revolutionize market research by providing instant feedback on new products or policy changes. This study specifically explores this possibility within the context of the payment industry, where the tension between technological convenience and data privacy—known as the “privacy calculus”—is a defining characteristic of consumer behavior. The central question is whether an AI, when assigned a specific persona, can truly replicate the nuanced and often contradictory nature of human thought.

Scripting the Digital Persona

To test the reliability of artificial agents, we designed a series of experiments comparing ChatGPT-4o’s responses to a real-world survey of the Dutch population conducted by Brits and Jonker (2023). The experiments involved creating prompts that assigned the AI specific roles based on variables such as age, gender, and whether the agent was a user of payment apps. A key component was the inclusion of the Westin Privacy Index, which categorizes individuals into three groups—Privacy Fundamentalists (highly concerned), Privacy Pragmatists (balanced), and the Privacy Unconcerned—based on their responses to three questions regarding data privacy and security practices in society. We also asked ChatGPT ten additional questions regarding the perceived benefits and risks associated with the usage of payment apps. By layering these psychological profiles onto the AI, we observed whether the model adjusted its perceived benefits and perceived risks of payment apps in a manner aligned with human logic. For instance, human respondents categorized as fundamentalists tend to report low benefit and high risk scores regarding the use of payment apps. This approach allowed the study to measure not only whether the AI could answer questions, but also whether AI agents exhibit the internal consistencies typically found in human data.

Internal Consistency and Risk Averse Tendency

We simulated responses based on several different settings. In Case I, we just specified age, gender and user/non-user profile in prompts. Figure 1 illustrates a consistent trend where less-concerned types report higher benefit scores and lower risk scores. Moreover, the standard errors—which are higher for benefits and lower for risks within each Westin category—closely mirrors the actual survey data. In addition, the simulated agents correctly intuited that payment app users generally view technology as highly beneficial and low-risk, while non-users maintain the opposite view, even when the AI was not explicitly told to correlate these factors (Figure 2). However, the data also revealed a significant “diversity deficit.” While a human population displays a wide range of opinions, the AI-generated responses were much more uniform, resulting in data that lacked variability (Figure 1). Furthermore, the AI exhibited a persistent “privacy fundamentalist” bias, consistently rating the risks of data misuse higher than actual Dutch citizens did as shown in Figure 3. This suggests that the AI’s internal safety training, designed to make it cautious, inadvertently makes its digital personas more risk-averse than the average person.

Because the distribution of Westin types did not match the actual survey data, we explicitly specified the Westin category in the Case 2 prompts. This ensured that the proportion of each type exactly mirrored the original study, alongside age, gender, and user/non-user profiles. As illustrated in Figure 4, this resulted in a more pronounced divergence among respondents of different Westin types regarding the risks and benefits of payment apps. However, the distinction between users and non-users became more ambiguous, as shown in Figure 5. For example, 23.4% of the ChatGPT-simulated non-users showed a high benefit score paired with a low risk score—a higher proportion than in the original survey. This suggests that specifying certain characteristics in prompts can have unintended effects on the responses, even though the user and non-user persona prompts remained identical to those in Case 1.

Figure 1. Risk and Benefit Scores by Westin type in Case 1

Figure 2. Perceived benefits and scores for Case 1 by user and non-user

Figure 3. Share of respondents by Westin type

Figure 4. Risk and Benefit Scores by Westin type in Case 2

Figure 5. Perceived benefits and scores for Case 1 by user and non-user

A Copilot, Not a Replacement

The implications of this research for the financial and market research sectors are two-fold. First, Generative AI proves to be an invaluable copilot for the early stages of research, such as pre-testing survey questions, identifying potential red flags in product design, or simulating extreme what-if scenarios at a minimal cost. Second, however, the study serves as a warning against the complete replacement of human samples with artificial ones. Because AI tends to produce an echo chamber of polarized or overly cautious views, relying solely on digital agents could lead to skewed results that ignore the diverse, and often irrational, perspectives of a real-world audience. For now, the most effective strategy remains a hybrid approach where AI sharpens the tools, but real humans provide the final, authentic pulse of the market.

References

Brits, Hans and Jonker, Nicole (2023). The Use of Financial Apps: Privacy Paradox or Privacy Calculus?. De Nederlandsche Bank Working Paper No. 794.

Takahashi, K., and Park, J. S. (2026). Generative AI for Surveys on Payment Apps: AI View on Privacy and Technology. BIS Working Paper No 1333.

About the authors

Koji Takahashi

Koji Takahashi is a Professor at Seikei University. Previously, he served as the head of the Economic Studies Group at the Bank of Japan’s Institute for Monetary and Economic Studies (2024–2026) and was a visiting economist at the Bank for International Settlements (2021–2024).

Joon Suk Park

Joon Suk Park is a Senior Manager at the Bank of Korea and served as a visiting economist at the Bank for International Settlements (2022-2024).

More on these topics

Tags:
Privacy Overview

This website uses cookies so that we can provide you with the best user experience possible. Cookie information is stored in your browser and performs functions such as recognising you when you return to our website and helping our team to understand which sections of the website you find most interesting and useful.