How Accurate Are Synthetic Respondents? What the Evidence Actually Shows

· 18 min read

Synthetic respondents are accurate enough to rank options, compare messages and map objections, and not accurate enough to produce a number you can defend. The published evidence supports directional and comparative use, not treating simulated answers as survey data. Which of the two you are doing decides whether the method helps you or quietly misleads you.

Key takeaways

Most writing here comes in two registers: vendors asserting parity, or researchers calling the method illegitimate in principle. This article concedes what should be conceded and states when synthetic users produce something worth acting on. If the category is new, start with how synthetic personas are built — the accuracy argument turns on it.

Why this question is hard to answer

"Are synthetic users accurate?" is unanswerable as posed, because accurate silently means three incompatible things.

Distribution match. Does the synthetic sample reproduce the population's answer distribution — the same share choosing option B, the same mean on a 1–7 scale? The strictest standard, and the one market sizing needs.

Directional agreement. Does it pick the same winner, order options the same way, move the same direction between two conditions? Weaker, and what most product decisions rest on.

Individual prediction. Can an agent built from one person's data reproduce that person's answers? A frontier question, largely irrelevant commercially — but the one with a sane yardstick. Toubia and colleagues, building their Twin-2K-500 dataset at Columbia, repeated tasks in a final survey wave so twin accuracy could be scored against test–retest: how well people reproduce their own answers. Perfect agreement is not the target, because humans do not achieve it.

A vendor claiming high accuracy usually means directional agreement; a critic reporting failure usually means distribution match. Both can be right.

What the research actually found

The literature is young, uneven, and better than the marketing on either side suggests. The table maps what has been attempted, not a scorecard.

Line of evidence What it did What tends to hold up What does not
Sarstedt, Adler, Rau & Schmitt's review of silicon samples (Psychology & Marketing, 2024) Reviewed studies comparing result patterns from silicon and human samples; the systematic review found 28 articles reporting 285 comparisons across seven domains and 96 tasks Personality traits, framing effects, political attitudes and party preferences — including Argyle and colleagues' silicon-sampling work, which the review counts among the replications Endowment effect, mental accounting, sunk-cost fallacy; overall, results "vary considerably across different domains"
Sun and colleagues' random silicon sampling (2024) Prompted a model with group-level demographic distributions only, then compared against US public-opinion polls Aggregate response distributions "remarkably similar to the actual U.S. public opinion polls" Replicability "varies depending on the demographic group and topic of the question", attributed to societal biases in the model
Toubia and colleagues' Twin-2K-500 dataset (2025) Surveyed 2,058 US participants, 2.42 hours each on average, over 500 questions and four waves, the last repeating tasks to fix a test–retest baseline Individual- and aggregate-level prediction show promise with deep per-person grounding; they report that interview-based twins built by Park and colleagues matched survey answers 85% as accurately as participants matched their own answers two weeks later The grounding is hours of interview and questionnaire data per person — not what a commercial panel supplies
Lin's work on synthetic respondents and the illusion of human data Reviewed AI contamination of online samples and its statistical consequences Nothing about synthetic quality — the finding is that internal coherence is no longer evidence of anything LLM-mediated text is far more homogeneous than human text (45% similarity on summarisation against 27% for human summaries), and sanitisation "truncates precisely the distributional tails"
Ong's review of LLM research paradigms (2024) Separated GPT-ology, LLMs-as-computational-models and silicon sampling, and examined the inference each supports The paradigms are distinguishable and their claims assessable on their own terms Closed- versus open-source models, invisible training data and no conventions for prompt "hyperparameters" leave reproducibility unsettled

None of these are SynthFolk results, and none licenses a general "synthetic respondents are X% accurate" claim — accuracy is conditional on question type, market, model and grounding. That conditionality is a marketing-science norm, not a synthetic-research excuse: Wind and Sharp, surveying advertising's empirical generalisations, note that even well-established laws suffer from inadequate knowledge of the conditions under which they generalise. Sarstedt and colleagues land where this article does: silicon samples "hold particular promise in upstream parts of the research process such as qualitative pretesting and pilot studies", much less so in main studies.

Where synthetic respondents match real ones

Against the directional standard, with reasonable grounding:

Where they do not

Against the distribution standard, the failures are predictable in advance:

What the critics get right

The most-cited pieces here are critical, and on mechanism they are right.

Rosala and Moran at the Nielsen Norman Group argue that AI-generated "users" cannot substitute for real ones, conceding one narrow use: rehearsing a discussion guide in an unfamiliar domain before meeting real participants. Papangelis, in ACM Interactions, names the "synthetic persona fallacy" — the slide from plausible artefact to stand-in for a constituency — plus a bias-laundering problem in which the training data's dominant voices come back as the population. Chapman on the Quant UX Blog notes that survey statistics do not apply to samples drawn from a model, and MeasuringU's review of published experiments agrees: hypothesis generation, not final decisions.

Not all resistance is about validity — Castelo, Bos and Lehmann show algorithms are trusted less for tasks that merely seem subjective, whatever their performance, and qualitative insight is the archetype — so scepticism and evidence must be argued separately. Four criticisms, though, are structural rather than teething problems.

1. Framing dependence. Models answer the question they are handed, in the terms they are handed it. Ask "would you use this feature?" and the base rate of yes is a property of the prompt. Toubia's group cites work showing answers can be dominated by prompt architecture — option labelling and ordering alone. Lin's case is starker: one instruction never to answer negatively about a country moved a synthetic sample's identification of America's primary adversary from 86% China to 88% Russia.

2. Distribution flattening. Synthetic samples under-disperse. In a replication effort the Sarstedt review covers, six of fourteen studies showed a "correct answer effect" — the model answering in a highly uniform way with none or almost no variation. Lin's homogeneity and tail-truncation findings show the same shape from the other side, and Valenzuela and co-authors name the mechanism: parametric reductionism, representing a person by a few parameters and losing all they do not carry. Any statistic depending on spread is invalid.

3. Hallucinated specificity. Asked for detail, models supply it: a brand they "always" buy, a price they paid. Well-formed and invented. The sharpest evidence comes from clinical screening. Fernandez, Berner and Shevlin generated 2,106 synthetic personas from 13 diagnostic profiles and put them through seven validated psychiatric instruments. Given only brief descriptions, and no instrument content in the prompts, the model produced diagnosis-congruent scores above clinical cutoffs, scaling monotonically with severity. Coherent responding above threshold, they conclude, "can no longer serve as a proxy for authentic participation". The danger is not fiction with bad grammar. It is fiction that passes the validated instrument.

4. An uneven default opinion profile. Training corpora over-represent some populations, as Papangelis argues, and demographic prompting shifts the default only partially. Sun's team measures it: the method that reproduces national polls well loses fidelity unevenly across subgroups and topics, tracking biases inherent in the model. The familiar shorthand — Western, educated, industrialised, rich, democratic — gives the direction; the problem is that output never says which cells are unreliable.

What actually determines synthetic-respondent quality

The decisive lever is not model choice. It is what the persona is conditioned on.

An ungrounded prompt describing "a 34-year-old teacher in Lyon" returns the model's prior about French teachers. Ten of them agree, because they are one distribution sampled ten times, and most bad experiences with synthetic research are experiences with that setup. Sun's result is the constructive half: feeding a model the distribution of a group rather than a caricature is what moved the synthetic answers close to real polls. Unanimity is a report about the model, not the market.

SynthFolk conditions personas on the European Social Survey and Eurobarometer — large, documented probability surveys — plus demographic mapping and five European cultural regions, so a Polish and a Portuguese respondent differ on measured dimensions, not stereotype. What that fixes:

Failure mode Does grounding address it?
Distribution flattening Partly. It restores between-persona variance; it does not make dispersion statistics valid.
Uneven default opinion profile Partly, within Europe. Outside ESS and Eurobarometer coverage the underlying model skew is unchanged.
Framing dependence No. A model-behaviour property, managed by question design.
Hallucinated specificity No. Grounding constrains attitudes, not invented anecdotes.
Thin data on rare populations, post-cutoff events No — grounding makes the limit explicit rather than removing it.

Three "no"s and two "partly"s. Grounding is the difference between a method that is sometimes useful and one that never is — not the difference between simulation and measurement.

Where synthetic research is fit for purpose today

Use it for Do not use it for
Early exploration and hypothesis generation Final go / no-go decisions
Concept and message screening — twenty down to three Pricing research and willingness-to-pay
Piloting an interview guide or questionnaire Sensitive, stigmatised or low-incidence populations
Creative and ad pretests, where output is a ranking Regulated claim substantiation and compliance
Cross-market first passes, to place fieldwork budget Any figure destined for a board deck as an estimate
Objection mapping and stakeholder alignment Usability, observed behaviour, lived experience

The left column is the Sarstedt review's own recommendation — upstream pretesting and piloting — and its output is comparative, so being wrong redirects a next step rather than shipping a decision. SynthFolk's quantitative module is built for that column: 50 to 250 respondents, read as comparisons, not estimates.

How to validate synthetic results on your own studies

Category-level benchmarks tell you little about your category. Golder and colleagues' case for empirics-first research applies here: work grounded in a real phenomenon, using data, producing valid insight without waiting on theory.

1. Back-test a study you already ran. Take real research from the last twelve months and re-run its central question synthetically, blind to the result if you can. Compare the ranking, not the numbers. Directional agreement on studies you trust is your actual accuracy rate.

2. Run one parallel human study, against the right baseline. For a single live decision, field both: the synthetic study and a small real panel on the same quotas. The baseline makes or breaks this. Sharp, Danenberg and Bellman's PNAS letter is the cautionary case: re-reading a celebrated psychological-targeting result, they show the original tested targeted ads against deliberately mis-targeted ads rather than against untargeted advertising, that targeting won only two of five experiments — about what chance predicts — and that creative quality went uncontrolled. A weak baseline manufactures a result.

3. Test–retest the synthetic side. Re-run the identical study, then again with questions reordered and paraphrased. This is the Twin-2K-500 logic turned on your own work: instability bounds the resolution the method has for your question. If A and B swap when reordered, the difference was never real. Record the model version — on Ong's point, a re-run on a new version is a new study.

4. Check subgroups and spread. Do demographic cells differ the way population data says they should? Do distributions have tails? Collapsed subgroups and missing tails mean the panel is not doing the work you think.

5. Confirm the decision with humans. Whatever survives goes to real users before it ships. Not optional, and cheaper than before: the synthetic pass removed the wrong candidates and fixed the broken questions.

Start with steps 1 and 3. They cost a handful of credits — pricing is published, unlike most of this market — and tell you more than any vendor benchmark.

Our own benchmark (forthcoming)

We are running a head-to-head study: one five-question survey fielded twice, through SynthFolk's quantitative module and a real panel on matching quotas in one European market. Questions are pre-registered and deliberately include a type the literature says should fail — price sensitivity. We will publish distribution deltas, rank-order correlations, directional-agreement rates, subgroup deltas, both cost figures and the raw data.

It does not exist yet. When it does, it will be linked here. Until then, treat every claim on this page as reasoning from the literature listed below and our operating experience, not as a measured result.

The honest conclusion

Synthetic research is a complement, not a replacement — a phrase worn thin, so concretely: it changes the economics of questions you were never going to fund, and sharpens the research you do fund by removing wrong turns before you pay for fieldwork. It does not produce evidence about people. It produces a fast, cheap prior about what people might say, which you then check.

A team that replaces its research function with a synthetic panel has not saved money; it has bought an articulate machine for confirming what it already believed. A team that runs synthetic first and human second can afford better human research.

Frequently asked questions

Are synthetic users accurate enough for a real decision?

For decisions that narrow a set — which three concepts advance, which market to research first — yes, with the protocol above. For decisions committing budget, headcount or a launch, no.

Can synthetic respondents replace a panel?

No, and the framing is the problem. They replace the studies you were not going to run at all. Where a panel is genuinely required — anything measured, regulated, sensitive or behavioural — there is no substitute.

What is the error margin?

There is none, and any vendor quoting one is misusing the term. Margin of error is a property of probability sampling; synthetic respondents have no sampling frame. What you can measure is agreement with your past results and stability across re-runs — steps 1 and 3.

Which question types should I never ask?

Exact prices and willingness-to-pay, unaided brand or ad recall, frequency of a past behaviour, anything after the training cutoff, anything stigmatised, and any answer you intend to quote as a percentage. Behavioural frequency is the clearest case: Tanusondjaja and colleagues established Pareto ratios for a department-store chain from over 550 million transactions. No prompt substitutes for that.

How does this compare to a low-quality human panel?

The fairest comparison, and the one changing fastest. Cheap panels carry inattentive respondents and fraud, which more responses partly cure; synthetic panels add correlated bias, which more responses worsen by tightening confidence around the wrong answer. But the two are converging from the human side. Lin reports that roughly 34% of Prolific respondents say they use AI on open-ended questions, rising to 73% on Mechanical Turk for some task types, and that an autonomous agent passed 99.8% of 6,000 attention checks built to catch careless humans. "Real panel" is no longer a clean baseline.


Run the validation protocol on your own question. Take a study you already have the answer to and re-run it synthetically — a small qualitative study costs about ten credits, and new accounts start with enough. Open the dashboard and back-test one question. Whether the panel reproduces what you already know is the only benchmark that should move you.

References

Castelo, N., Bos, M. W., & Lehmann, D. R. (2019). Task-Dependent Algorithm Aversion. Journal of Marketing Research. DOI: 10.1177/0022243719851788.

Chapman, C. Synthetic Survey Data? It's Not Data. Quant UX Blog. https://quantuxblog.com/synthetic-survey-data-its-not-data

European Social Survey. https://www.europeansocialsurvey.org/

Eurobarometer, European Commission. https://europa.eu/eurobarometer/

Fernandez, K., Berner, L. A., & Shevlin, B. R. K. The threat of synthetic respondents extends to clinical mental health screening. Submitted manuscript, University of California, Los Angeles, and Icahn School of Medicine at Mount Sinai.

Golder, P. N., Dekimpe, M. G., An, J. T., van Heerde, H. J., Kim, D. S. U., & Alba, J. W. (2023). Learning from Data: An Empirics-First Approach to Relevant Knowledge Generation. Journal of Marketing, 87(3), 319–336. DOI: 10.1177/00222429221129200.

Lin, Z. Synthetic respondents and the illusion of human data. Preprint, Department of Psychology, Yonsei University.

Macdonald, E. K., & Sharp, B. M. (2000). Brand Awareness Effects on Consumer Decision Making for a Common, Repeat Purchase Product: A Replication. Journal of Business Research, 48, 5–15. DOI: 10.1016/S0148-2963(98)00070-8.

MeasuringU. A Review of Experiments with Synthetic Users. https://measuringu.com/review-of-experiments-with-synthetic-users/

Ong, D. C. (2024). GPT-ology, Computational Models, Silicon Sampling: How should we think about LLMs in Cognitive Science? arXiv:2406.09464.

Papangelis, K. The Synthetic Persona Fallacy: How AI-Generated Research Undermines UX Research. ACM Interactions. https://interactions.acm.org/blog/view/the-synthetic-persona-fallacy-how-ai-generated-research-undermines-ux-research

Romaniuk, J., & Sharp, B. (2004). Conceptualizing and measuring brand salience. Marketing Theory, 4(4), 327–342. DOI: 10.1177/1470593104047643.

Rosala, M., & Moran, K. (2024). Synthetic Users: If, When, and How to Use AI-Generated "Research". Nielsen Norman Group, 21 June 2024. https://www.nngroup.com/articles/synthetic-users/

Sarstedt, M., Adler, S. J., Rau, L., & Schmitt, B. (2024). Using large language models to generate silicon samples in consumer and marketing research: Challenges, opportunities, and guidelines. Psychology & Marketing, 41, 1254–1270. DOI: 10.1002/mar.21982.

Sharp, B., Danenberg, N., & Bellman, S. (2018). Psychological Targeting. Letter, Proceedings of the National Academy of Sciences. DOI: 10.1073/pnas.1810436115.

Sun, S., Lee, E., Nan, D., Zhao, X., Lee, W., Jansen, B. J., & Kim, J. H. (2024). Random Silicon Sampling: Simulating Human Sub-Population Opinion Using a Large Language Model Based on Group-Level Demographic Information. arXiv:2402.18144.

Tanusondjaja, A., Romaniuk, J., Nenycz-Thiel, M., Sakashita, M., & Viswanathan, V. (2023). Examining Pareto Law across department store shoppers. International Journal of Market Research. DOI: 10.1177/14707853221145851.

Toubia, O., Gui, G. Z., Peng, T., Merlau, D. J., Li, A., & Chen, H. (2025). Twin-2K-500: A dataset for building digital twins of over 2,000 people based on their answers to over 500 questions. arXiv:2505.17479.

Valenzuela, A., Puntoni, S., Hoffman, D., Castelo, N., De Freitas, J., Dietvorst, B., Hildebrand, C., Huh, Y. E., Meyer, R., Sweeney, M. E., Talaifar, S., Tomaino, G., & Wertenbroch, K. (2024). How Artificial Intelligence Constrains the Human Experience. Journal of the Association for Consumer Research, 9(3). DOI: 10.1086/730709.

Wind, Y. (J.), & Sharp, B. (2009). Advertising Empirical Generalizations: Implications for Research and Action. Journal of Advertising Research, 49(2). DOI: 10.2501/S0021849909090369.