Synthetic Users Explained: What They Are, What They Can and Can't Do
Synthetic users are AI-generated research participants: language-model personas defined by a demographic and attitudinal profile, asked the questions you would ask a real panel. They are not people, and they do not produce opinions. They produce a fast, cheap, directional simulation of how a described population might respond — useful for some research questions and actively misleading for others. This page explains which is which.
Key takeaways
- Synthetic users are simulated respondents generated by a large language model and conditioned on a described population. Nothing they say originates from a person.
- They are strongest on comparative and directional questions, weakest on absolute numbers, novel behaviour and specialist populations. Published review evidence draws the same line: agreement between silicon and human samples "var[ies] considerably across different domains", with the promise upstream, in pretesting and pilot studies (Sarstedt et al., 2024).
- Quality is set mostly by grounding data, not model size: group-level demographic conditioning alone can reproduce real poll distributions closely, though replicability varies by group and topic (Sun et al., 2024).
- They are a complement to real research, not a replacement — marketing scholarship reaches the same conclusion about AI generally: it augments human managers better than it replaces them (Davenport et al., 2020).
What are synthetic users?
Synthetic users are artificial research participants created by prompting a large language model to answer in character as a member of a defined population. Each carries a profile — age, market, income, values, attitudes, media habits — and answers interview questions, survey items or creative stimuli as that profile plausibly would. The literature calls the same construct a silicon sample (Sarstedt et al., 2024). The output looks like research data. It is simulation.
That distinction is the whole subject. A synthetic user has no memory of buying anything — only a statistical model of how text about people like that tends to read. Sometimes that is a good enough proxy to decide with; often it is not. The term is used loosely across the category, so it helps to fix the vocabulary first.
Synthetic users vs. synthetic personas vs. synthetic respondents
Treated as synonyms, they are better understood as one thing at three stages.
| Term | What it names | Where you meet it | Typical output |
|---|---|---|---|
| Synthetic persona | The profile of a simulated individual | Study setup | A character sheet: who this person is, what they value |
| Synthetic user | The persona in a session — answering, reacting to a stimulus | Qualitative interviews, concept tests, creative evaluation | Transcripts, verbatims, themes |
| Synthetic respondent | The persona in a survey — one row in a dataset | Surveys of tens to hundreds | Distributions, cross-tabs, rank orders |
A fourth term, digital twin, names a simulated model of one specific real person built from their own data — a different technique with far heavier requirements: Toubia et al. (2025) surveyed 2,058 people across 500 questions, 2.42 hours each, simply to build a public testbed for individual-level twins. It is not the population-level sampling described here.
For how the profiles are constructed, see the guide to synthetic personas. This page stays with what happens once one starts answering.
How synthetic users are generated
Four stages, and the second is where most of the quality difference between tools lives.
1. Population definition. Describe who you want to hear from — market, age range, income, and the traits relevant to the question: the brief you would hand a recruitment agency.
2. Grounding and sampling. The system builds profiles matching that brief. Asking the model to invent them is fluent and has a specific defect: of 14 psychological replications reviewed by Sarstedt et al. (2024), six showed a "correct answer effect" — the model answering in a highly uniform way with almost no variation. The alternative is sampling attributes from real population data. Sun et al. (2024) show why that is the lever: conditioned only on group-level demographic distributions, a model produced response distributions "remarkably similar to the actual U.S. public opinion polls". SynthFolk conditions personas on the European Social Survey and Eurobarometer across five cultural regions, so a persona's institutional trust or price sensitivity reflects a measured distribution rather than a model's impression of one.
The rule that follows is ours rather than a published finding, but it falls out of the mechanism: disagreement is the signal. Unanimity reports the model's default, not the market's spread.
3. Independent querying. Each synthetic user is generated and asked separately, not role-played together inside one context window. LLM-mediated text is already more homogeneous than human text — 45% similarity on summarisation tasks against 27% for human summaries (Lin) — so a shared context only compounds the pull toward the mean. The group format is also a weak instrument with real people: across 227 marketers' most recent pack redesigns, Caruso et al. (2025) found focus groups among the methods associated with less successful outcomes, while research identifying which design elements consumers already link to the brand was associated with more successful ones — an association in one category rather than a law, but a poor case for simulating the group discussion.
4. Aggregation. Themes and verbatims for qualitative work, distributions and breakdowns for quantitative, comparative scoring for A/B stimuli.
A consequence of stage two, as reasoning rather than measurement: a smaller, better-grounded panel should beat a larger ungrounded one. More respondents from the same collapsed distribution add confidence, not information.
SynthFolk's qualitative module turns a plain-language audience description into 3–8 grounded personas.
What synthetic user research is used for
Synthetic research earns its place where value lies in breadth, speed and comparison rather than in discovery. Sarstedt et al. (2024) locate its promise in "upstream parts of the research process such as qualitative pretesting and pilot studies", and assess main-study use far more critically.
Concept screening. Eleven positioning statements, budget to test two. A synthetic panel gives you an ordering to argue with — enough to decide what enters real testing, a call most teams currently settle by argument in a meeting room.
Message and creative comparison. A/B comparisons of headlines, value propositions or ad creative suit the method better: a relative judgement is less exposed to the calibration problem than an absolute one. What it replaces is not the real test but the shortlist for it — a properly powered stimulus study looks like Caruso et al. (2026), where 484 US and 491 UK consumers evaluated 48 redesigned packs. SynthFolk's media module runs that screening pass as an A/B mode.
Survey and interview piloting. Run your instrument against synthetic respondents before spending recruitment money. Ambiguous questions break synthetic answers as they break human ones, and you will find the bad items in an afternoon.
Cross-market first passes. A dozen markets is prohibitive with real panels and trivial with synthetic ones. The per-market numbers are not defensible — Sun et al. (2024) find replicability varies by demographic group and topic, attributing it to societal biases in the models — so read the divergence pattern as a hypothesis about where fieldwork is needed.
Objection mapping. Asking a diverse panel what would stop them buying surfaces an objection inventory fast — not how common each is, only that none blindsides you later.
Every one is a filtering task upstream of a real decision, not the decision itself — which also sets the boundary with the broader AI market research category, where AI-assisted analysis and AI-moderated interviews with real people get conflated with fully synthetic work.
The case against synthetic users
The most-read pages on this topic are not written by vendors. Nielsen Norman Group's critique of synthetic users (2024) argues they cannot replace what studying real people teaches, and that they often return shallow or overly favourable feedback. Papangelis (2025), in ACM Interactions, makes the related "synthetic persona fallacy" argument: marketing statistical pattern-matchers as simulated cognition borrows the authority of research while abandoning its standards. Quantitative UX practitioners put it bluntest — simulated responses are not data, because nothing was measured.
Three parts of that critique are correct, and no amount of grounding fixes them.
- There is no measurement. A synthetic distribution is generated, not observed; reporting it with the confidence-interval conventions of survey research is a category error. Nor can coherence stand in: Lin documents autonomous AI agents passing 99.8% of 6,000 attention checks while producing psychometrically sound data indistinguishable from careful human work.
- Models have systematic response biases. Sun et al. (2024) tie subgroup replicability failures to societal biases in the models. Lin is sharper: sanitisation "truncates precisely the distributional tails — expressions of prejudice, ambivalence, extreme views". Valenzuela et al. (2024) name the mechanism parametric reductionism — represent a person by a few parameters and you lose whatever they do not carry.
- Substitution is a real institutional risk. The failure mode is not a bad study but a research budget quietly replaced by a subscription, and a team that has stopped talking to anyone.
The honest response is not to rebut this but to design around it. How well grounded personas track real respondents is treated in how accurate synthetic respondents are.
When to use real users instead
| Question type | Use synthetic | Use real users |
|---|---|---|
| Which of these 10 concepts is worth testing? | ✅ Screening pass | Final validation |
| Is this headline clearer than that one? | ✅ Directional A/B | Confirm the winner |
| What percentage will buy at €29? | ❌ No calibration | ✅ Required |
| How do people actually complete this task? | ❌ No behaviour | ✅ Usability testing |
| What do users of a brand-new category want? | ❌ No grounding data | ✅ Discovery |
| How do 40 paediatric oncology nurses rate this? | ❌ Data too thin | ✅ Required |
| Does this claim hold up for regulators? | ❌ Never | ✅ Required |
Five hard limits sit behind that table:
Absolute numbers. A synthetic panel reporting 63% purchase intent produces a number with no calibration to reality. How far synthetic rows move an estimate is not hypothetical: in Lin's simulations of 2024 US presidential polls, inserting 10–52 synthetic respondents into samples of 1,600 flipped which candidate led. Read distributions comparatively, never as a forecast.
Genuinely novel products. Grounding data predates the model, and the training data is not inspectable — an open problem Ong (2024) flags for research using LLMs as stand-ins for people. Where no attitudinal distribution exists the model interpolates confidently; the further from established behaviour, the less the output means.
Underrepresented and specialist populations. Grounding data is thinnest exactly where accuracy matters most, and the resulting fiction is convincing — the dangerous part. Fernandez et al. generated 2,106 synthetic personas from 13 DSM-informed diagnostic profiles and ran seven validated psychiatric screening instruments; from brief descriptions alone, the model produced clinically differentiated, severity-sensitive scores. Coherent, threshold-crossing responding is no longer a proxy for authentic participation. Plausibility is not your check.
Lived experience and observed behaviour. Synthetic users generate what a described person might say. They cannot show you the workaround someone invented or the hesitation before a click, and two classes of measurement are out of reach for the same reason. Behavioural frequency comes from behavioural records: Tanusondjaja et al. (2023) derive heavy-buyer distributions from over 550 million department-store transactions, not from asking. Memory-based constructs belong to a buyer's memory network in a buying situation rather than to text about a brand, which makes brand salience a measurement problem rather than a question for a model (Romaniuk & Sharp, 2004).
Regulatory, safety or legal exposure. If a wrong answer produces a recall, a fine or harm, use humans.
Which is the position: synthetic research is a complement, not a replacement. It changes the economics of questions you were never going to fund, and sharpens real research by clearing the obvious wrong turns first. A team that dismisses its research function and buys a synthetic panel has bought a way to confirm its own priors.
Synthetic users in practice
A first study should be small, comparative and checkable.
Setup. Qualitative module. Audience in plain language — grocery shoppers aged 30–50 in two European markets, mixed income. Three personas, the smallest panel that can disagree. From 6 credits, published on the pricing page.
Instrument. Four or five open questions on a decision you have already made, plus one probe you do not know the answer to.
What comes back. A transcript per persona, then a synthesis: recurring themes, where the personas split, the verbatims behind each. Read the split first. If all three gave the same answer in three registers, your question was leading or your audience description too narrow — both fixable in the next run, not after six weeks of fieldwork.
The check that matters. Pick a question you answered with real research last year, run it synthetically, and see whether the panel reproduces the finding. Two things make that a serious test. Running your own comparison on your own phenomenon generates real knowledge, not a lesser substitute for a general theory of accuracy (Golder et al., 2023). And the yardstick is not exact agreement: humans do not answer identically twice either, which is why Toubia et al. (2025) built a repeated wave into their study to establish a test–retest baseline. The human comparison set is also no longer clean — roughly 34% of Prolific respondents and up to 73% on MTurk report using AI on open-ended items (Lin).
Teams that get sustained value converge on one loop: diverge synthetically, pressure-test the instrument synthetically, validate with humans on the survivors, then check your calibration against what the real study found. Step four separates serious use from a rationalisation engine, and nobody can do it for you.
FAQ
Are synthetic users real people?
No. They are language-model simulations conditioned on population-level statistics: no individual represented, no personal data processed, no response given by a human. Any tool implying otherwise misrepresents the method.
Are synthetic users accurate?
There is no single accuracy number, and the review literature is explicit about why: agreement with human samples varies considerably across domains (Sarstedt et al., 2024), and replicability varies by group and topic even where aggregate distributions land close (Sun et al., 2024). Results also move with prompt wording and model version — unresolved "hyperparameters" of this work (Ong, 2024). Treat the output as directional and uncalibrated; the evidence is reviewed in our accuracy guide.
What do synthetic users cost?
SynthFolk prices in credits, published openly: a three-persona qualitative study starts at 6 credits; larger panels and quantitative samples of 50–250 respondents cost proportionally more. Compare against a recruited panel study for the same question: typically weeks, and four figures.
Do synthetic users replace usability testing?
No — the clearest boundary in the method. Usability testing measures what people do under observation; synthetic users generate what someone might say. There is no substitute for watching a person fail to find a button.
Which languages and markets are supported?
SynthFolk runs in 12 interface languages with personas grounded across five European cultural regions, which is what makes a cross-market first pass practical. Coverage tracks the underlying survey data — where that is thin, treat the output as thinner.
Run one yourself. A three-persona study starts at 6 credits, and new accounts begin with free credits — enough to test the method on a question you already know the answer to. Open the dashboard: whether the panel reproduces what you know is the benchmark that counts.
References
- Caruso, W., Romaniuk, J., Page, B., Anesbury, Z. W., & Williams, J. (2025). The role of market research in pack redesign performance. International Journal of Market Research, 67(1), 17–32. https://doi.org/10.1177/14707853241296656
- Caruso, W., Romaniuk, J., Page, B., Anesbury, Z. W., Saeed, R., & Williams, J. (2026). The packaging redesign modernisation dilemma: The relationship with familiarity, likeability, and its effect on purchase intent. Journal of Retailing and Consumer Services, 92, 104800. https://doi.org/10.1016/j.jretconser.2026.104800
- Davenport, T., Guha, A., Grewal, D., & Bressgott, T. (2020). How artificial intelligence will change the future of marketing. Journal of the Academy of Marketing Science, 48, 24–42. https://doi.org/10.1007/s11747-019-00696-0
- Eurobarometer. Public opinion surveys of the European Commission. https://europa.eu/eurobarometer/
- European Social Survey. https://www.europeansocialsurvey.org/
- Fernandez, K., Berner, L. A., & Shevlin, B. R. K. The threat of synthetic respondents extends to clinical mental health screening. Submitted manuscript, University of California, Los Angeles, and Icahn School of Medicine at Mount Sinai (undated).
- Golder, P. N., Dekimpe, M. G., An, J. T., van Heerde, H. J., Kim, D. S. U., & Alba, J. W. (2023). Learning from data: An empirics-first approach to relevant knowledge generation. Journal of Marketing, 87(3), 319–336. https://doi.org/10.1177/00222429221129200
- Lin, Z. Synthetic respondents and the illusion of human data. Preprint, Department of Psychology, Yonsei University (undated).
- Nielsen Norman Group. (2024, June 21). Synthetic users: If, when, and how to use AI-generated "research". https://www.nngroup.com/articles/synthetic-users/
- Ong, D. C. (2024). GPT-ology, computational models, silicon sampling: How should we think about LLMs in cognitive science? arXiv:2406.09464. https://arxiv.org/abs/2406.09464
- Papangelis, K. (2025, December 17). The synthetic persona fallacy: How AI-generated research undermines UX research. ACM Interactions (blog). https://interactions.acm.org/blog/view/the-synthetic-persona-fallacy-how-ai-generated-research-undermines-ux-research
- Romaniuk, J., & Sharp, B. (2004). Conceptualizing and measuring brand salience. Marketing Theory, 4(4), 327–342. https://doi.org/10.1177/1470593104047643
- Sarstedt, M., Adler, S. J., Rau, L., & Schmitt, B. (2024). Using large language models to generate silicon samples in consumer and marketing research: Challenges, opportunities, and guidelines. Psychology & Marketing, 41, 1254–1270. https://doi.org/10.1002/mar.21982
- Sun, S., Lee, E., Nan, D., Zhao, X., Lee, W., Jansen, B. J., & Kim, J. H. (2024). Random silicon sampling: Simulating human sub-population opinion using a large language model based on group-level demographic information. arXiv:2402.18144. https://arxiv.org/abs/2402.18144
- Tanusondjaja, A., Romaniuk, J., Nenycz-Thiel, M., Sakashita, M., & Viswanathan, V. (2023). Examining Pareto Law across department store shoppers. International Journal of Market Research. https://doi.org/10.1177/14707853221145851
- Toubia, O., Gui, G. Z., Peng, T., Merlau, D. J., Li, A., & Chen, H. (2025). Twin-2K-500: A dataset for building digital twins of over 2,000 people based on their answers to over 500 questions. arXiv:2505.17479. https://arxiv.org/abs/2505.17479
- Valenzuela, A., Puntoni, S., Hoffman, D., Castelo, N., De Freitas, J., Dietvorst, B., Hildebrand, C., Huh, Y. E., Meyer, R., Sweeney, M. E., Talaifar, S., Tomaino, G., & Wertenbroch, K. (2024). How artificial intelligence constrains the human experience. Journal of the Association for Consumer Research, 9(3). https://doi.org/10.1086/730709