Synthetic respondents
A synthetic respondent is an AI-simulated survey participant: one row in a quantitative dataset, generated by a language model conditioned on a demographic and attitudinal profile instead of collected from a person. Where a synthetic persona is a richly specified character built for long qualitative conversations, a synthetic respondent is a thinner unit run at scale — tens to hundreds at a time — because the output you want is a distribution rather than a narrative.
How accurate are they?
The question is unanswerable as posed, because accurate silently means three incompatible things.
Distribution match asks whether the synthetic sample reproduces the population's answer distribution — the same share picking option B, the same mean on a scale. This is the strictest standard and the one synthetic respondents fail. Directional agreement asks whether they pick the same winner and order options the same way. This is weaker, it is what most product decisions actually rest on, and with reasonable grounding it holds up. Individual prediction asks whether an agent built from one person's data can reproduce that person's answers; it is an active research question and largely irrelevant commercially.
A vendor claiming high accuracy usually means directional agreement. A critic reporting failure usually means distribution match. Both can be right at once.
The documented failure modes
Four are structural rather than teething problems. Models trained as assistants agree with the framing they are handed, so a question phrased as a proposal gets an inflated yes rate. Synthetic samples under-disperse, which invalidates any statistic depending on spread — significance tests, confidence intervals, segment separation. Asked for detail, models supply invented specifics that read exactly like recalled experience. And training corpora skew Western, educated and higher-income, which demographic prompting shifts only partially.
There is no margin of error for a synthetic sample. Margin of error is a property of probability sampling, and synthetic respondents have no sampling frame. What you can measure instead is agreement with your own past studies and stability across re-runs — which is the only benchmark that should change how much weight you give the method.
Read the full guide: How Accurate Are Synthetic Respondents? What the Evidence Actually Shows →