How to Validate Synthetic Research Against Human Evidence
Validate synthetic research by comparing a predefined output with relevant human evidence that the generating process did not receive. Also check a simple baseline and repeat the simulation. Agreement with people and stability across runs answer different questions; neither should be replaced by a convincing transcript or a single overall accuracy claim.
Our review of synthetic respondent accuracy covers the broader debate. This guide proposes an applied protocol for a specific task. It is not a completed SynthFolk benchmark or a guarantee that any task will pass.
1. Define the target and the decision
Specify exactly what the simulation should reproduce: a ranking of concepts, a response distribution, particular objections or individual answers. Do not use the same metric for all four. A model could identify the most preferred option while getting the size of the preference wrong.
State what a useful result would let you do. If the purpose is questionnaire preparation, the useful outcome may be identifying ambiguities subsequently observed in a human pretest. If the purpose is predicting an aggregate distribution, the comparison must measure distributional error. Those tasks require different reference material.
Choose acceptable error before examining the output. The threshold should reflect the decision: reversing the top two concepts may matter more than a small discrepancy among concepts already rejected. There is no universal numerical threshold that makes synthetic research valid for every use.
2. Separate allowed inputs from evaluation answers
Prepare the brief, audience definition, stimulus and questions. Keep the human answers used for scoring out of the material supplied to the simulation. If you use a past study to improve prompts, reserve another study or part of the evidence for evaluation.
The distinction is explicit in the consulted version of Twin-2K-500 by Toubia and colleagues (2025). It separates persona information, held-out evaluation answers and later human retest answers. Its dataset and digital-twin task differ from a generic commercial simulation, but the separation clarifies what an independent check requires.
Consider possible training exposure as well. A widely published survey result may have been present in a model's training material. You cannot always exclude that possibility for a proprietary model. Document it and prefer a suitable private or newly collected reference where permitted, rather than claiming that any historical reproduction proves generalisation.
3. Check the human reference
Human data are a reference, not automatically flawless ground truth. Inspect eligibility, recruitment, response quality, exclusions, instrument wording and the period of collection. A survey of one customer group cannot validate a prediction for a different population without additional assumptions.
Make conditions comparable. If people saw a visual stimulus and the model received only a researcher-written summary, the comparison changes both respondent type and input. That may be a useful workflow comparison, but it should be labelled as such.
Keep uncertainty in the human estimate visible. If two human estimates are imprecise, small differences between them and the model may not support a strong conclusion. Conversely, a low-quality human panel should not be used to excuse arbitrary model error.
4. Compare against a simple baseline
A complex simulation must add something beyond a simpler alternative. For categorical answers, a baseline might predict the most frequent category in separate training data. For a ranking task, it could use an established historical ranking where that is relevant. For instrument preparation, it might be an analyst's checklist review.
Choose the baseline using information available at prediction time. A baseline tuned to the evaluation answers is no longer a fair reference. Do not deliberately weaken it to make the simulation look useful.
Record both absolute performance and improvement over the baseline. If both choose the same winner, the simulation may add explanations, but those explanations need their own assessment. Correct classification does not establish that the generated reasoning describes how people decided.
5. Use a small set of interpretable measures
| Target | Useful comparison | Important limitation |
|---|---|---|
| Categorical distribution | Percentage-point differences for every option | Aggregate agreement can hide subgroup errors |
| Ranking | Order agreement and whether the leading option changes | Correlation can conceal a consequential reversal near the top |
| Open-ended concerns | Human review of matched, missed and unsupported themes | A fluent explanation can still be unsupported |
| Individual answers | Held-out prediction against a defined baseline | Requires individual reference data and a suitable task |
| Repeatability | Variation across repeated model runs | Measures model stability, not human validity |
Do not collapse these measures into an unexplained score. Keep the original outputs so readers can inspect the errors. For themes, distinguish a plausible additional hypothesis from an invented claim about a participant. The former may help plan research; the latter is not an observation.
Sun and colleagues (2024) found that demographic-conditioned simulation varied by subgroup and question in their opinion-survey application. That supports looking beyond one aggregate metric. It does not prescribe a universal subgroup test or establish performance in your category.
6. Test stability without selecting a favourable run
Repeat the same configuration and retain every run included in the evaluation. Then make planned changes to wording or order that should not fundamentally change the task. Record whether conclusions move. Keep this sensitivity analysis distinct from the initial performance estimate.
Ong (2024) highlights the importance of reporting prompts, procedures, model settings and access dates. Apply that principle to the available controls. If a provider does not expose a setting or silently changes a model, document the limitation rather than inventing a value.
Do not tune until the expected answer appears and then report that run as a successful replication. Tuning is development. It needs a subsequent evaluation against material not used to tune.
Human test–retest agreement is another distinct quantity. In Twin-2K-500, people answer some tasks again to establish a human reference for predictability. Repeating an AI survey does not supply that human reference, and there is no automatic justification for dividing one agreement number by another without a clearly defined measure.
7. Turn the result into an explicit use boundary
Suppose, in an invented example, a simulation reproduces the overall winner in a packaging comparison but misses the winner among occasional category buyers. If that subgroup matters to the launch, the overall match is insufficient. The next action might be further human research or rejecting the simulation for that decision.
Alternatively, suppose a synthetic rehearsal repeatedly identifies a confusing eligibility question that people also misinterpret in a pretest. That is evidence of usefulness for instrument preparation under those conditions. It is not evidence that the same system predicts purchase rates.
Write the conclusion narrowly: task, population, stimulus, configuration, reference and observed failures. List uses the evaluation did not cover. Sarstedt and colleagues (2024) review domain-dependent outcomes and discuss appropriate roles for silicon samples; their review does not remove the need for this task-specific judgement.
Begin with one checkable task. Prepare the brief, retain the human reference separately, and use the appropriate SynthFolk workflow. Review pricing and the sample-size guide before planning repeated runs. This protocol itself does not establish that SynthFolk has passed any external benchmark.
Sources and editorial method
The seven-step protocol and packaging scenario are proposed applications. The cited research supplies methodological distinctions; no benchmark results are claimed here. Preprint references identify the versions consulted from the local bibliography.
- Toubia, O., et al. (2025). Twin-2K-500: A dataset for building digital twins of over 2,000 people based on their answers to over 500 questions. arXiv:2505.17479, version 1.
- Sun, S., et al. (2024). Random Silicon Sampling: Simulating Human Sub-Population Opinion Using a Large Language Model Based on Group-Level Demographic Information. arXiv:2402.18144, version 1.
- Ong, D. C. (2024). GPT-ology, Computational Models, Silicon Sampling: How should we think about LLMs in Cognitive Science? arXiv:2406.09464, version 1.
- Sarstedt, M., Adler, S. J., Rau, L., & Schmitt, B. (2024). Using large language models to generate silicon samples in consumer and marketing research: Challenges, opportunities, and guidelines. Psychology & Marketing. DOI: 10.1002/mar.21982.