Synthetic Survey Sample Size: Why More Responses Do Not Prove Accuracy

· 7 min read

A synthetic survey's sample size counts generated responses. Increasing that count may stabilise an estimate of what a configured model produces, but it does not automatically improve accuracy about people. Choose the number of runs according to the task and the uncertainty being examined, rather than applying a human-survey margin-of-error rule to generated rows.

This distinction complements our review of synthetic respondent accuracy. The practical issue is how to spend a simulation budget without confusing repetition with additional evidence about the market.

Ask what varies when you generate another answer

In a human survey, an additional participant can contribute another person's response, subject to the sampling design and data quality. In a synthetic survey, another row is produced by the generation process. Its variation may arise from different persona inputs, random model sampling, prompt differences or changes in the model itself.

These sources of variation are not equivalent. Repeating one persona's prompt explores a different quantity from changing the audience profile. Changing both the persona and the model makes it harder to identify what caused a difference in the answer.

Write down the unit you intend to study. It could be a generated response under a fixed configuration, a persona scenario or an entire repeated simulation. Do not call each unit an independent consumer simply because it has a unique identifier.

Separate model variation from error about people

Imagine a generator that tends to select option A about 60% of the time under one fixed setup. More draws can help estimate that tendency. If the corresponding proportion among relevant people is 40%, a stable estimate near 60% remains wrong for the human question.

The example is invented to explain the distinction. It is not an estimate of SynthFolk's behaviour or a claim about a particular model. The general issue is that reducing random variation does not necessarily remove a systematic difference between the generating process and the target population.

For an idealised set of independent binary draws with probability 0.60, the standard error of their mean is sqrt(0.60 × 0.40 / n). It is about 4.9 percentage points at 100 draws and 1.5 points at 1,000. These figures describe that mathematical sampling model. They are not margins of error for customers represented by AI personas, and independence cannot simply be assumed for a real synthetic workflow.

The fact that a spreadsheet permits the calculation does not establish that its assumptions describe the research process.

Why a larger synthetic sample may still omit important people

Increasing the number of generated responses does not add a missing buying situation to the brief. If every persona assumes easy internet access, generating more of them will not reveal the experience of people without it unless the process contains a way to represent that difference.

Sun and colleagues (2024) examined simulations conditioned on demographic distributions and found variation across questions and demographic groups. Their application involved US opinion data, and the consulted preprint notes limitations including possible training exposure. It does not establish that a larger simulation can overcome an unsuitable audience representation in another domain.

Check whether the audience distinctions are relevant to your question before increasing the count. A quota on age alone may do little for a question driven by product access, category experience or household constraints. Filling the chosen quotas shows that the configuration followed them; it does not validate the generated answers.

Allocate runs to the uncertainty you need to inspect

Research task Useful reason for additional runs What additional runs cannot establish alone
Rehearsing a questionnaire Look for more possible ambiguities and answer paths How often people will misunderstand an item
Checking a simulated ranking Observe whether the ordering changes under the fixed setup Which option real customers prefer
Examining audience assumptions Compare explicitly defined scenarios The scenarios' real population frequencies
Testing wording sensitivity Compare planned prompt or order variants Which wording produces the most truthful human estimate
Assessing external validity Repeat a predefined comparison with human evidence Generalisation to untested markets or tasks

Use the smallest initial run that lets you inspect the material meaningfully. Increase it when the extra output addresses a named uncertainty. This is a resource-allocation proposal, not a universal recommended sample size.

If your main question is whether a ranking corresponds to people, another well-designed human comparison may be more informative than a large increase in generated rows. If your main question is whether the simulation itself is unstable, repeated model runs are directly relevant.

A worked example with a fixed simulation budget

Suppose a team has a budget for several exploratory runs of three service descriptions. It could spend everything generating many responses to one brief, or reserve runs for repeated configurations and plausible audience alternatives.

The team first fixes the descriptions, questions and audience brief. It repeats that setup to see whether the leading option is stable. It then changes one audience assumption, such as whether the service is accessible during working hours, and examines what changes. These are scenario differences, not measured market segments.

If the preferred option reverses when the wording changes slightly, the team records that instability. It does not select the run that supports the desired launch. If the ranking remains stable, the next question is still whether it agrees with relevant people.

The example is hypothetical. No particular number of runs guarantees a useful result, and the platform's available controls determine which comparisons are possible. The point is to name the question each additional run answers.

Report the configuration alongside the count

At minimum, record the audience brief, stimulus, exact questions, response options, date and available model information. Distinguish the number of persona profiles from repeated responses and repeated whole studies. Record exclusions and failed runs rather than presenting only completed answers.

Ong (2024) discusses reproducibility problems and the need for detailed prompts, procedures and settings. If a tool does not expose a parameter, mark it unavailable. An invented temperature setting is worse than an explicit limitation.

Report findings in language that matches the evidence: “The ranking was unchanged across the recorded runs” describes model stability. “The ranking matched the held-out human study under the specified conditions” describes a particular external comparison. Neither alone establishes universal reliability.

Set a stopping rule before interpreting the output

Decide when another run would no longer change the next research action. You might stop questionnaire rehearsal when repeated outputs no longer add a new issue worth checking, while documenting that this is an editorial stopping rule, not proof of complete coverage.

For a validation exercise, specify the runs and comparisons beforehand and retain failures. Do not keep expanding the sample until a preferred result appears. Sarstedt and colleagues (2024) place silicon sampling within a broader research process; a larger synthetic count does not erase the boundaries of that use.

Budget for the uncertainty, not the row count. Use the brief template, inspect current pricing and plan the quantitative workflow. For claims about people, add the external validation step your decision requires.

Sources and editorial method

The probability calculation is an explicitly idealised mathematical example. The service scenarios and stopping suggestions are original applications, not observed results or validated sample-size recommendations for SynthFolk.