AI Focus Groups: Human Discussion or Simulated Interviews?
“AI focus group” can mean a discussion among real people assisted by AI, separate human interviews conducted by an AI moderator, or a set of simulated participants. Establish which one a service provides before interpreting its findings. Only the first contains an observed group discussion; the third contains no human participants.
The distinction builds on our explanation of synthetic users. This guide focuses on the practical consequence: choosing a format that can provide the information you need, while using simulated interviews to prepare questions rather than manufacture testimony.
What makes a focus group a group?
In a human focus group, participants hear and respond to one another. A remark can trigger recognition, disagreement, a correction or a story someone would not otherwise have offered. That interaction is part of the material. Several independent interviews can answer related questions, but they do not observe the same process.
Similarly, asking a language model to write dialogue between five named personas does not observe five people's interaction. The dialogue is an output of the generating system. Giving each persona a separate conversation can change how the simulation works, but it does not establish independent human viewpoints.
This matters when the research question concerns social negotiation, shared terminology or how people express disagreement. A plausible synthetic conversation may help a researcher anticipate possibilities. It cannot show how an actual group handled the issue.
Identify the format before buying the service
| Format | Who supplies the answers? | What you can inspect | Key limitation |
|---|---|---|---|
| Human group with AI assistance | People interacting in a group | Recordings, transcripts and AI-produced analysis | Recruitment, moderation and interpretation still require scrutiny |
| AI-moderated individual interviews | People responding separately | Questions, follow-ups and individual responses | No observed interaction between participants |
| Independent synthetic interviews | A model conditioned on persona inputs | Briefs, generated responses and model information | No observation of those people's lived experience |
| Simulated group dialogue | A model or coordinated model process | Generated conversation and its setup | Apparent agreement or conflict is generated, not socially observed |
For example, Outset describes AI-moderated research with people. SynthFolk's qualitative workflow produces simulated interviews. These descriptions identify the evidence source; they do not compare the products' effectiveness.
Ask the provider to show an actual input-to-output path and to identify where people participate. Do not infer human participation from names, avatars, a transcript layout or natural conversational phrasing.
Use simulated interviews for a bounded preparation task
One defensible starting task is reviewing an interview guide. You can ask a simulation to respond to the questions, look for ambiguous wording and note assumptions that should be challenged in a human pilot. Sarstedt and colleagues (2024) identify upstream pretests and pilots as promising uses of silicon sampling; they do not claim that synthetic conversations replace all qualitative research.
Set a narrow output contract. Instead of “tell us what customers think”, ask for possible interpretations of the supplied offer, missing information and questions to investigate. Require uncertainty to remain visible. If the simulation introduces an unsupported factual claim about the market, move it to a verification list.
Maintain a distinction between an objection worth exploring and an objection known to exist among customers. This distinction should survive into presentations. A quotation marked “simulated response” in a transcript must not become “customer feedback” on the executive slide.
Write a guide that allows disagreement
Start with the respondent's interpretation before introducing your explanation. In a hypothetical study of a bicycle-repair subscription, “What do you understand to be included?” is more informative about comprehension than “Would you enjoy the convenience of unlimited repairs?” The second question supplies both a benefit and a potentially misleading description.
Ask about practical conditions, alternatives and reasons to decline. Avoid embedding the desired response in the persona brief: “budget-conscious cyclists who want a subscription” selects interest before the study starts. A broader brief can distinguish frequent riders, occasional riders and people who already maintain their own bicycles.
For a human interview, accounts of a recent repair can provide useful context. In a simulation, a purported recent repair is fictional unless grounded in supplied evidence; treat it as a scenario, not a recalled event.
Use the research-brief template to separate known facts, audience assumptions and the decision. The concept-testing questions offer wording for comprehension and barriers.
A worked preparation workflow
First, write the subscription offer with clear exclusions and ask the simulated interview to paraphrase it. If the answer assumes replacement parts are included when they are not, inspect whether the description invites that reading. Revise the description if necessary.
Second, ask for situations in which the offer would be inconvenient. An answer about travel distance suggests a question for human research: how far are relevant riders willing and able to travel for repairs? The synthetic answer does not measure that distance.
Third, repeat the exercise with the same brief while changing question order. Keep all runs. If the prominence of an objection depends on which benefit was mentioned first, report the sensitivity rather than selecting the most persuasive transcript.
Fourth, take the revised guide to relevant people. Record which anticipated issues actually appeared, which did not and what the simulation missed. The bicycle example is illustrative; no customer findings or performance results are being claimed.
Do not confuse stable dialogue with validated insight
Ong (2024) discusses problems of inference and reproducibility when using LLMs, including prompt and procedure reporting. For an applied study, retain the model identification where available, date, brief, exact questions and stimulus. These records make a discrepancy diagnosable; they do not make the output correct.
Repeated agreement among simulated personas may reflect similar model assumptions. It should not be interpreted as evidence that a view is widespread. Likewise, deliberately making personas disagree can create a more interesting transcript without making its disagreements more representative.
Use external reference material to assess usefulness. If the task is finding possible misunderstandings, compare the rehearsal with human pretest results. If the task is predicting a ranking, evaluate against held-out human responses. The validation protocol treats these as different targets.
What the evidence on human focus groups does not say
Caruso and colleagues (2025) found that use of focus groups was associated with less successful reported pack redesign outcomes in their marketer survey. That study has an important boundary: it did not randomly assign research methods, examine every focus-group use or compare human groups with AI simulations.
It therefore cannot justify a claim that AI focus groups outperform human groups. Its useful lesson here is to examine whether the method and measures fit the decision. A conversation about liking a pack and a task measuring recognition of that pack are not interchangeable simply because both can be called consumer research.
Use the simulation to improve the next conversation with people. Prepare a bounded task, review current pricing and start a qualitative exercise. Keep the generated material labelled throughout the research handoff.
Sources and editorial method
The bicycle-subscription workflow is our proposed example, not a reported study. Product descriptions were checked in September 2026. The academic sources support the methodological distinctions and do not validate SynthFolk.
- Sarstedt, M., Adler, S. J., Rau, L., & Schmitt, B. (2024). Using large language models to generate silicon samples in consumer and marketing research: Challenges, opportunities, and guidelines. Psychology & Marketing. DOI: 10.1002/mar.21982.
- Ong, D. C. (2024). GPT-ology, Computational Models, Silicon Sampling: How should we think about LLMs in Cognitive Science? arXiv:2406.09464, version 1.
- Caruso, W., Romaniuk, J., Page, B., Anesbury, Z. W., & Williams, J. (2025). The role of market research in pack redesign performance. International Journal of Market Research, 67(1), 17–32. DOI: 10.1177/14707853241296656.
- Outset: product description.