AI Market Research: What It Actually Replaces (and What It Doesn't)

· 16 min read

AI market research is the use of machine learning — in practice, large language models — to run or accelerate parts of a study: designing the instrument, sampling respondents, collecting responses, analysing what comes back. It covers three genuinely different things: AI analysis of data from real people, AI-moderated conversations with real people, and fully synthetic respondents that generate the data themselves.

Almost every unproductive argument about this category comes from collapsing those three into one. A research director saying "we already do AI market research" usually means their survey platform writes the summary. A vendor saying it may mean nobody was interviewed at all. This guide separates them, so you can pick by question rather than by the tooling you own.

Key takeaways

What "AI market research" actually means

Four distinct levels share the label, and most tool comparisons silently compare across them. Davenport and colleagues draw the distinction formally in the Journal of the Academy of Marketing Science — AI's marketing impact varies with the level of intelligence and the type of task — and conclude it works better augmenting managers than replacing them.

Level 0 — AI features inside existing tools. Your survey platform clusters open-text answers, writes the summary, suggests a chart. Real respondents, real data, AI on the last mile: faster analysis, unchanged cost and fieldwork time.

Level 1 — AI-assisted analysis at scale. Coding thousands of open responses, theme extraction across tickets, reviews or transcripts, cross-source synthesis. Still real human data; the AI does work an analyst would do more slowly. This is where most market research automation delivers today, and it is the best-evidenced end of the spectrum: in Marketing Science, Timoshenko and Hauser found machine-selected user-generated content at least as valuable a source of customer needs as professional experiential interviews, and more productive per unit of analyst effort.

Level 2 — AI-moderated research with real people. A model runs the interview, probes, adapts the guide. You still recruit, still incentivise, still wait for fieldwork; what you save is moderator time and scheduling, and you gain a bigger sample than one moderator could handle. Whether people answer an automated interlocutor as they would a person is contingent, not settled: Gelbrich and colleagues' Journal of Marketing meta-analysis of 943 effect sizes from 327 studies finds customers sceptical of automated agents yet buying from them much as from a human, with contingencies that do not generalise across robots, chatbots and algorithms.

Level 3 — synthetic respondents. No humans participate. A model generates respondents conditioned on demographic and attitudinal profiles, and they answer your questions. Recruitment cost goes to zero, fieldwork to minutes, and the link to any real person becomes indirect, running through whatever data the personas were grounded in. Reviewing studies that compared silicon and human samples, Sarstedt and colleagues report agreement varying considerably across domains, and place the method's promise in upstream pretesting and pilot work rather than in main studies.

Synthetic market research and synthetic data market research both point at Level 3, though the second also covers statistically similar copies of an existing dataset. One simulates respondents; the other anonymises data you already hold.

The research workflow, stage by stage

The useful question is not whether AI does market research, but what it does to each stage.

Stage What AI changes What does not change
Framing the question Nothing — the least automatable step there is A bad question gives useless data at any speed
Designing the instrument Fast drafting; piloting against synthetic respondents finds broken items in an afternoon You still have to recognise a leading question
Sampling and recruitment Level 3 removes it; Levels 0–2 do not touch it Incidence rate governs cost wherever humans are involved
Fielding Minutes instead of weeks at Level 3 Real fieldwork takes real time
Analysis Substantially automated at every level Interpretation stays yours
Reporting Largely automated Nobody is persuaded by a report nobody trusts

The stages AI compresses hardest — sampling and fielding — are exactly the ones that make research slow and expensive. The stage that decides whether the study was worth running is untouched — and it matters more than most teams assume. Caruso and colleagues surveyed 227 marketers in the International Journal of Market Research about their most recent pack redesign: doing research was no guarantee of a better outcome, and the most common methods and metrics — focus groups, brand attitudes — were associated with less successful redesigns, while research identifying which design elements consumers already link to the brand was associated with more successful ones.

Choosing an approach by research question

Start from the question, not the tool.

Your question Best fit Why
Which of eight concepts go into real testing? Synthetic (Level 3) You need the bottom half eliminated, not a precise ranking
Which headline reads clearer to a sceptic? Synthetic (Level 3) Relative judgements resist model bias better than absolute ones
What share of the UK market would pay £39 a month? Human panel Absolute figure, commercial consequences; simulation has no calibration to willingness to pay
Why did churn spike last quarter? Level 1 on your own data, then interviews The evidence already exists in tickets and behaviour
What objections will we hit in a new market? Synthetic first, then human Cheap objection inventory, verified by people who live there
How do users complete this task? Usability testing Observed behaviour has no synthetic equivalent
40 in-depth interviews in three weeks AI-moderated (Level 2) Real participants, moderator capacity no longer the constraint
Substantiating an advertising claim Human panel, documented Regulated evidence needs real respondents and an audit trail
Which of twelve markets deserve fieldwork? Synthetic (Level 3) Prohibitive with panels; you only need divergence signals

The pattern: synthetic simulation is strong where you need breadth, speed and comparison, and weak where you need calibration, observation or defensibility.

Done properly, the stimulus comparison in row two looks like Caruso and colleagues' 2026 packaging study — 484 US and 491 UK consumers rating 48 redesigned packs — which a synthetic run shortlists candidates for rather than replaces.

Where synthetic respondents fit

At the Level 3 end, a synthetic respondent is a language-model persona defined by a profile and prompted to answer in character. What separates a useful one from a worthless one is grounding: whether the profile was invented by the model or sampled from real population data.

Ungrounded personas converge: ask ten the same question and you get one answer restated ten times, all from the same point in the model's training distribution. The convergence is measurable — Lin's review of synthetic respondents reports AI-mediated text markedly more homogeneous than human writing, 45% similarity in summarisation against 27%, with sanitisation truncating precisely the tails where prejudice, ambivalence and extreme views sit. Valenzuela and colleagues call the mechanism parametric reductionism: represent a person with a handful of parameters and you lose what the parameters do not carry.

Grounded personas — seeded from large-scale probability surveys such as the European Social Survey and Eurobarometer — diverge, because they start from genuinely different attitudinal positions. Sun and colleagues showed how far grounding alone goes: conditioning a model on group-level demographic distributions produced response distributions remarkably similar to real US opinion polls, though replicability varied by group and topic. Disagreement inside the panel is the signal; unanimity is a warning.

For the mechanics in depth, see our guides to synthetic users and synthetic personas.

A practical workflow, end to end

1. Define the audience. Write the brief you would hand a recruitment agency: market, age band, income, and the two or three attitudinal traits that matter. "UK grocery shoppers, 25–55, mixed income, ordering online at least weekly, split on price sensitivity" works. "Consumers" does not.

2. Generate the panel. The system samples personas from the grounded distributions matching that brief. Qualitative work uses three to eight personas, each with enough biographical texture to hold a position through a long conversation; quantitative work uses 50 to 250, because you want distributions rather than narratives.

3. Run the instrument. Qualitative studies run as in-depth interviews with probing. Quantitative studies run as surveys of up to five questions — rating scales, multiple choice, yes/no — each respondent answering independently rather than inside one simulated group. A simulated focus group collapses toward consensus for the same reason a real one does, minus the dissent that resists it.

4. Read the output. Themes and quotes for qualitative, distributions and demographic breakdowns for quantitative, comparative scoring for A/B stimuli. Read every number as a comparison, never as a forecast.

5. Validate. Check one finding against something you already know — prior research, sales data, a known segment difference — then take the surviving decisions to real people. Teams skip this; it is not optional.

Try step three yourself: SynthFolk's quantitative module fields a survey to 250 synthetic respondents in one run; the media module does the same for creative.

What it costs

The comparison is not "cheaper" versus "dearer": the two cost models have different shapes.

Panel research has a high marginal cost per respondent, driven by sample size, incidence rate (how many people you screen to find one who qualifies) and interview length. A mainstream consumer audience in one European market is typically a four-figure line item for a few hundred completes; a low-incidence B2B audience, several times that.

Simulation has a near-flat marginal cost. SynthFolk prices in credits, published openly rather than behind a demo request:

Study Credits
Qualitative study, 3 personas from 6
Qualitative study, 8 personas from 16
Quantitative survey, 250 respondents 50

Plans are €9.99 for 75 credits a month, €24.99 for 250, and €79.99 for 750, so a 250-respondent survey takes a fifth of the mid-tier allowance. The full breakdown, including media-evaluation runs, is on the pricing page.

A flat marginal cost changes which questions are worth asking at all: when a comparison costs less than a coffee, you run it on questions you would previously have settled by argument in a meeting — which is why validation matters more, not less.

AI consumer insights: what you can and cannot infer

"AI consumer insights" is the commercial framing of the same spectrum, and it invites one specific overreach.

You can infer direction. Which message lands better, which concept survives a sceptic, which objection recurs everywhere. Rank orders and relative gaps hold up well.

You cannot infer magnitude. A synthetic panel reporting 63% purchase intent has produced a number calibrated to nothing, and there is no correction factor to apply: agreement between silicon and human results varies by domain rather than sitting at a stable offset, so a figure that lands close in one category tells you nothing about the next. At best it is a comparison point against another stimulus measured identically.

You cannot infer much about thin populations. Grounding data is sparsest where teams want most precision: small minorities, rare clinical groups, niche professional roles. Simulated subgroups collapse toward the majority-culture default — confident, fluent and wrong. Fernandez, Berner and Shevlin show how convincing that failure looks: from brief diagnostic descriptions alone, a model produced 2,106 personas whose answers to validated psychiatric screening instruments were clinically differentiated and rose with assigned severity. Plausibility is not validity.

How accurate are synthetic respondents summarises the published validation work, including where it found the method failing.

Personal data, GDPR and the EU AI Act

Factual point, not legal advice — and one of the clearer differences between the approaches.

A conventional study processes personal data at several points: recruitment screening, contact details, incentive payment, recorded interviews, responses linked to an identifiable individual. Each attracts obligations: lawful basis, data minimisation, retention limits, participant rights, and extra care around special-category data such as health.

A synthetic study collects nothing from anyone: no data subject, no consent to manage, no retention schedule, no cross-border transfer when you run the same study across twelve markets. The grounding behind the personas is published aggregate statistics, not individual records.

For EU and UK teams that removes a whole class of process, but not every consideration. Your own inputs still count: uploading a customer verbatim or an unreleased creative asset is a processing decision in its own right. And transparency expectations apply to how you describe the method, so say plainly in the report that the respondents were synthetic; Hardcastle, Vorster and Brown's Journal of Advertising study of AI-driven personalised journeys finds people judging AI mediation on autonomy and disclosure, not efficiency. Our GDPR page covers how SynthFolk handles the data you do provide.

What AI market research does not replace

Stated plainly, because the category earned its credibility problem:

Synthetic research is a complement, not a replacement. It changes the economics of questions you were never going to fund, and sharpens the research you do fund by clearing the wrong turns first. A team that closes its research function and buys a simulation licence has bought an efficient way to confirm its assumptions.

How to choose a tool

Six criteria, in weighting order:

  1. Published pricing. Most of this category is demo-gated: if you cannot see what a study costs before a sales call, you cannot compare.
  2. Grounding data, named. Ask which datasets the personas come from. "Proprietary" usually means invented by the model.
  3. Model transparency. Know which model generated your respondents — your results change when it does. Ong's review of LLM-based inference puts reproducibility, invisible training data and prompt-level "hyperparameters" at the top of the unresolved list.
  4. Independent respondents. Sampled and queried separately, or one context window role-playing a focus group? The second collapses variance.
  5. Data handling and residency. Where your inputs go, how long they are kept, whether they train anything.
  6. Export and languages. Findings you cannot export, or cannot run in the markets you sell in, are findings you will not use.

Frequently asked questions

Is AI market research reliable?

It depends which level you mean. Level 0–1 analysis of real human data is as reliable as the data underneath it. Level 3 synthetic respondents are reliable for directional and comparative questions, unreliable for absolute figures, thin populations and novel categories. Read the validation evidence first.

What does AI market research cost?

At the analysis end, usually nothing extra — a feature of software you already pay for. At the synthetic end, SynthFolk prices in credits: from 6 for a three-persona qualitative study, 50 for a 250-respondent survey, against plans of €9.99 for 75 credits, €24.99 for 250 and €79.99 for 750 a month. Panel research is quoted per project.

Can synthetic respondents replace a panel?

For screening, prioritisation and instrument piloting, they replace the panel study you would otherwise have run first — the upstream role Sarstedt and colleagues describe. For the decision itself — anything with a number attached — they do not. Synthetic to narrow, human to confirm.

Is synthetic market research GDPR-compliant?

Synthetic respondents are not people, so no personal data is collected from participants and the obligations attached to processing it do not arise. Your own inputs are a separate question; describing the method in the report is a transparency matter, not a data protection one. Factual description, not legal advice.

How many synthetic respondents do I need?

Fewer than intuition suggests for qualitative work: three to eight well-grounded personas surface more than twenty badly grounded ones, because respondents from the same collapsed distribution add confidence rather than information. For quantitative work, 250 makes demographic breakdowns readable; 50 is enough for a two-option comparison.


The fastest way to judge any of this is to rerun a question you answered with real research last year and compare. A three-persona study starts at 6 credits, a 250-respondent survey costs 50, and new accounts start with enough credits to run one. Open the dashboard and put the method through its own test.

References

Caruso, W., Romaniuk, J., Page, B., Anesbury, Z. W., & Williams, J. (2025). The role of market research in pack redesign performance. International Journal of Market Research, 67(1), 17–32. https://doi.org/10.1177/14707853241296656

Caruso, W., Romaniuk, J., Page, B., Anesbury, Z. W., Saeed, R., & Williams, J. (2026). The packaging redesign modernisation dilemma: The relationship with familiarity, likeability, and its effect on purchase intent. Journal of Retailing and Consumer Services, 92, 104800. https://doi.org/10.1016/j.jretconser.2026.104800

Davenport, T., Guha, A., Grewal, D., & Bressgott, T. (2020). How artificial intelligence will change the future of marketing. Journal of the Academy of Marketing Science, 48, 24–42. https://doi.org/10.1007/s11747-019-00696-0

European Social Survey ERIC. (n.d.). European Social Survey. https://www.europeansocialsurvey.org/

European Union. (n.d.). Eurobarometer. https://europa.eu/eurobarometer/

Fernandez, K., Berner, L. A., & Shevlin, B. R. K. (n.d.). The threat of synthetic respondents extends to clinical mental health screening. Submitted manuscript, University of California, Los Angeles, and Icahn School of Medicine at Mount Sinai.

Gelbrich, K., Roschk, H., Miederer, S., & Kerath, A. (2026). Automated versus human agents: A meta-analysis of customer responses to robots, chatbots, and algorithms and their contingencies. Journal of Marketing, 90(2), 1–26. https://doi.org/10.1177/00222429251344139

Hardcastle, K., Vorster, L., & Brown, D. M. (2025). Understanding customer responses to AI-driven personalized journeys: Impacts on the customer experience. Journal of Advertising, 54(2), 176–195. https://doi.org/10.1080/00913367.2025.2460985

Lin, Z. (n.d.). Synthetic respondents and the illusion of human data. Preprint, Department of Psychology, Yonsei University.

Ong, D. C. (2024). GPT-ology, computational models, silicon sampling: How should we think about LLMs in cognitive science? arXiv:2406.09464. https://arxiv.org/abs/2406.09464

Sarstedt, M., Adler, S. J., Rau, L., & Schmitt, B. (2024). Using large language models to generate silicon samples in consumer and marketing research: Challenges, opportunities, and guidelines. Psychology & Marketing. https://doi.org/10.1002/mar.21982

Sun, S., Lee, E., Nan, D., Zhao, X., Lee, W., Jansen, B. J., & Kim, J. H. (2024). Random silicon sampling: Simulating human sub-population opinion using a large language model based on group-level demographic information. arXiv:2402.18144. https://arxiv.org/abs/2402.18144

Tanusondjaja, A., Romaniuk, J., Nenycz-Thiel, M., Sakashita, M., & Viswanathan, V. (2023). Examining Pareto Law across department store shoppers. International Journal of Market Research (advance online publication). https://doi.org/10.1177/14707853221145851

Timoshenko, A., & Hauser, J. R. (2019). Identifying customer needs from user-generated content. Marketing Science (Articles in Advance). https://doi.org/10.1287/mksc.2018.1123

Valenzuela, A., Puntoni, S., Hoffman, D., Castelo, N., De Freitas, J., Dietvorst, B., Hildebrand, C., Huh, Y. E., Meyer, R., Sweeney, M. E., Talaifar, S., Tomaino, G., & Wertenbroch, K. (2024). How artificial intelligence constrains the human experience. Journal of the Association for Consumer Research, 9(3). https://doi.org/10.1086/730709