Concept Testing Tools: Choose the Method Before the Platform
Concept testing tools help present an idea, collect reactions and compare alternatives. The right tool depends on what you need to establish: understanding, relevance, preference or actual behaviour. A platform that generates a ranked list of concepts does not necessarily provide evidence that customers will buy the winner.
The broader AI market research guide distinguishes analysis, collection and simulation. Here, the practical question is which combination supports your next concept decision and how to evaluate it before investing in a larger study.
Decide what the test must resolve
Write down the uncertainty in a form that can be checked. “Is the idea good?” combines several questions. People may understand an idea but have no need for it; want it but be unable to access it; like its appearance but fail to recognise the brand. Each requires a different observation.
Consider a hypothetical refillable cleaning product. A team may need to discover whether shoppers understand that the bottle is reusable. That is a comprehension task. Whether shoppers can find the refill in a store is a different task. Whether they buy and use it again requires behavioural evidence over time.
Specify the decision following each possible result. If confusion leads to a rewritten description, the initial test can be relatively narrow. If the result will authorise national distribution, a simulated preference score is not an adequate evidential basis. The cost of being wrong determines how much additional evidence the decision needs.
Match the uncertainty to a research design
| Uncertainty | Suitable evidence | Useful platform capability | What remains unresolved |
|---|---|---|---|
| Do people understand the offer? | Their explanation in their own words | Open-ended responses and access to originals | Whether they will buy |
| Which objections arise? | Relevant human accounts and follow-up | Interviews, probes and transcript inspection | How common each objection is in the population |
| How do concepts compare on a stated measure? | Comparable responses under a defined exposure design | Assignment, order control and consistent scales | Behaviour outside that study |
| Can people locate or use it? | Observed task performance | Appropriate stimulus presentation and behavioural recording | Long-term adoption |
| Does the change improve a business outcome? | Actual outcomes under a suitable comparison | Reliable exposure and outcome data | Transfer to other contexts |
A synthetic workflow can help prepare descriptions and challenge assumptions before these tests. Its output belongs to the preparation stage until a relevant comparison establishes more. Sarstedt and colleagues (2024) identify pretests and pilots as promising applications of silicon sampling while documenting variation in agreement with human evidence.
Check the stimulus controls
The platform needs to present the material the decision concerns. Text-only feedback on a pack description cannot establish how a pack performs on a crowded shelf. A static screenshot may be sufficient for a question about wording but insufficient for a question about a multi-step interaction.
Give alternatives comparable amounts of information. If one concept includes a familiar brand, an attractive image and a clear price while another has only a working title, the study compares those presentations too. Record which differences are intentional.
Ask whether you can control presentation order, see what each participant was shown and retain the exact stimulus version. If a vendor revises the wording between runs, apparent differences may reflect the revision. A report that does not preserve the stimulus is difficult to interpret later.
Include the existing offer where it is a meaningful alternative. A tool will usually rank the options you give it even when all are weaker than the current solution. Choosing a winner within an incomplete set is not the same as improving the business.
Measure something connected to the decision
For packaging, liking is only one possible outcome. Caruso and colleagues (2025) surveyed 227 marketers about recent pack redesigns and found associations between the research used and reported redesign performance. Their findings favour attention to elements consumers already link to the brand, while some familiar methods and attitude measures were associated with poorer reported outcomes.
This was an observational study with self-reported outcomes and non-probability recruitment. It does not establish that a particular method causes a redesign to fail. It does show why a research purchase should start with a defensible measurement question, rather than the reassurance of having conducted a test.
A later Caruso et al. study (2026) examined modernity, familiarity, likeability and purchase intention in consumer evaluations of redesigned packs. Those are separate constructs. Its structural model concerns measured relationships in that study; it is not evidence of actual sales effects or of a synthetic tool's ability to reproduce them.
For the refill example, a practical scorecard could therefore separate correct understanding, identification of the brand, perceived relevance and unresolved concerns. Label these as distinct outcomes instead of compressing them immediately into one success score.
Evaluate the platform with a small real task
Prepare a short brief, two comparable concepts and a scoring plan. Before choosing a vendor, inspect whether it can preserve the inputs and produce the evidence your scorecard requires. The concept-testing question template provides adaptable wording.
Examine at least one response that supports a conclusion and one that challenges it. Check how incomplete answers, uncertainty and contradictory responses enter the summary. A system that quietly removes them may make the result easier to present and harder to trust.
Compare the total workflow, not only collection speed. Record preparation time, participant recruitment where applicable, analyst review, export quality and necessary follow-up. For synthetic runs, record model information and the full brief where available. For human research, record eligibility, recruitment and completion rules.
Do not treat every extra capability as valuable. If the task is a narrow comprehension check, a complex scoring system may contribute less than readable original responses. If the task requires controlled concept exposure, a conversational interface without assignment controls may be inadequate even if its summaries are strong.
Worked example: revising the refill concept
Imagine that a human pretest of the refill description reveals uncertainty about whether the first purchase includes a bottle. The immediate action is to clarify the offer and check understanding again. It would be premature to interpret low consideration as evidence that shoppers reject refillable products.
After revision, the team may compare the offer with its current product. A simulated run could suggest additional objections about storage or handling. Those are hypotheses to include in follow-up, not observed consumer complaints. If they do not appear in human research, retain that discrepancy rather than rewriting the history of the test.
If the concept advances, a limited commercial pilot can examine real purchasing and repeat use. The success criteria should include operational feasibility as well as customer response. A concept can perform well in a questionnaire and fail because the refill is unavailable when people need it.
These are illustrative decisions, not results of a SynthFolk experiment. Their purpose is to show how the next evidence requirement changes as a concept moves from description to real use.
Where synthetic concept testing belongs
Use simulation to explore possible misunderstandings, rehearse questions and identify assumptions worth checking. Do not describe its respondents as customers who tried the product. Their purchase intentions are generated text or ratings, not demand estimates from a sampled market.
If you use SynthFolk's media workflow for a creative first pass, preserve the supplied stimulus and label the output. If you use the qualitative workflow, keep objections attached to their simulated context. The validation guide explains how to compare a repeatable synthetic task against human evidence.
Choose the smallest test that resolves the next uncertainty. Write the question, select the evidence source, and then compare platforms. You can review SynthFolk pricing and start a labelled exploratory run before taking the resulting questions into fieldwork.
Sources and editorial method
This guide proposes a selection process; it does not rank vendors from a hands-on trial. The refill example is fictional. The packaging studies concern human research and do not validate synthetic concept scores.
- Caruso, W., Romaniuk, J., Page, B., Anesbury, Z. W., & Williams, J. (2025). The role of market research in pack redesign performance. International Journal of Market Research, 67(1), 17–32. DOI: 10.1177/14707853241296656.
- Caruso, W., Romaniuk, J., Page, B., Anesbury, Z. W., Saeed, R., & Williams, J. (2026). The packaging redesign modernisation dilemma: The relationship with familiarity, likeability, and its effect on purchase intent. Journal of Retailing and Consumer Services, 92, 104800. DOI: 10.1016/j.jretconser.2026.104800.
- Sarstedt, M., Adler, S. J., Rau, L., & Schmitt, B. (2024). Using large language models to generate silicon samples in consumer and marketing research: Challenges, opportunities, and guidelines. Psychology & Marketing. DOI: 10.1002/mar.21982.