Concept Testing Questions: A Questionnaire You Can Adapt
Concept testing questions should reveal what people understand, whether the proposed benefit matters in their circumstances, and what prevents them from considering the idea. Ask these separately. A high liking score cannot tell you whether respondents understood the offer, could access it or would choose it over an existing alternative.
In AI market research, you also need to identify who answers. The questionnaire below can support a human concept test. Running it with synthetic respondents is a rehearsal that generates possible interpretations, not a measurement of customer demand.
Define the concept before writing the questions
A concept is a proposed offer described precisely enough to evaluate. Include its intended use, relevant features, availability and price when price is part of the decision. Do not compare a detailed, polished description with a vague alternative and attribute every difference to the underlying idea.
As an illustrative example, imagine a neighbourhood shop considering a weekly collection service for household essentials. The decision is whether to develop a paid pilot, revise the offer or drop it. A description should explain collection times, order deadlines, what is included and what the fee covers. Describing it only as “a convenient new service” embeds the desired conclusion before anyone responds.
Specify the target population and exclusions. Existing subscribers to a related service may help diagnose operational issues, but they are not the only relevant audience for a growth decision. Trinh, Dawes and Sharp (2024) found much of the examined brands' growth headroom among light and non-brand buyers. That is a reason to question a loyal-customer-only brief, not a formula for the response quotas of every concept test.
A sequence of questions with distinct jobs
The following wording is original and illustrative. Adapt it to the category and check comprehension with relevant people before fielding. It is not a validated measurement scale.
| Purpose | Example question | What to watch for |
|---|---|---|
| Category eligibility | “When, if ever, did you last buy household essentials for your home?” | Include people relevant to the decision; do not screen solely for enthusiasm. |
| Unaided understanding | “In your own words, what would this service let you do?” | Interpretations that differ from the intended offer. |
| Missing information | “What, if anything, would you need to know before considering it?” | Gaps in the description rather than flaws in the underlying idea. |
| Relevant situation | “In what situation, if any, would this be useful to you?” | Concrete occasions rather than general praise. |
| Current alternative | “What would you do in that situation today?” | The actual alternative, including doing nothing. |
| Perceived advantage | “What, if anything, would be better than your current approach?” | A benefit meaningful to the respondent. |
| Perceived disadvantage | “What, if anything, would be worse?” | Trade-offs hidden by a positive-only questionnaire. |
| Practical barrier | “What could prevent you from using it?” | Access, timing, household constraints or another obstacle. |
| Credibility | “Which part of the description, if any, is difficult to believe?” | Claims that need explanation or evidence. |
| Consideration | “Given the stated fee and collection times, how likely would you be to consider trying it?” | A stated intention conditional on the shown offer. |
| Reason | “What is the main reason for that answer?” | Interpretation of the rating, not just its magnitude. |
| Final correction | “What have we misunderstood or left out?” | Issues the researcher's framework did not anticipate. |
Do not automatically ask all twelve. Keep only questions that contribute to the decision. A screening question is useful in a human study because it helps establish eligibility; it does not authenticate a generated persona's invented purchasing history.
Keep wording and response options neutral
Avoid questions such as “How useful is our affordable and convenient service?” They combine several claims and make disagreement awkward. Ask about usefulness, the fee and the schedule separately when each matters. Include a genuine way to express no need, uncertainty or insufficient information.
For a consideration scale, make the endpoints clear and the options balanced. Do not present a neutral answer as a negative one in the report. If respondents cannot judge the idea without information you omitted, distinguish that response from rejection.
Pew Research Center's questionnaire guidance discusses how wording, response options and question order affect answers, and why pretesting matters. A synthetic rehearsal can flag possible ambiguities, but only a human pretest shows how relevant people interpret the actual questionnaire.
Ask for an initial explanation before listing possible benefits. Once you show “saves time, reduces effort, improves planning”, later answers may repeat the supplied language. Record which ideas respondents raised without prompting and which appeared only after a prompt.
Compare alternatives without confusing the design
Showing each respondent one concept lets them evaluate it without seeing its competitors in the study. Showing several supports direct comparison but introduces order and contrast effects. Neither design is universally correct. Choose according to whether the real decision is stand-alone comprehension or a choice among alternatives.
If you show several concepts, vary their order and give them comparable information. Preserve the same question wording. Do not add a price to one option and omit it from another unless that asymmetry is intentionally part of the offer being tested.
Include the current solution where relevant. Choosing the most popular of three new ideas does not establish that any improves on what people already use. The concept-testing tools guide explains how the research design determines the platform requirements.
Worked example: turn an answer into a next step
Suppose a respondent in a hypothetical pilot interprets “weekly collection” as home delivery. That answer points first to a comprehension problem. Rewrite the description and retest understanding before treating their purchase-intent response as an evaluation of the intended service.
Suppose another respondent understands the offer but cannot collect during the available hours. That suggests an access constraint. Adding more persuasive language would not remove it. A third person might find the schedule workable but already buy everything during another regular trip; the service may solve little for them.
These are invented examples of interpretation, not reported findings. The useful output is a table connecting each observed issue to a possible change and the evidence needed to evaluate it. Keep competing explanations visible. An objection about a fee might reflect low relevance, unclear value or inability to pay; the objection alone does not identify which explanation is correct.
What a synthetic rehearsal can and cannot establish
Sarstedt and colleagues (2024) discuss silicon sampling as a promising aid to pretests and pilots. Applied here, a simulation can propose misunderstandings, expose assumptions in the brief and help rehearse follow-up questions. These are candidates for inspection.
Do not convert “eight of ten simulated personas said yes” into an estimate of household demand. A generated response to “what did you buy last week?” is not a purchase record. Agreement across reruns can reveal stable model output while leaving its correspondence with people unknown.
Use the rehearsal to revise the instrument, then test the revised questions with people. Keep a change log: original question, simulated concern, editorial decision and human-pretest result. This prevents a model's plausible objection from silently becoming a fact about the market.
Decide what happens after the answers arrive
Define the next decision before collection. Comprehension failures may trigger a rewrite. Repeated practical barriers may require a changed offer. Evidence that the idea matters in a real buying situation may justify a behavioural pilot. A positive stated intention still needs to be distinguished from an actual purchase.
Avoid treating liking as the only outcome. In an observational study of 227 marketers' recent pack redesigns, Caruso and colleagues (2025) found that research methods and metrics were associated with different reported outcomes. It was not a randomised comparison proving one method caused success. Its practical relevance is to make the measured outcome explicit rather than assume that any favourable research score reduces decision risk.
Build a questionnaire around your decision. Start with the research brief, use the quantitative workflow for a labelled synthetic rehearsal, and review current pricing. Take the revised instrument into human research before making a demand claim.
Sources and editorial method
The questionnaire and shop example are original teaching material. Academic sources support the audience and method distinctions; they did not validate this questionnaire or test SynthFolk.
- Trinh, G. T., Dawes, J., & Sharp, B. (2024). Where is the brand growth potential? An examination of buyer groups. Marketing Letters, 35, 95–106. DOI: 10.1007/s11002-023-09682-7.
- Sarstedt, M., Adler, S. J., Rau, L., & Schmitt, B. (2024). Using large language models to generate silicon samples in consumer and marketing research: Challenges, opportunities, and guidelines. Psychology & Marketing. DOI: 10.1002/mar.21982.
- Caruso, W., Romaniuk, J., Page, B., Anesbury, Z. W., & Williams, J. (2025). The role of market research in pack redesign performance. International Journal of Market Research, 67(1), 17–32. DOI: 10.1177/14707853241296656.
- Pew Research Center. Writing Survey Questions.