AI Ad Creative Testing: From Simulated Feedback to a Real Test
AI ad creative testing can help inspect an advertisement's wording, possible interpretations and unanswered questions before fieldwork. Generated feedback does not measure attention, brand memory or sales among people. Use it to form testable hypotheses, then choose a human or behavioural test that matches the advertising decision.
Within AI market research, creative evaluation is one application of the distinction between analysing material and measuring audience response. A system may accurately describe what an image contains while still being unable to predict what a buyer notices or remembers.
Decide which outcome the creative needs to improve
“Which ad is best?” combines several possible outcomes. One advertisement may communicate the offer clearly; another may be easier to connect with the brand; a third may produce more clicks from people who never buy. Define the business question and the intermediate measures before choosing a score.
For example, a local retailer may need its advertisement to make a seasonal offer easy to understand and identify the retailer. A creative test should not reduce that task to aesthetic preference. If the eventual objective is incremental sales, the evaluation must eventually include outcomes and a credible comparison capable of addressing incrementality.
Separate diagnostic questions from success measures. Asking what seems confusing can help revise an ad. It does not directly estimate how much the revision will change sales. A useful diagnostic should lead to a testable change, not an unsupported forecast.
Keep brand recognition separate from liking
Romaniuk and Sharp (2004) distinguish brand salience from an attitude-only perspective, defining it in terms of a brand being noticed or coming to mind in buying situations. That directs attention to buyers' memory structures and the situations that activate them.
A text model's ability to name a logo or describe a brand is not a measurement of those structures in customers. Likewise, a simulated persona saying “I remember this advertisement” has not experienced the intended media exposure. Treat that answer as generated content, not measured recall.
For a practical creative review, ask whether the execution preserves the elements people actually associate with the brand. Establishing that association requires relevant human evidence. The model can point to candidate elements or describe changes; it cannot certify their familiarity in the target population merely by recognising them itself.
Use a diagnostic matrix before asking for a winner
| Question | What a simulation may help inspect | Evidence needed for the stronger claim |
|---|---|---|
| Is the offer understandable? | Possible readings, ambiguity and missing conditions | Comprehension among relevant people |
| Is the brand identifiable? | Presence and clarity of brand cues in the material | Human recognition or attribution under suitable exposure |
| Does it address a buying situation? | Plausible links between the message and the stated scenario | Evidence that the situation matters to buyers |
| Is the creative noticed? | Visual or textual elements worth examining | Appropriate attention or exposure measurement |
| Does it change behaviour? | Hypotheses about possible mechanisms | Actual outcomes and a suitable comparative design |
This matrix is our proposed planning aid. It does not imply that every item can be measured accurately by a synthetic panel. Choose only diagnostics that contribute to the decision and preserve the distinction between a suggestion and an observed result.
Prepare comparable creative variants
If the intended test concerns a headline, hold other consequential features reasonably consistent. If several elements change together, describe the comparison as one execution against another; do not attribute the outcome solely to the headline.
Record image, copy, format and the offer conditions. Compare versions in a context relevant to the intended placement. A full-size image inspected at leisure differs from a mobile advertisement passed quickly in a feed. A model examining the uploaded image does not recreate that human exposure by default.
Include the current execution where relevant. Ranking three new designs only identifies a preferred option within that set. It does not establish improvement over the advertisement already running.
For question wording, use the concept-test questionnaire. For platform selection, see concept-testing tools.
A worked example: a retailer's seasonal offer
Imagine a retailer comparing two fictional advertisements for the same offer. Version A gives the discount prominent placement but makes eligibility hard to find. Version B explains eligibility but reduces the prominence of a familiar brand element. These are hypothetical design choices, not real test results.
A synthetic review may interpret A as applying to all products when it applies to selected items. The team should inspect the wording directly and ask relevant people what they understand. If the ambiguity is confirmed, revise it. The simulated misunderstanding helped identify a question; the human check established whether it occurred among the tested people.
The model may prefer B's cleaner appearance. That does not resolve whether buyers recognise the retailer as quickly. Preserve the brand question for an appropriate human comparison rather than treating the aesthetic judgement as a substitute.
After fixing comprehension, the team can test viable variants in the intended channel. Define exposure, outcome and the comparison before examining performance. A difference in platform-reported conversions is not automatically incremental sales; audience allocation, delivery and measurement need to support the claim being made.
What packaging research contributes to creative decisions
In a survey of 227 marketers, Caruso and colleagues (2025) found associations between pack-redesign research choices and reported outcomes. Their findings draw attention to retaining elements consumers link with the brand. The study's observational design and self-reported measures limit causal conclusions.
Caruso and colleagues (2026) separately examined modernity, familiarity, likeability and purchase intention in human evaluations of redesigned packs. The lesson for an advertising brief is to avoid treating these constructs as synonyms. The paper concerns packaging and stated intentions, so applying its distinctions to advertising is a planning inference, not direct evidence that a particular ad change will increase sales.
Neither study demonstrates that generated creative ratings reproduce human research. Cite them for the measurement questions they support, not as endorsements of an AI evaluation score.
Turn the review into a record of testable changes
For each proposed edit, write the issue, its evidence status, the change and the next check. “The model suggested the condition may be unclear” is a hypothesis. “Several people in the pretest interpreted the condition differently” is a finding about that pretest. “The revision increased sales” requires outcome evidence beyond either of those observations.
Keep rejected suggestions too, especially if they would remove established brand elements without supporting evidence. A long list of possible edits is not a reason to change everything. Prioritise the changes that address a consequential uncertainty and can be evaluated.
If the same simulation is used repeatedly, assess it against human references for the relevant task. Record misses and false alarms as well as useful suggestions. The validation protocol gives a way to do this without selecting only successful examples.
Use SynthFolk as an exploratory step
SynthFolk's media workflow provides simulated evaluations of supplied material. Keep the original stimulus and label the generated responses when sharing the report. Do not present a simulated score as a prediction of click-through rate, recall or revenue without external validation appropriate to that prediction.
Sarstedt and colleagues (2024) describe promising upstream roles for silicon sampling. In this setting, the bounded use is to prepare a more focused human or behavioural test, with a clear handoff from hypotheses to evidence.
Start with one creative question. Review current pricing, run an exploratory evaluation, and record what real-world observation would confirm or reject each proposed change.
Sources and editorial method
The diagnostic matrix and retailer example are original applications. The cited packaging studies do not test advertisements or SynthFolk. Their transfer to creative planning is explicitly limited to choosing and distinguishing measurement questions.
- Romaniuk, J., & Sharp, B. (2004). Conceptualizing and measuring brand salience. Marketing Theory, 4(4), 327–342. DOI: 10.1177/1470593104047643.
- Caruso, W., Romaniuk, J., Page, B., Anesbury, Z. W., & Williams, J. (2025). The role of market research in pack redesign performance. International Journal of Market Research, 67(1), 17–32. DOI: 10.1177/14707853241296656.
- Caruso, W., et al. (2026). The packaging redesign modernisation dilemma: The relationship with familiarity, likeability, and its effect on purchase intent. Journal of Retailing and Consumer Services, 92, 104800. DOI: 10.1016/j.jretconser.2026.104800.
- Sarstedt, M., Adler, S. J., Rau, L., & Schmitt, B. (2024). Using large language models to generate silicon samples in consumer and marketing research: Challenges, opportunities, and guidelines. Psychology & Marketing. DOI: 10.1002/mar.21982.