Ad creative pretesting checks whether an ad gets its message across before the media budget goes behind it. The questions are plain. Does the viewer notice the brand? Can they say what the ad is offering? Do they believe it, and does anything in it put them off? A synthetic audience can give a first answer to each of these in minutes, across several audience definitions, with the asset frozen so the comparison is fair. It cannot tell you click-through rate, cost per acquisition, or how a platform's delivery system will treat the ad. Those numbers come from a live split test. Pretesting earns its place by sending fewer weak ads into that test, and by giving the ones that go a clear hypothesis.
Creative is the part of a campaign teams test least, and it matters most. NCSolutions analysed nearly 500 FMCG campaigns across TV, digital, mobile, print, and radio, and Nielsen published the results in 2017. Creative was still the largest single driver of sales response, ahead of reach, targeting, and timing. See When It Comes to Advertising Effectiveness, What Is Key?
Most teams still spend more time on targeting than on checking whether the ad makes sense.
What a pretest can and cannot tell you
A pretest measures how an ad is read. A live test measures what the ad does.
A pretest can tell you:
- Whether the brand registers, and how early.
- What the viewer thinks is on offer.
- Which claim draws doubt.
- Whether the tone suits the audience or grates on it.
- Whether the call to action matches what the ad has earned.
- Which of several variants is clearest on one defined dimension.
A pretest cannot tell you:
- Click-through rate, conversion rate, or cost per result.
- How the platform will split delivery between variants.
- How fast the ad wears out with repeated exposure.
- Whether a real person scrolling at speed would stop at all.
Keep the two lists apart in every report. A pretest score quoted as a predicted CTR is a made-up number.
Why a live split test is not enough on its own
Ad platforms have good tools for causal tests. Meta's A/B testing divides the audience into random groups so each group sees one version. See About A/B Testing. Google Ads custom experiments split a campaign's traffic and budget between the original and a trial. See About custom experiments.
These tools tell you which version did better under live delivery. They do not tell you why. They also spend money on every arm, including the ones that were confusing from the start. On a small budget, a four-way split test can run for weeks and still fail to separate the variants.
A pretest costs little and cuts the number of arms. It also gives the live test a reason to exist. "Variant B names the price in the first frame, so we expect fewer low-intent clicks" is a hypothesis. "Variant B looks better" is not.
A six-step creative pretest
1. Freeze the asset
Test what will actually run. For a static ad, that means the image, headline, body copy, and call to action as they will appear. For video, it means the final cut, with a note on how it plays muted.
Change one thing between variants when you can. If variant A and variant B differ in hook, offer, and colour, you cannot tie a difference in response to any one of them.
2. Define the audience by state of mind
Start with the audience the campaign will target, then add what you know about how those people think. "Runners aged 25 to 45 in the UK" is where targeting stops. "Recreational runners who have bought shoes online and distrust performance claims" is where a pretest starts.
Add a second audience when the ad has to work for two groups, such as existing customers and cold prospects. They often read the same ad in different ways.
3. Ask what the ad said before asking whether it was good
Show the ad, then ask the audience to:
- Name the brand.
- Say what is on offer, in its own words.
- Say who the ad seems to be for.
- Name the claim it found hardest to believe.
- Say what it would do next, if anything.
Ask for a rating last. Liking, on its own, tells you the least of anything an ad test can measure. People like plenty of ads they cannot attribute to any brand.
4. Test the first seconds on their own
Much ad viewing is a scroll past. Test a cut-down version as well, such as the first frame of a video or the image with the body copy hidden. If the brand and offer do not survive that, the full ad carries a message most viewers never see.
5. Sort reactions by what they mean
| Signal | What it sounds like | Likely fix |
|---|---|---|
| Brand miss | "Some kind of fitness app?" | Show the brand earlier and larger |
| Offer miss | "I am not sure what they are selling." | State the offer in the first line or frame |
| Disbelief | "Nobody gets results in a week." | Add proof or soften the claim |
| Wrong audience | "This is for serious athletes." | Change casting, language, or setting |
| Tone | "It feels like it is shouting at me." | Adjust voice or pacing |
| Weak next step | "Why would I click that now?" | Match the call to action to the stage |
6. Pick two and write the hypothesis
Take the two variants that pass on comprehension and credibility into the live test. Write down what you expect to differ and why. Choose the platform metric in advance and decide how long the test runs before anyone looks at results.
Reading simulated responses honestly
Synthetic audiences help here for the same reasons they help anywhere early. They are fast, they keep the stimulus fixed, and they give reasons rather than a single number.
They also have known weak spots. A model reads the whole ad, while a person scrolling a feed may take in one frame, so a simulated audience notices things a real one skips. Demographic prompting can overstate differences between groups. See When Synthetic Users Fail. And reactions to humour, music, and visual style are less reliable than reactions to claims and offers, because the first set depends on taste and the second on reasoning.
Treat comprehension and credibility findings as the strongest output. Treat appeal scores and segment gaps as hypotheses.
A worked example
Imagine a DTC coffee subscription brand with four static ads for a first-order discount. All four share a layout and image. Only the headline changes. One leads with price, one with freshness, one with convenience, and one with a sustainability claim.
The team defines two audiences in Thesia. The first buys supermarket coffee and has never subscribed to anything coffee-related. The second cancelled a coffee subscription in the past year.
Both audiences explain the price and convenience ads correctly. The freshness ad reads as generic, and the most common response is some version of "every coffee brand says that." The lapsed subscribers read the sustainability ad as a price premium in disguise. The team drops those two, runs price against convenience in a Meta A/B test for two weeks, and uses cost per first order as the primary metric.
This is a protocol example, not a reported Aetherya outcome.
When to skip simulation
Go straight to a live test when:
- You have the budget to test every variant at meaningful spend.
- The variants differ only in visual style, where taste decides.
- The ad is a small refresh of a proven creative with a performance history.
- The claim is regulated, as in health, finance, or alcohol, and needs legal review first.
Final answer
Pretest ad creative to learn whether viewers get the brand, the offer, and the reason to believe, before you pay to learn whether they click. Freeze the asset. Define the audience by state of mind. Ask what the ad said before asking whether it was liked, and test the first seconds on their own. Send two variants with a written hypothesis into a platform split test, and let that test pick the winner.
An ad nobody understood fails the same way at every budget.
Next step: Explore Thesia to pretest ads and social posts with defined audiences before launch.
Sources
- NCSolutions, published by Nielsen (2017). When It Comes to Advertising Effectiveness, What Is Key?
- Meta Business Help Center. About A/B Testing.
- Google Ads Help. About custom experiments.
- When Synthetic Users Fail: A Cross-Domain Benchmark of LLM-Simulated Human Survey Responses. arXiv preprint, July 2026.
