AI concept testing puts a product idea in front of computational audience models before anyone builds it. The team writes the concept the way a buyer would meet it, defines who that buyer is, and asks each model to explain the idea, name what it would replace, and say what would stop it from trying. The output is a sorted list of objections and misreadings. It is not a demand forecast. That is enough to kill weak concepts early and to sharpen the strong ones before a human study or a smoke test. The method fails when teams read simulated enthusiasm as purchase intent, or when the concept is so new that the model has nothing to reason from. Screen with simulation. Confirm with people. Measure demand with something that costs the buyer effort, such as a waitlist, a deposit, or a paid pilot.
Most product ideas die slowly. A team debates them for weeks, builds a prototype, and learns in month three what a buyer could have said in minute one.
What a concept test is for
A concept test answers a narrow question. Given a short description of a product that does not exist yet, does the intended buyer understand it, want it, and believe it?
Those are three separate failures. A buyer can understand an idea and not want it. A buyer can want it and not believe the team can deliver it. Good concept questionnaires ask about comprehension, appeal, and believability as separate items for this reason, instead of folding them into one "would you buy this?" score.
The decision at the end is usually one of four:
- Drop the concept.
- Rewrite it and test again.
- Take it to human research.
- Build the smallest version that can measure real demand.
A concept test that cannot push a team toward one of these asked the wrong question.
Why run it on a synthetic audience first
Human concept tests are slow for boring reasons. Recruitment takes days, screeners leak, and a panel of 200 people for one concept costs real money. So teams test late, test few concepts, and test the version they have already fallen for.
A synthetic audience removes most of that cost at the screening stage. The team can run ten rough concepts against the same audience definitions in one afternoon, keep the procedure fixed, and compare the objections side by side.
There is published evidence that this works for some questions. Brand, Israeli, and Ngwe asked GPT-3.5 for hundreds of survey responses per prompt. The responses produced downward-sloping demand curves and willingness-to-pay estimates of realistic size, in line with human consumer studies. See Using GPT for Market Research. Li, Castelo, Katona, and Sarvary built brand perceptual maps from language model output and reported agreement with human survey data above 75%. See their 2024 Marketing Science paper.
Both papers study familiar categories, with brands people already know and products that already have market prices. That is the easy case. A concept test is the hard case, because the product is new by definition.
Where simulation misleads on new concepts
Three failure patterns come up again and again.
Politeness. Language models lean agreeable. Ask one whether an idea is appealing and it will usually find something to like. A concept that scores well on appeal and badly on everything else is usually a bad concept with a friendly audience.
Missing context. The model knows what people have written about products that exist. It does not know how a nurse on a night shift would fit a new app into a handover that already runs late. The less a concept resembles anything on the market, the less the model has to work with.
Exaggerated segments. Demographic prompting can make groups look more different than real people are. A July 2026 preprint found this pattern across General Social Survey and World Values Survey tasks. See When Synthetic Users Fail. If a simulated concept test says under-30s love the idea and over-50s hate it, treat that split as a hypothesis to check, not a finding.
These are reasons to design the test with care. They are not reasons to skip it.
A concept test protocol that holds up
1. Write the concept the way a buyer would meet it
Do not test a strategy memo. Test what a buyer would actually read: a headline, two or three sentences on what it does and who it is for, a price or price range if one exists, and what the buyer would have to do to start.
Keep every concept to the same length and format. If one concept has a price and another does not, price is now the variable.
2. Define the audience by the problem
"Women 25 to 40" says little about whether someone needs a meal-planning service. "Parents who cook most weeknight dinners, tried a meal-kit subscription, and cancelled it" says a lot.
Give each audience the traits that bear on the decision. That means the current workaround, what they spend on it now, why they switched or gave up before, and who else has a say. Leave out decorative details. A persona's favourite coffee shop does not make the test more valid.
3. Ask interpretation questions before preference questions
Order matters. Ask the audience to:
- Explain the concept in its own words.
- Name what it would replace or compete with.
- Name the first thing that would stop it from trying.
- Say what evidence it would need to believe the main claim.
- Only then, rate appeal and likelihood to try.
Most of the value sits in the first four answers. If half the audience explains a concept wrongly, the concept has a comprehension problem, and its appeal score describes a product that does not exist.
4. Record objections by type
Sort what comes back by kind of failure, not by good and bad sentiment.
| Failure type | What it sounds like | What to do |
|---|---|---|
| Comprehension | "So it is a kind of CRM?" when it is not | Rewrite the description |
| Relevance | "We already solved this with a spreadsheet." | Narrow the audience or the job |
| Credibility | "Nobody can do that in ten minutes." | Show the mechanism or soften the claim |
| Switching cost | "Moving our data would take a quarter." | Test a lighter first step |
| Price | "That costs more than the tool it replaces." | Test price as its own variable |
| Authority | "My manager would never approve this." | Add the approver as a second audience |
The table is the deliverable. It tells the team what to change. A single appeal score never does.
5. Run it more than once
Run the same concept against the same audience at least twice and check whether the main objections repeat. A top objection that shows up in one run and vanishes in the next is noise. One that shows up every time is worth a human conversation.
6. Set the gate for the next stage
Decide before the run what a concept has to show to move forward. A reasonable gate:
- Most of the audience explains the concept correctly.
- The top objection is one the team can address or test.
- The concept clearly beats a named workaround.
- The main claim has a believable mechanism behind it.
Concepts that pass go to human interviews or a demand test. Concepts that fail get one rewrite or get dropped.
What comes after simulation
Simulated interest is not demand. Nobody in a synthetic audience has a budget, a calendar, or a boss.
The next step should make a real person pay some cost. That might be an interview where the buyer walks through how they handle the problem today, a landing page that asks for an email or a deposit, or a paid pilot with a narrow scope. Each step costs the buyer a bit more and tells the team more than the step before.
Keep the simulation record next to the real results. Over a few cycles, the comparison shows which objections the simulation caught and which ones only real buyers raised. That record is how a team learns how far to trust the next screen.
A worked example
Imagine a small B2B team with four concepts for a feature that summarises customer calls for account managers. The concepts differ on one thing, which is what the summary is for. One targets handovers, one renewal risk, one coaching, and one the CRM record.
The team defines two audiences in Thesia. The first is account managers at mid-sized SaaS companies who each handle 40 or more accounts. The second is their sales managers, who approve new tools. Each concept is three sentences long, the same length, with no price.
The account manager audience explains three of the four concepts correctly. The coaching concept reads to them as surveillance, and the top objection is some version of "my manager will use this to grade my calls." The sales manager audience likes the coaching concept most. That split is the finding. It goes into six human interviews, three with each role, before anyone designs a screen.
This is a protocol example, not a reported Aetherya outcome.
When to skip simulation
Go straight to people when:
- The concept depends on a physical experience, such as taste, fit, or feel.
- The audience is small and specialised, and you can reach ten of them this week.
- The concept changes a product people already use, and you have the usage data.
- The product is regulated, such as a financial or medical one.
Simulation should remove uncertainty it can address. It should not delay a conversation you could have tomorrow.
Final answer
AI concept testing is a fast, cheap screen. Use it to find concepts that buyers misread, disbelieve, or cannot fit into their work, and to rewrite the ones worth keeping. Ask interpretation questions before preference questions. Sort objections by type. Run each test twice. Then take the survivors to real buyers and ask them to spend something, even if it is only twenty minutes.
A concept that only works on a polite audience does not work.
Next step: Explore Thesia to screen concepts against defined audiences before your next build cycle.
Sources
- Brand, J., Israeli, A., and Ngwe, D. (2023). Using GPT for Market Research. Harvard Business School Working Paper 23-062.
- Li, P., Castelo, N., Katona, Z., and Sarvary, M. (2024). Frontiers: Determining the Validity of Large Language Models for Automated Perceptual Analysis. Marketing Science, 43(2), 254-266.
- When Synthetic Users Fail: A Cross-Domain Benchmark of LLM-Simulated Human Survey Responses. arXiv preprint, July 2026.
