A synthetic audience test asks computational audience models to respond to a specified message, product, price, or digital journey. Its best use is early decision support. It can help a team reject weak concepts, find likely objections, compare reactions across specified groups, and decide what deserves human or market validation. It cannot observe real customers, prove future conversion, or replace research into lived experience. A sound workflow treats simulation as a disciplined pre-test, keeps the population and assumptions visible, and follows important findings with human research, behavioral data, or a controlled experiment.
That boundary is not a minor disclaimer. It is the reason the method can be useful without becoming misleading.
What decision does synthetic audience testing help with?
The method is most useful when a team has several plausible options and limited time or traffic.
A growth team might have six landing-page concepts but enough traffic to run one serious A/B test. An agency may need to explain why a client's new proposition feels vague before recruitment can begin. A product team may want a first view of objections to a feature that does not yet exist.
In each case, the immediate decision is narrower than "Will this work?" The practical questions are:
- Which idea appears least credible?
- What does each audience think the offer means?
- Where does hesitation begin?
- Which assumption should we investigate with real people?
- Which two variants deserve production and live traffic?
Synthetic testing can reduce the option set. It does not settle the market outcome.
A precise definition
A synthetic audience is a specified set of computational agents used to model responses under stated conditions. The specification may include demographic traits, decision style, risk tolerance, prior knowledge, needs, constraints, and exposure history. The system then places those audience models in a defined task, such as reading a pricing page or choosing between two claims.
This differs from asking a general chatbot, "What would customers think?" A useful simulation needs a population, a stimulus, a decision context, a repeatable procedure, and an evidence record. Without those elements, the result is informed brainstorming dressed as research.
It also differs from synthetic data used to protect privacy in statistical datasets. The shared word is confusing. Synthetic audience testing concerns modeled responses to a decision. Privacy-preserving synthetic data concerns generated records that resemble the statistical properties of a source dataset.
What the research says
The academic record is mixed, which is exactly why broad accuracy claims are inappropriate.
Aher, Arriaga, and Kalai tested language models against established experiments in economics, psycholinguistics, and social psychology. Recent models reproduced several known findings, but the authors also identified a "hyper-accuracy distortion" in one task. Their larger contribution was methodological: evaluate a simulator against a specific human behavior and experiment, not against a vague idea of humanness. See the ICML 2023 paper.
Other work shows why caution is necessary. Bisbee and colleagues compared synthetic responses with American National Election Study data. Overall averages sometimes looked plausible, while variance, subgroup relationships, and reproducibility were weaker. Small prompt changes and model changes also altered results. Their conclusion was not that models are useless. It was that researchers cannot assume synthetic samples support the same inferences as human survey data. See the 2024 Political Analysis paper.
A July 2026 preprint adds a sharper warning. Across General Social Survey and World Values Survey tasks, demographic prompting exaggerated differences between groups under the tested protocols. The models often treated identity as more predictive of attitudes than it was in the human data. The paper is recent and not a final judgment on every system, but its failure pattern matters for any product that reports segment differences. See When Synthetic Users Fail.
The buyer-friendly conclusion is straightforward. Plausible language is not the same as valid population inference. A simulator should be judged by the decision, population, and behavior it claims to model.
Where synthetic audience testing is useful
Eliminating weak options
Early teams often debate concepts that fail for obvious reasons once a buyer tries to interpret them. A simulation can reveal missing context, conflicting claims, unclear ownership, or an implausible promise before design and media costs accumulate.
Finding objections
Objection discovery is an exploratory task. A team can test whether risk, setup effort, price, trust, switching cost, or organizational politics dominates the response for different audience definitions. These are hypotheses for follow-up, not measured market prevalence.
Preparing human research
Recruiting real participants is expensive enough that the interview guide should not waste time on questions a desk review could have exposed. Simulation can help refine stimuli, identify probes, and find areas where respondents may disagree. Human sessions can then focus on lived context and surprises.
Prioritizing live experiments
A/B testing is strongest when variants are mature and the product has enough traffic to measure a meaningful effect. Synthetic testing can help teams decide which variants merit that scarce traffic. The live experiment still provides the causal evidence.
Where it is a poor fit
Do not use synthetic audiences as the sole evidence for:
- Estimating a production conversion rate.
- Questions that depend on what it is like to live through trauma, disability, or discrimination.
- Choices in medicine, law, lending, employment, and government that bear directly on people's rights or welfare.
- Claiming prevalence in a population without observed data and a valid sampling design.
- Discovering behavior caused by conditions absent from the model.
- Replacing usability sessions where physical interaction, accessibility technology, or environmental context matters.
The failure is not that the model lacks eloquence. The failure is that the decision asks for evidence the method does not observe.
A five-part workflow
1. Write the decision before the prompt
Start with the choice the team must make. "Choose two of six value propositions for customer interviews" is testable. "Understand our customers" is not.
State what a wrong choice costs. That may be media spend, development time, recruitment budget, or launch delay.
2. Define the applicable population
Describe who the result concerns and who it does not concern. Use decision-relevant traits, not decorative demographics. A director buying CRO services may care about client risk, proof standards, delivery speed, and white-label reporting. Eye color adds nothing.
Avoid assuming that demographic identity determines an opinion. If segment differences matter, compare them later with human or behavioral data.
3. Freeze the stimulus and procedure
Store the exact page, image, copy, price, and context each audience receives. Record model version, system version, run count, exclusions, and analysis rules. If the procedure changes after results appear, say so.
4. Separate output from interpretation
Raw responses are generated observations within the simulation. Themes and recommendations are analyst interpretations. Neither is observed customer behavior. Label all three layers.
5. Choose the next evidence source
End with a test plan. Use interviews for meaning and lived context. Use analytics for existing behavior. Use a randomized experiment for causal impact. Use field data to calibrate the simulation over time.
How this works in Aetherya
Consider an agency choosing between three hero statements for a DTC client's landing page.
The team defines two buyer cohorts and records why those cohorts matter. It uploads the three variants to Thesia, keeps the visual treatment constant, and asks each audience to complete the same interpretation task. The review focuses on message comprehension, trust, hesitation, and stated objections. It does not report a forecast conversion rate.
The evidence passport records the audience definition, stimulus, procedure, system version, output status, and limitations. The agency uses the result to remove the least coherent version and carries the remaining two into five human interviews or a live A/B test.
This example is a decision protocol, not a claimed experiment result. Before publication as a case study, Aetherya should attach a real run, preserve the raw aggregate output, and state whether later human or market evidence agreed.
How do the methods compare?
| Decision | Synthetic audience | Human research | Live experiment |
|---|---|---|---|
| Remove obviously weak messages | Strong fit | Optional | Later |
| Find likely objections | Strong exploratory fit | Strong fit | Usually later |
| Understand lived experience | Limited | Required | Sometimes useful |
| Estimate production conversion | Directional at most | Limited | Strongest fit |
| Learn why users abandon an existing flow | Useful for hypotheses | Strong fit | Pair with analytics |
| Validate a regulated journey | Supporting role only | Required | Required where appropriate |
What to ask before you buy
A useful sales conversation begins with evidence.
- How is the audience population specified and grounded?
- Which model and system versions produced the result?
- Can I inspect the stimulus, procedure, exclusions, and raw aggregate output?
- How does the system represent disagreement and uncertainty?
- Which claims have been compared with human or market outcomes?
- What would make the result invalid?
- Will a system update stop me from reproducing this decision?
Any accuracy figure needs a label: which behavior, which people, which dataset, and what dates?
The practical rule
Use synthetic audience testing to improve the questions and options that reach the market. Do not use it to pretend the market has already answered.
Teams that keep this distinction can move faster without confusing speed with certainty. Teams that ignore it risk producing precise reports about people they never observed.
Next step: Explore Thesia to frame a bounded audience simulation, or review Aetherya's calibration approach before choosing a method.
Sources
- Aher, G. V., Arriaga, R. I., and Kalai, A. T. (2023). Using Large Language Models to Simulate Multiple Humans and Replicate Human Subject Studies. ICML.
- Bisbee, J., Clinton, J. D., Dorff, C., Kenkel, B., and Larson, J. M. (2024). Synthetic Replacements for Human Survey Data? The Perils of Large Language Models. Political Analysis, 32(4), 401-416.
- Chen, Z., Zhu, D., and Zheng, L. N. (2026). When Synthetic Users Fail: A Cross-Domain Benchmark of LLM-Simulated Human Survey Responses. Preprint.