Synthetic users and real user research answer different questions. A synthetic user is a computational audience model that responds under specified conditions. Real user research gets its evidence directly from people. Sometimes that means an interview. In other studies, the researcher watches somebody try the product, sends a survey, or sits beside them at work. Synthetic users work best while ideas are still cheap to change. Teams can use them to reject confusing concepts, uncover possible objections, and find holes in an interview guide before committing the research budget. Real people are necessary when the decision depends on lived experience, current behaviour, physical interaction, accessibility, organizational context, or accountable testimony. Use simulation to narrow the problem. Use human research to learn what the model could not know. Use market experiments when the claim concerns causal impact on real behaviour.
The argument about replacement starts with the wrong question. Method choice should begin with the decision.
What is a synthetic user?
A synthetic user is a model configured to respond as part of a defined audience and task. The setup may include role, goals, prior knowledge, constraints, decision style, risk tolerance, and exposure to a stimulus.
The useful unit is not a colourful persona card. It is the full procedure: population, stimulus, task, system version, repeated runs, analysis rule, and evidence boundary.
Ask a general chatbot what "users" think and you get plausible brainstorming. Run a recorded procedure against a specified audience and you have a simulation that can be inspected and compared. Neither form observes a human participant.
What counts as real user research?
Real user research covers several methods, and they do not all answer the same question.
Interviews collect testimony about experience, memory, meaning, and current practice. Usability sessions observe someone attempting a task. Field research examines work in context. Surveys estimate responses in a sampled population when the design supports that inference. Controlled experiments estimate the effect of a change on behaviour.
Calling all of this "talking to users" hides the reason each method exists.
The decision each method can support
Suppose a product team has twelve onboarding concepts.
The team can use synthetic users to identify obvious comprehension problems, missing information, and likely objections. It may reduce twelve concepts to four before a designer builds high-fidelity prototypes.
Real usability sessions can then reveal whether people notice the controls, understand the sequence, use assistive technology successfully, and recover from mistakes. Those sessions may expose conditions nobody described in the simulation.
A production experiment can compare the final variants on activation and retention. That is a third source of evidence.
The methods form a sequence because the uncertainty changes as the product moves closer to market.
What published research tells us
The evidence does not support a universal verdict.
Aher, Arriaga, and Kalai proposed "Turing Experiments" that compare model simulations with specific human-subject experiments. Recent models reproduced several established findings and showed a systematic distortion in another task. Their work supports behaviour-specific evaluation. It does not show that a model can stand in for any person in any study. See the ICML 2023 paper.
Bisbee and colleagues compared millions of generated responses with American National Election Study data. Synthetic averages sometimes appeared close to the survey, but the generated sample had less variation, different regression relationships, prompt sensitivity, and changes across collection dates. Forty-eight percent of the synthetic regression coefficients differed significantly from the human estimates, and some effects changed sign. See the Political Analysis article.
Recent UI research is trying to measure where synthetic evaluation does and does not transfer. The 2026 Sycamore preprint compared grounded and ungrounded synthetic personas in a specialist genomics-visualization task. Grounding moved feedback toward concerns documented among real users, yet both synthetic conditions missed a preference found in the expert study. See Sycamore.
That pattern matters. Grounding can improve relevance without guaranteeing that the simulation will discover the same issue as a person.
Where synthetic users fit well
Screening concepts
Rough ideas often contain basic flaws. The proposition is vague. The proof does not support the claim. The call to action asks for too much. A simulation can expose these problems before the team spends recruitment or engineering budget.
Building an objection map
Synthetic users can produce hypotheses about risk, price, trust, switching work, procurement, and relevance. The team can organize those hypotheses and test the important ones with customers.
Preparing human research
An interview guide improves when the researcher has already stress-tested assumptions and stimuli. Simulation can find leading questions, weak comparisons, and missing probes.
Comparing a fixed stimulus across audiences
A recorded task can apply the same page or concept to several audience definitions. The output helps the team inspect how assumptions about context change the response.
Working before the product exists
Early strategy often has no interface, analytics, or traffic. Synthetic users can help examine a proposition or decision brief while changes remain cheap.
Where real people are required
Lived experience
A model has no employment history, disability, family obligation, migration experience, or memory of a failed implementation. It can generate a plausible account. It cannot supply lived evidence.
Discovery
Researchers often learn that the original question was wrong. A person can describe an improvised workflow, political constraint, or unmet need that the team did not encode.
Physical and environmental context
Hardware, movement, lighting, noise, assistive technology, and interruptions shape many experiences. Text generation cannot observe those conditions.
Rare or poorly represented groups
Models may reduce an unfamiliar population to common language patterns or stereotypes. Direct recruitment and relevant expertise matter when the audience is small, local, specialist, or marginalized.
Accountable consultation
Public policy, health, employment, credit, and other consequential decisions may require documented participation by affected people. Synthetic responses cannot provide consent or representation.
Comparison by evidence need
| Question | Synthetic users | Real user research | Live experiment |
|---|---|---|---|
| Which rough concept is hardest to understand? | Strong pre-test | Useful | Usually later |
| What objections should we investigate? | Strong exploratory fit | Strong fit | Not needed first |
| How does this workflow operate in practice? | Limited | Required | Sometimes useful |
| Can users complete the task with assistive technology? | Poor fit alone | Required | Useful after accessibility review |
| Which variant changes conversion? | Directional at most | Limited | Strongest fit |
| How common is an attitude in the market? | Not sufficient alone | Survey with valid sampling | Depends on outcome |
| Why did current users abandon? | Useful for hypotheses | Strong fit | Pair with analytics |
A practical sequence
Stage 1: define the decision
Write the choice the team faces and the consequence of being wrong. "Select two price explanations for customer interviews" is clear. "Understand pricing" is not.
Stage 2: run synthetic exploration
Freeze the audience, stimulus, and task. Record generated observations, disagreement, and unstable answers. Remove concepts that fail basic interpretation.
Stage 3: recruit real people
Use the simulation to write better screening questions and probes. Ask people about actual behaviour and context. Pay attention when their evidence contradicts the model.
Stage 4: test in the market
Where the decision concerns real behaviour, run a controlled experiment or rollout with suitable instrumentation. Compare the observed direction with the earlier simulation.
Stage 5: update the model
Store where the simulation agreed, missed, or exaggerated. A synthetic research program improves through comparison with reality, not through confidence in fluent output.
A worked Aetherya protocol
Imagine a software company deciding how to explain a new approval workflow.
The team creates three explanations. It defines two audiences in Thesia: an operations manager who requests approval and a finance controller who reviews it. Each audience receives the same interpretation task. The team records what each version seems to do, who owns the next step, what risk remains, and which terms cause confusion.
One version fails because both audiences interpret "automated approval" as removing human control. The team rewrites it as "automated routing with approval retained by finance." That revised concept goes into six human interviews.
During the interviews, controllers explain that audit export matters more than approval speed. The synthetic pre-test did not find that issue. The product team changes the prototype and later measures adoption in a controlled rollout.
This example shows the division of labour. Simulation removed a wording failure. People uncovered a workflow requirement. Market data measured use.
It is a protocol example, not a reported Aetherya result.
Questions to ask before trusting synthetic users
- Which population does the model claim to represent?
- What behaviour or decision is being simulated?
- Which information grounds the audience?
- Can I inspect the exact stimulus and procedure?
- How does the system show disagreement?
- Which human or market outcome has been used for comparison?
- What changed between system versions?
- What finding would trigger human research instead?
If the answer to every limitation is "the model is highly accurate," the method is not ready for a serious decision.
Final answer
Synthetic users are useful before research and between research cycles. They help teams screen, prepare, and prioritize. Real users provide experience, context, surprise, and accountable testimony. Live experiments measure what a change does in the market.
Choose the evidence the decision needs. Do not ask one method to impersonate all three.
Next step: Explore Audience Chat in Thesia for a bounded synthetic pre-test, then carry the strongest questions into human research.
Sources
- Aher, G. V., Arriaga, R. I., and Kalai, A. T. (2023). Using Large Language Models to Simulate Multiple Humans and Replicate Human Subject Studies. ICML.
- Bisbee, J., Clinton, J. D., Dorff, C., Kenkel, B., and Larson, J. M. (2024). Synthetic Replacements for Human Survey Data? The Perils of Large Language Models. Political Analysis, 32(4), 401-416.
- Srinivasan, A., Boucher, M., and Stasko, J. (2026). Sycamore: Characterizing Synthetic Personas for Evaluating Genomics Visualization Retrieval. Preprint.