Aetherya
Industry InsightsSeptember 29, 2026Written by: Andrei Dan8 min read

Customer digital twins in market research: what the evidence supports

Individual and population digital twins differ in evidence, privacy, and fit. Learn what current research shows and how to validate a twin before trusting it.

A customer digital twin is a computational model of one person, or one kind of person, built from data about how they answer, choose, and behave. In market research the term covers two different things. One is a twin of a real individual, built from that person's own survey answers or interview. The other is a synthetic persona built from population data and a written profile. The first has stronger evidence behind it. A 2024 study built agents from in-depth interviews with 1,052 Americans and found they reproduced those people's survey answers at 83% of the rate the people reproduced their own answers when asked again. Agents built from demographics alone reached 74%. Twins are useful for screening messages and concepts before fieldwork. They do not replace asking the customer, and a twin is only as current as the data behind it.

"Digital twin" comes from engineering, where a twin of a jet engine takes live sensor data and gets checked against the real engine. Most marketing uses of the term keep the first half and drop the checking. This note is about that gap.

Two kinds of twin

Individual twins. Each model stands for one real person who gave consent and data. The twin is built from their answers to a long survey, a recorded interview, or their purchase history. You can check it, because you can ask the real person the same question and compare.

Population twins. Each model stands for a type of person, defined by traits drawn from census data, published surveys, or a customer segment. Nobody sits behind any single model. You can check the aggregate against real survey results, but not an individual twin.

Vendors often use the same word for both. Ask which one you are buying. The evidence, the privacy obligations, and the right uses differ.

What the research shows

The strongest study so far comes from Park and colleagues. They interviewed 1,052 Americans at length, built a language model agent for each person, and tested the agents on General Social Survey items the agents had not seen. They measured accuracy against a ceiling, which was how consistently each real person repeated their own answers when asked again later. Agents built from the interview reached 83% of that ceiling. Agents built from demographics alone reached 74%. The interview-based agents also narrowed accuracy gaps between racial and political groups. See Park et al., 2024.

Two things stand out. Rich personal data beats a demographic sketch. And even the best twins fell short of how well people predict themselves.

In 2025, Toubia and colleagues at Columbia released Twin-2K-500. It holds answers from 2,058 US participants to more than 500 questions across four survey waves, published so researchers can build twins and test them on held-out answers. See Twin-2K-500. Public benchmarks like this one are how the field will tell twins that work from twins that only sound plausible.

The caution comes from work on population models. Bisbee and colleagues found that synthetic survey samples could match human averages while getting variance and subgroup relationships wrong. See Political Analysis, 2024. A July 2026 preprint found demographic prompting tended to exaggerate differences between groups. See When Synthetic Users Fail.

Put simply, a twin built from what a person told you stands on firmer ground than a twin built from what people like them are supposed to think.

Where twins help

Screening before fieldwork

A team with ten message variants and budget for one human study can run all ten past a twin panel first. The twins will not pick the winner with certainty. They usually catch the variants that confuse, overclaim, or miss the audience, and the human study can then compare the two or three that remain.

Asking the question nobody asked

A survey closes when it closes. If the twins are built from that survey, a team can ask them a new question a month later and get a first read. This is where individual twins are most useful. It is also where they drift, because the person has moved on and the twin has not.

Rehearsing with hard-to-recruit groups

Some audiences are expensive to reach, such as senior buyers, clinicians, or people in small markets. A twin panel built from a past study of that group lets a team rehearse before spending recruitment budget. The report should call it rehearsal.

Where twins fail

  • New behaviour. A twin predicts from what the person already said or did. It is weakest on reactions to things the person has never seen.
  • Changed circumstances. A twin built in March does not know the person lost their job in June.
  • Social and physical context. Most twins answer alone, in text. Real purchases often involve a partner, a committee, or a shop floor.
  • Group differences. Population twins in particular can exaggerate how far apart segments are. Check any headline segment gap against real data before acting on it.

An individual twin is personal data. Under GDPR, a model that predicts a named person's opinions and choices is processing that data, and it may count as profiling. The person needs to know, the purpose needs a lawful basis, and deleting the source data should delete the twin.

Population twins avoid most of this, because no single model maps to a person. That is one reason many teams start with population twins and build individual twins only for customers who have opted into a research panel.

A validation routine

Whichever kind you use, check it before you trust it:

  1. Hold out a sample of real answers the twins never saw.
  2. Ask the twins the same questions.
  3. Compare at the level you plan to use, whether that is individual answers, segment averages, or the ranking of options.
  4. Repeat after every model or data change.
  5. Keep a record of where the twins were wrong, not only the overall score.

If a vendor cannot show you this for your kind of decision, the accuracy claim is marketing.

A worked example

Imagine a subscription app with 3,000 customers who completed an onboarding survey last year. The team wants to test three pricing-page explanations before a redesign.

It builds a population audience in Thesia from the survey's segments rather than from named customers, so no model stands for a real person. It runs all three explanations past that audience and asks each model to estimate its bill, name the plan it would choose, and state its biggest doubt. One explanation leads most of the audience to a plan that does not fit its usage. The team drops it.

The other two go to eight customer interviews and then a live test. The team records where the simulated doubts matched what customers said, so it knows how much weight to give the next screen.

This is a protocol example, not a reported Aetherya outcome.

Final answer

A customer digital twin models a person or a type of person. Its value depends on the data behind it and on whether anyone checked it. The best current evidence shows interview-grounded twins reproducing people's survey answers at about 83% of their own consistency, with demographic sketches doing worse. Use twins to screen and rehearse before human research. Validate them on held-out answers for the decision you care about. Treat individual twins as personal data.

A twin nobody compared with the real customer is a persona with a better name.

Next step: Explore Thesia to build an audience from your own customer segments and screen decisions before fieldwork.

Sources

Aetherya

Cognitive Simulation Research & Technology

Related