Research and methodology

Synthetic Audiences: When They Work and When They Don't

Synthetic audiences are neither magic nor useless. The research is now clear enough to say where they match real people, where they break, and how to use them without fooling yourself.

Dhruv Kashyap · 11 min read · October 3, 2026

Synthetic audiences are neither magic nor useless. The research is now clear enough to say where they match real people, where they break, and how to use them without fooling yourself. The debate about synthetic audiences usually has two camps. One says AI personas will replace market research. The other says they are autocomplete in a costume and cannot tell you anything real. Both are wrong, and the research is now clear enough to say why. The question is not whether synthetic audiences work. It is what they work for. This post walks through the evidence on both sides, with references, and ends with the rules we follow at Tesemble.

What is a synthetic audience?

A synthetic audience is a set of AI-generated personas, each defined by attributes such as role, industry, company size, geography, age or buying behavior. The whole audience is exposed to the same stimulus, such as a product concept, a message or a survey question, and the responses are compared across segments. The individual personas are not real people. The value, when there is value, is in the pattern across the audience.

Where synthetic audiences work

  1. Reproducing aggregate patterns across groups The foundational result came from political science. In "Out of One, Many" (Argyle et al., Political Analysis, 2023), researchers conditioned GPT-3 on thousands of real sociodemographic backstories from US surveys. The resulting "silicon samples" reproduced the response distributions of many human subgroups with surprising nuance. The authors called this property algorithmic fidelity. The key word is distributions. The model was not predicting what one person would say. It was reproducing how groups tend to respond.

  2. Ranking which treatment works better In July 2026, Nature published a study by Ashokkumar, Hewitt, Ghezae and Willer that tested whether GPT-4 could predict the results of 70 pre-registered, nationally representative survey experiments. Simulated responses predicted the actual treatment effects with accuracy comparable to pooled human forecasters, and held up for studies that were not published before the model's training cutoff. This is the result most relevant to GTM. Comparing two messages, two offers or two positionings against the same audience is a treatment comparison. Synthetic audiences are best at telling you which option is likely to do better.

  3. Purchase intent and product concepts Researchers at PyMC Labs and Colgate-Palmolive tested synthetic consumers against 57 real product concept surveys covering 9,300 human responses (Maier et al., 2025). When the models answered in free text and those answers were mapped to a rating scale, they reached 90% of human test-retest reliability and produced realistic response distributions. The same paper found an important failure: when models were asked directly for a number on a scale, the distributions were unrealistic. How you elicit the response matters as much as the model.

  4. Perceptual maps and brand positioning In Marketing Science (Li, Castelo, Katona and Sarvary, 2024), language models generated brand similarity judgments and product attribute ratings. The resulting perceptual maps agreed with human survey data at rates above 75% for the categories tested. That is useful for positioning work: understanding where your brand sits relative to alternatives.

  5. Realistic economic behavior, with caveats A Harvard Business School working paper, Using LLMs for Market Research (Brand, Israeli and Ngwe), found that GPT responses showed realistic economic patterns such as declining marginal utility and state dependence. It also found that willingness-to-pay estimates were sometimes comparable to human studies and often inaccurate, occasionally with the wrong sign. We come back to that below.

  6. Individuals, when grounded in real data Stanford researchers (Park et al., 2024) built agents from two-hour interviews with 1,052 real people. The agents replicated participants' answers on the General Social Survey 85% as accurately as the participants replicated their own answers two weeks later. Note what made this work: each agent was grounded in a long interview with a specific person. A persona built from a job title and a company size does not have that grounding, and should not be expected to perform like this.

Where synthetic audiences fail

  1. Predicting a specific individual A July 2026 benchmark, When Synthetic Users Fail (Chen, Zhu and Zheng), tested four models on the General Social Survey and the World Values Survey. At the individual level, no model beat the strongest simple baseline.

If you want to know what one particular buyer will do, a synthetic persona is the wrong tool.

  1. Too little variation In Synthetic Replacements for Human Survey Data? (Bisbee et al., Political Analysis, 2024), ChatGPT personas produced average scores close to a major US election survey, but with less variation than real respondents. The averages looked right; the spread did not. That makes the data unreliable for statistical inference.

  2. Stereotyping and flattening A Nature Machine Intelligence paper (Wang, Morgenstern and Dickerson, 2025) showed that language models standing in for human participants tend to both misportray and flatten demographic groups, representing them as more uniform and more stereotypical than they are. The 2026 When Synthetic Users Fail benchmark found the same pattern in a form that matters directly for GTM: models over-weight demographic identity. On segment-targeting tasks, they inflated the gaps between segments two to four times and pointed to the wrong segment in roughly half of the US cases tested. Larger models did not fix it. This is the single most important caveat for ICP work. A synthetic audience can point you toward a segment. It will likely exaggerate how different that segment is.

  3. Defaulting to dominant viewpoints Whose Opinions Do Language Models Reflect? (Santurkar et al., ICML 2023) compared model opinions with 60 US demographic groups and found substantial misalignment. When a persona is underspecified, models tend to fall back on the most common viewpoint in their training data. Thin persona definitions produce generic answers.

  4. Optimism and agreeableness Nielsen Norman Group ran a study comparing synthetic users with real ones for an online course product (Rosala and Moran, 2024). Synthetic users claimed to finish every course; real users admitted dropping off partway. Synthetic users praised discussion forums that real users found contrived. Synthetic needs came back as long, flat lists without priorities. Synthetic personas are prone to telling you what a reasonable person should do, not what real people actually do.

  5. Absolute numbers Two of the positive studies above carry the same warning. The Nature study found that simulated predictions systematically overestimated effect sizes. The HBS study found willingness-to-pay estimates that were often inaccurate. Direct numeric ratings produced unrealistic distributions in the PyMC Labs study. Synthetic audiences are much better at which than at how much.

  6. Knowing too much In a study of classic experiments (Aher, Arriaga and Kalai, ICML 2023), language models reproduced several well-known human findings, but in a wisdom-of-crowds task they showed a "hyper-accuracy distortion": the simulated people knew more than real people do. A synthetic buyer may understand your category better than a real buyer skimming an email.

The pattern

Put the evidence side by side and a clear shape appears.

Synthetic audiences are useful for Synthetic audiences are unreliable for
Comparing options against the same audience Predicting what a specific individual will do
Ranking which segment responds more strongly Measuring exactly how large a segment difference is
Surfacing likely objections and confusion Absolute numbers: demand, conversion rates, willingness to pay
Generating and prioritizing hypotheses Validating a hypothesis on their own
Exploring many ideas quickly before real research Replacing interviews, pilots and real campaigns
Well-documented categories and roles Novel categories and under-represented groups

You've got more GTM ideas than you can test.

Test them with Tesemble first.

Try for free

10,000 free credits at launch. No credit card.