A third of the time, synthetic research points the wrong direction

Vice President, Digital Experience and Engagement
  • Twitter
  • LinkedIn

James Bisbee and colleagues ran the cleanest test of this anyone has published. Writing in Political Analysis in 2024, they generated synthetic survey responses and compared them against the American National Election Studies, a large, rigorous, real-human dataset.

On averages, the synthetic data looked fine. Mean responses landed within one standard deviation of the real thing, which is exactly why this technique is easy to sell.

Then they looked at relationships rather than averages. Estimating how demographic characteristics related to opinion, 48 percent of the synthetic coefficients differed significantly from the real ones, and among those, the sign of the effect flipped 32 percent of the time.

Roughly a third of the time, the synthetic data did not merely get the magnitude wrong. It pointed the opposite direction.

Two further findings from the same study should stop anyone planning to use this in place of customer research. Synthetic responses showed far less variation than human ones, so severely compressed that a power analysis on synthetic data suggested 33 respondents would suffice where the real data required more than 280. And the same prompt produced materially different distributions in April, June, and July of 2023, with substantial mean reversion after a model update.

Read those together. Synthetic research understates disagreement, which makes you feel more certain on a tenth of the evidence, while your findings quietly change whenever the vendor ships an update.

That is the risk profile. Not that AI produces obviously bad research, which would be easy to catch. It produces plausible research with the variance sanded off and the direction occasionally reversed.

AI can still do a great deal for research teams. Used carelessly, it produces a research function that generates more documents and knows less.

 

Where AI is genuinely strong

Independent assessments converge on roughly the same boundary: AI helps most in planning and analysis, meaning the stages that involve organizing and retrieving information, and least during sessions, meaning the stages that require being present with a person. Nielsen Norman Group's testing reached the same conclusion, that AI "can speed up certain research tasks but is currently most helpful in the planning and analysis stages."

That boundary is worth taking literally, because the stages it excludes are where the value is created.

Before fieldwork, AI is excellent at the archaeology nobody has time for. Most organizations already know more than they can retrieve. Studies from two years ago live in a deck in someone's folder. Support ticket themes never made it to the design team. A prior team ran almost this exact study and nobody remembers.

Pointing AI at that corpus to answer "what do we already know about this audience, this journey, this question" routinely prevents duplicate studies. It is the least glamorous application and often the highest return, because a study you did not need to run saves six weeks, not six hours.

After fieldwork, AI handles the volume work well: transcription, translation, sanitizing personal information, preliminary coding, grouping similar statements, first-pass quantitative analysis. NN/g notes these "accelerate the initial steps in your analysis process." Note the word initial.

 

Where it fails, specifically

During sessions, it is close to useless. NN/g is direct about this: current tools "cannot understand and interpret users' actions or nonverbal interactions." Moderated research is a behavioral discipline. The value comes from noticing the pause before the answer, the hand that hovers and retreats, the participant who says the flow was fine while visibly having found it anything but fine. A transcript captures the words and discards the study.

In interpretation, frequency masquerades as importance. AI surfaces what was said most often. Research value frequently sits in the opposite place: the single participant whose experience exposes an accessibility barrier, a compliance exposure, or a trust failure that the other eleven never encountered because they are not in that situation. Nine people mentioning a slow load time is a performance ticket. One person unable to complete enrollment with a screen reader is a legal and moral problem. AI ranks the first higher.

You cannot fix representation by prompting a persona. This is the finding that should end the most common sales pitch in the category. Santurkar and colleagues at Stanford evaluated language model opinions against 60 US demographic groups and found substantial misalignment with the real views of those groups, with older and widowed respondents among the worst represented. The critical result: that misalignment persisted even after explicitly steering the models toward the specific demographic group in question.

So instructing a model to answer as a 62-year-old procurement manager in a mid-sized manufacturer does not produce that person's perspective. It produces the model's stereotype of that person, and the gap does not close when you ask more precisely. Zhicheng Lin, writing in Advances in Methods and Practices in Psychological Science, names this the identity-essentialization fallacy: treating demographic labels as fixed, homogeneous categories, which reinforces stereotypes rather than capturing how people actually differ.

The teams most tempted by synthetic participants are the ones researching hard-to-recruit populations. That is precisely where training data is thinnest and confident fabrication is most likely.

Sycophancy compounds all of it. Synthetic participants agree. They return favorable feedback that real users withhold, and they generate undifferentiated lists of considerations without any sense of which one actually decides the matter. Real people prioritize because real people have constraints, deadlines, and a limited tolerance for effort. NN/g's testing found the same pattern in a UX context: synthetic users reporting they had completed all seven online courses in a series that real participants had abandoned after three, citing time pressure.

 

The legitimate use of synthetic participants

There is one, and it is narrower than most vendors suggest.

Lin's framing is the most useful available: LLMs are pragmatic simulation tools for hypothesis testing and rapid prototyping, and their outputs require human validation, which is precisely what undermines the case for using them as replacements. If you have to check it against real people, you have not saved the research.

In practice that means: pressure-test a discussion guide before spending on recruitment, explore how different roles might read a concept in order to decide what to ask, and generate hypotheses that fieldwork will then confirm or kill.

The test is simple: synthetic output should change what you investigate. It should never change what you conclude. The moment a synthetic finding appears in a stakeholder deck as evidence of customer need, the research function has started manufacturing confidence rather than reducing uncertainty, and someone will eventually build a roadmap on it.

Label it in every artifact. Not as a compliance formality. Because six months later, nobody will remember which findings came from people.

 

Making existing research findable

The most underrated application has nothing to do with generating anything.

Research loses most of its value to retrieval failure. Findings live in decks, folders, repositories, and individual memory. A designer starting work on enterprise administrator onboarding has no practical way to discover that a relevant study exists, so the organization pays twice: once for the study nobody found, once for the study they run instead.

An AI layer over a well-maintained repository changes the interaction from browsing to asking. The questions your designers should be able to put to it directly, in their own words:

What do we already know about onboarding for enterprise administrators? Which usability issues have recurred across multiple studies? Which journey stages have the least research coverage? What did we learn last time we tried this?

If a designer on your team cannot get an answer to those in under a minute today, that is the gap.

Two conditions determine whether this works, and both are unglamorous. The underlying research has to be structured and consistently documented, and permissions have to be correct, because research repositories contain participant data. Neither condition is created by buying a tool. This is a repository discipline problem that AI makes worth solving, not a problem AI solves.

 

The controls that make this defensible

Research data is among the most sensitive material a design organization handles. Recordings, transcripts, personal information, and unreleased product plans, frequently in the same file.

The minimum standard: named approved systems with appropriate data handling, personal information removed before anything enters a model, human validation of any finding that will influence a decision, source documentation so a claim can be traced to a participant, and a hard visible separation between real and synthetic material.

One more that teams forget: check what the analysis missed. AI summarizes toward the center of a dataset. Ask explicitly which perspectives are underrepresented in this sample and which participant experiences did not fit the pattern. The outliers are usually where the next study lives.

 

What this does to the researcher's job

It removes the administrative floor. Less transcript wrangling, less deck reformatting, less rebuilding a summary for the fourth audience.

What fills the space is the part that was always the actual job and rarely got enough of the calendar: more time with customers, more investigation of behavior that does not make sense yet, more presence in the product decisions where research either lands or gets ignored.

NN/g's working metaphor is that AI behaves like an intern. Useful with clear instructions, capable of real volume, and never the person who signs off on the conclusion.

That is the right relationship. Teams that get this wrong do not usually fail loudly. They just gradually start knowing less about their customers while producing more documents about them, and nobody notices until a launch goes badly.

Building the AI-enabled experience

Building the AI-enabled experience design workflow for in-house teams

Read our perspective

 

Sources

Primary research

 

Practitioner corroboration