Why “Synthetic” User Testing Fails the Product

AI-driven “synthetic” user testing is being pitched as a replacement for UX research: AI-generated “customers” answer questions, and AI agents perform usability tasks as if they were real people. The promise is that these personas mimic actual customer behavior inside the real product, giving you UX research without the users. The reality is that this approach is dangerous, risky, and expensive — and it typically erodes the very user value it claims to protect.

The Illusion of Fast, Cheap, and Plausible

The appeal of synthetic testing is obvious. It is fast, cheap, and easy to re-run. It requires no participant recruitment, no lengthy debates, and no uncomfortable questions that challenge existing assumptions. It can spin up thousands of AI personas at once to uncover common journeys and navigation patterns.

But that convenience masks a fundamental flaw. LLMs are trained to produce the most “plausible” output — statistically average behavior drawn from patterns on the web. Those patterns do not represent your customer base, and those people do not exist. Without careful scoping and curation, the AI’s default output is a generic guess that matches no real user segment.

Even with perfect prompting, LLMs cannot generate unexpected findings. They can only recombine what you already asked about. Real UX research, by contrast, often reveals insights that shift priorities or reframe the problem entirely. Those insights come from observing behavioral clues, emotions, and contradictions — a person doing the opposite of what they said. That cannot be replicated with text generation.

Not “Better Than Nothing”

Some argue that synthetic research, while imperfect, is preferable to no research at all. That reasoning is flawed. As Pavel Samsonov points out, things that sound like what customers might say are worthless. Things that customers actually have said, done, or experienced carry inherent value — even if they require interpretation. AI-produced “insights” create an illusion of customer experience that never happened. At best they are guesses; at worst they are misleading and non-applicable.

Relying on AI-generated insights alone is not a pragmatic substitute. It is closer to reading tea leaves — confident, but disconnected from reality.

The Hidden Cost of Mechanical Decisions

Automation always carries a cost. In the case of AI research, that cost shows up as mechanical decisions that favor uniformity and erode quality. As Maria Rosala and Kate Moran note, AI research will almost certainly be misrepresentative, and without real research, you will not catch those inaccuracies. Decisions made without talking to real customers are not just risky — they are harmful and expensive to correct later.

Synthetic testing also assumes people fit neatly into defined boxes. Human behavior is shaped by experiences, situations, and habits that cannot be reduced to text generation. The result is that AI strengthens existing biases, supports hunches, and amplifies stereotypes instead of challenging them.

Use AI for Triangulation, Not Verification

AI can offer useful starting points early in the process. The danger is treating its output as verified insight, presented with an unwarranted level of certainty. A more reliable path is to start with human research conducted with real customers using the real product. After that, you can apply AI to check whether you missed something critical in interviews. AI can enhance UX research; it cannot replace it.

Avoid the temptation to “validate” AI findings with a few user tests. Once a seed of insight is planted, it is easy to see evidence for it everywhere — whether or not it is actually there. Instead, study actual customers first, then triangulate with other data sources. If analytics and AI desk research align with your findings, you have far stronger grounds to move forward.

Real Users Before Synthetic Ones

The push to automate UX work overlooks what good design actually requires: critical thinking, observation, and planning. Cleaning up after AI-generated output frequently takes more time than doing the real work. The value of speaking directly to people who use your product remains immense — one day with a real customer is worth more than an hour with a thousand synthetic users pretending to be human.