The prototype that talks back: test with a synthetic user before the real one, never instead of them
Before scheduling the expensive test with real people, simulate a synthetic user navigating the prototype. It catches gross flow errors in minutes, but carries the bias of the model simulating it: it's for filtering beforehand, never for replacing the real thing.
You redesigned a flow after identifying the pattern that was getting users stuck, and now you want to test it before spending the expensive time of scheduling real people. Simulating a synthetic user, a persona based on what you already know from real research, running the flow for the first time, catches the gross error in minutes: the ambiguous button, the step nobody understands right away. Except that synthetic user will never react with the same surprise or the same uncomfortable silence as a real person, because it carries the biases of the model simulating it. It's for filtering beforehand, never for replacing.
Before running a new expense approval process with the real team, you ask AI to simulate a typical employee trying to follow the new flow, with the real profile you've already mapped in previous interviews. It gets stuck exactly on the cost center field, ambiguous for months. You fix it before exposing the process to the whole team. But the decision to release the process to production never rests on that simulation alone: a pilot with a real department is still the test that counts.
Before taking a new contract draft to the client, you ask AI to simulate a typical client's reading, one without legal training, pointing out where a clause gets confusing. It flags two ambiguous clauses even before you schedule the review meeting. That saves a whole round of back-and-forth. But the draft's final approval never rests on the simulation alone: the real client, with their real doubt, still reviews it before signing.
Before launching a new campaign, you already use the synthetic customer to test offer and message, the way the marketing track teaches. The same logic applies here, with the same risk: the synthetic customer catches obvious clarity errors quickly and cheaply, but carries the bias of the model simulating it, and never replaces the test with a real audience before the launch counts for good.
Before rolling out a new employee onboarding process, you ask AI to simulate a typical new hire going through the first steps, with the profile exit interviews have already mapped. It gets stuck at the same point old reports pointed to. You fix it before the first real hire goes through it. But validating the process never rests on the simulation alone: the first real group of new hires is still the test that counts.
Before scheduling the usability test with a real customer, you ask AI to simulate the persona the team has already validated in research, navigating the new feature's prototype. It flags the same friction thirty real sessions already suggested, in minutes. That confirms it's worth spending the real user's expensive time testing this flow. But the decision to launch the feature never rests on the simulation alone: the test with real people is still what decides.
Before taking a new sales pitch to the client, you ask AI to simulate the skeptical buyer who usually stalls the negotiation, reacting to the script. It flags the exact point where the pitch sounds forced. The salesperson adjusts before the real call. But closing the deal never rests on the simulation alone: the real client, with their real objection, still decides whether to buy.
Before rolling out a new procedure with the operations team, you ask AI to simulate a typical operator following the written step-by-step, pointing out where the instruction is ambiguous. It gets stuck at the same step last month's incident had already revealed. You fix the procedure before the real shift follows it. But validating the procedure never rests on the simulation alone: the first real shift that runs the procedure is what confirms whether it works.
Before publishing a new policy, you ask AI to simulate a typical employee reading the text and pointing out where they'd be unsure about what's allowed. It flags a real ambiguity in the gifts section. You fix it before the policy takes effect. But the final approval never rests on the simulation alone: training with real employees, and their real questions, is still the test that confirms whether the policy came out clear.
Before exposing a new API to partner teams, you ask AI to simulate a typical developer trying to integrate, following only the documentation. It gets stuck at the same endpoint the team already suspected was poorly described. You fix the documentation before the first real partner tries to integrate. But validating the API never rests on the simulation alone: the first real partner, with their real environment, is who confirms whether the integration works.
You just redesigned the second step of onboarding after confirming, with a traceable citation, that the problem was cognitive overload on the initial form. The prototype is ready in Figma, the test with a real user is only scheduled for next week because the recruiting calendar doesn't move fast. Instead of waiting, you build a synthetic persona based on the real profile that showed up across the 84 interviews: someone without much technical familiarity, anxious to finish quickly, on their phone. You ask AI to "be" that person and navigate the prototype step by step, narrating what they feel on each screen. In fifteen minutes it gets stuck exactly where you suspected, and flags two more friction points you hadn't seen. You don't cancel next week's real test: you go into it with three hypotheses already tested, ready to confirm or knock down with real people.
Before presenting a new decision model to the board, you ask AI to simulate a skeptical advisor reading the rationale, pointing out where the argument is weak. It finds two holes in the logic before the real meeting. That saves a round of rework. But the strategy's approval never rests on that simulation alone: the flesh-and-blood board, with their own questions, still decides.
Whoa, let me tell you about a good shift that still hides a trap inside it. The prototype is ready, but the test with a real user only happens next week, because the recruiting calendar doesn't move fast. You can wait around, or you can run the flow right now with a simulated person, based on what you already really know about who uses your product. The problem is that simulated person isn't a human, and treating them as if they were is the fastest way to make a confident, wrong decision.
The core idea of this lesson. After confirming the theme with a traceable citation (previous lesson), you build the prototype and want to test the flow. AI today lets you simulate a synthetic user: a persona based on real research data, simulated by a model, navigating your prototype and narrating what it feels. It's a fast, cheap filter, that catches gross flow errors in minutes, before you spend the expensive time of scheduling real people. It isn't a substitute: the synthetic user carries the same biases as the model simulating it, and never reacts with the surprise, frustration, or uncomfortable silence a real human reveals. This lesson's golden rule is short: synthetic before the real one, never instead of the real one.
01What the synthetic user catches well, and catches fast
Let's start with the gain, which is real. You build a persona based on the profile your research has already validated, not a generic mannequin, and ask AI to navigate the prototype as if it were that person, narrating what they understand, what confuses them, where they get stuck. In fifteen minutes, without scheduling anyone, without waiting for recruiting, you've already caught the ambiguous button, the confusing label, the step that requires knowledge the persona doesn't have. It's exactly the kind of gross error that would come up in a real test, just earlier and cheaper.
The value here isn't replacing research, it's filtering before it. Think about the cost of scheduling a test with five real people: recruiting, calendars, their time, your time moderating. If the prototype has an obvious flow error, you don't want to discover that in the expensive session with real people, you want to arrive there already past the gross stumbles, using that expensive time to catch what only a real human reveals.
02Why it's never the real thing: the bias it carries
Here's the caution that separates whoever uses this well from whoever fools themselves. The synthetic user isn't a person with their own opinion formed by a lifetime of experience: it's a language model representing what it learned that a person of that profile would say. There's research showing that language models, when simulating a group's opinion, tend to pull toward what's most common and best represented in the training data, and flatten the real variation that exists between real people. Translating that to your prototype: the synthetic user tends to react in a more predictable, more "average" way than a real person would.
And there's a gap no model simulates well: the uncomfortable silence. In a real session, you see the person hesitate, wonder if they're doing something dumb, or give up quietly because they don't want to look stupid in front of you. It's an extremely rich UX signal, and it doesn't exist in the synthetic user, because it doesn't feel embarrassment, doesn't fear judgment, doesn't have a body that freezes up. It gives you an articulate narration of what it "understands" or finds "strange." It doesn't give you the moment when a real person simply stops trying.
03The golden rule: before the real thing, never instead of the real thing
Put both halves together and the rule is short to write and easy to forget in the rush: use the synthetic user before the real test, never instead of it. It isn't a rule of methodological purity for its own sake. It's protection against a specific, expensive mistake: launching a flow change to all of production based only on a simulation that came back clean, and finding out later, with real customers, the friction only a real body reveals.
In practice, this becomes a simple discipline to follow. Every time the synthetic user passes a flow cleanly, ask: does this mean the flow is good, or does it mean the model that simulated this person doesn't know how to reproduce the doubt a real person would have there? If the answer leaves you in doubt, that's a sign next week's real test is still what decides, not a formality to check off after you're already sure.
Now build
Take the flow you want to test in your real task (or your real prototype) and build a synthetic persona based on a profile your research has already validated: level of technical familiarity, hurry, device, whatever is relevant to your case.
Ask an AI to "be" that person and navigate the flow step by step, narrating what they understand, what confuses them, where they hesitate, exactly as if they were thinking out loud during a usability test.
At the end, list the three friction points it flagged, in the order they appeared. For each one, mark: is it a gross flow error (label, button, confusing step) that can be fixed before the real test? Or is it a point where you suspect only a flesh-and-blood person, with their frustration or their silence, will reveal the real size of the problem?
Fix whatever is gross before next week's real test. Carry the rest as a hypothesis to confirm with real people. Never cross the real test off the calendar because the synthetic one "came back clean."
Practice
1. What's the main value of testing a flow first with a synthetic user?
2. Why does the synthetic user never replace the test with a real person?
3. The synthetic user passed cleanly through a new checkout flow. A colleague suggests canceling next week's real test to save time. What do you say?
4. What kind of UX signal does the synthetic user have the most trouble reproducing for real?
Fair? Let's close the message of this lesson together. The synthetic user is a real filter: it catches gross flow errors in minutes, at the cost of a conversation, not a whole recruiting effort. But it's made of the same material as the model simulating it, with the same biases, and without the body that hesitates, fears, and sometimes gives up in silence. The golden rule doesn't change: synthetic before the real one, never instead of the real one. Use the filter to arrive cleaner at the test that really decides, and never let the filter decide for you. Next.
For the board
On the gainit is a fast filter that clears the obvious before the expensive session with real people.
On the limitthe model flattens the variation between people and feels no embarrassment, no fear of judgement, no hesitation.
On the golden rulesynthetic before the real thing, never instead of it. Passing the filter cleanly does not excuse the real decision.
Thanks for the feedback. It helps sharpen the next lesson.