From draft to testable prototype, in hours
AI builds a clickable prototype in hours, what used to take weeks of design. But too much fidelity fools the test just as much as too little. Deciding what to test, and what can stay fake, is still yours.
You have an idea for a new flow to fix a drop in activation, and need something to test with users by Friday. Before, that meant two weeks of design: wireframe, review, high fidelity, clickable prototype. Today AI generates a navigable prototype from a text description, in hours. The problem shows up when the prototype comes out too polished, full of refined visual detail, and the user reacts to the design, not to the idea you wanted to test.
Every Monday morning it's the same thing. You open the spreadsheet, pull the week's numbers, build the same table as always. Two hours that evaporate, the next week again, from scratch.
A client sends a service contract at five on Thursday afternoon asking for a review by Friday morning. AI drafts the clause in seconds; putting together a coherent contract is still yours.
You need a landing page prototype to test a new positioning by tomorrow. AI generates the full layout in minutes, but it comes out so polished the test ends up being about the design, not about whether the message convinces.
You need to build an onboarding flow prototype to test with a pilot group of new hires. AI generates the process screens fast, but deciding what to actually test in this pilot is still yours.
You have an idea for a new flow to fix a drop in activation, and need something to test with users by Friday. Before, that meant two weeks of design: wireframe, review, high fidelity, clickable prototype, handoff. Today AI generates a navigable prototype from a text description, in hours, with screens, transitions, and even sample data filled in. The problem shows up when the prototype comes out too polished, full of refined visual detail and micro-animation, and the test user reacts to the quality of the design, praising the aesthetics, instead of reacting to the core idea you needed to validate: whether that new flow solves the activation problem or not.
You need an interactive proposal prototype to test with a big client before locking the final format. AI builds the draft fast, but deciding what the test needs to prove is still yours.
You need to simulate a new distribution center layout before investing in the physical change. AI generates the simulation fast, but deciding what to test first is still yours.
You need to build a consent-form prototype to test clarity with real users before publishing. AI generates the draft fast, but making sure the test reflects the real situation is still yours.
You need a functional prototype of a new API to validate with the consuming team before coding it for real. AI generates the stub fast, but deciding what this prototype needs to prove is still yours.
You need a prototype to test a new checkout flow with users by tomorrow. AI generates the clickable flow fast, but choosing the right fidelity so as not to distort the test is still yours.
You need to simulate three market-entry scenarios for Friday's board meeting. AI builds the model fast, but each scenario's premise is still yours.
Let me tell you something about prototypes. Before, putting together something clickable to test an idea was the whole process's bottleneck: two weeks of design, handoff, adjustment, just to arrive at a screen a user could actually click. AI changes this radically. It generates the navigable flow from a text description, in hours. Except that speed brings a new trap nobody had before: the prototype can come out too good, and "too good" ruins a test too.
The core idea of this lesson. AI drastically speeds up the mechanical part of building a prototype: generating the screens, the clickable flow, the sample data. What's still yours, and what decides whether the test is worth anything, is choosing the right fidelity for the question you're asking, and making sure the user reacts to the idea, not to visual polish that shouldn't be there yet.
01What AI speeds up: from description to clickable flow
Think about what changed. You describe the flow you want to test, "an onboarding screen with three steps, the second step asking to connect an account", and AI generates the screens, the transitions between them, and even fills in sample data to make it look real. What used to require a dedicated designer for days now comes out as a navigable draft in hours.
That's great, and it would be a mistake not to use it. The speed gain lets you test far more ideas per month than before, and testing more ideas is exactly what separates a product that learns fast from one that takes forever to find out what works.
02The decision that's yours: which fidelity serves this question
Here's the point AI doesn't decide on its own. Prototype fidelity isn't "the prettier, the better". It's "the right level of finish for the question you're testing".
If the question is "does this flow concept make sense to the user", low fidelity is enough, and sometimes it's better: plain screens, with no defined color, no micro-animation, let the user react to the flow's logic, not the visuals. If the question is "is this specific design clear enough", then yes, you go up to high fidelity, because it's the finish itself that's being tested.
03The risk: an overly polished prototype distorts the user's reaction
Here lives the new danger that ease brought. Since AI can generate high fidelity almost as fast as low fidelity, the temptation is always to ask for the most polished version, "since it comes out easy anyway". The problem is a well-known research effect: users react well to pretty things, and that positive reaction bleeds into their evaluation of the concept, even when the concept itself has a serious problem. You walk out of the test thinking you validated the idea, when you actually only validated that the design was pleasant.
The opposite also exists and is rarer, but happens: a prototype that's too rough, with no visual cues at all, can make the user get stuck on things that don't matter for the test, like "what color is this button", and never get around to reacting to the idea itself. Your job is to calibrate: enough fidelity to feel real enough and not distract, without polish that isn't the question yet.
Learn more: the "halo effect" in prototype testing
The halo effect is when one good, visible quality (here, pretty design) contaminates the evaluation of another quality you wanted to measure separately (here, whether the idea solves the problem). It's a well-documented bias in user research: people report liking a concept more when it's well designed, even when asked specifically about the flow's logic, not about the aesthetics. The simplest defense against it is to deliberately lower the visual fidelity down to the minimum level that still leaves the flow understandable, especially in the early stages when you're testing whether the concept makes sense, not whether the button looks nice.
04The choreography: from idea to test, with the right script
Put it all together in a four-step flow. First, you decide the question the test needs to answer, "does this three-step flow activate more than the two-step one?", before asking for any prototype. Second, you ask AI for the prototype at the fidelity that question calls for, no more, no less, being explicit: "generate in low fidelity, with no defined color, focused only on the flow's structure". Third, you write the test script with the central question at the center, not generic questions like "did you like it?". Fourth, you audit the result: separate reaction to the concept from reaction to the finish before deciding whether the idea works.
Skipping the first step is the most common mistake. Without a clear question beforehand, you end up asking for the "most complete possible" prototype, because it feels safer, and you throw away the chance to test what actually mattered, fast and cheap.
Do it now
Take a real idea you want to test, your real task or another one (a new flow, a screen, a feature).
- Write the central question the test needs to answer, in one sentence.
- Decide the right fidelity for that question: low (tests concept and logic), medium (tests information structure), or high (tests the design itself).
- Ask AI for the prototype explicitly at that fidelity, without letting it "improve" on its own into a more polished version.
- Write two test-script questions that attack the central question from step 1 directly, avoiding generic questions like "what did you think?".
You just calibrated the test for the right question, instead of asking for the prettiest prototype just because it's now easy to get.
Practice
1. Why can always asking for a high-fidelity prototype, just because AI generates it fast, hurt a test with users?
2. What's the step most people skip when asking AI for a prototype, and that causes the biggest waste of the test?
For the board
On fidelityit is your decision, not a default. High fidelity just because it is fast is not a choice, it is laziness.
On the halo effectthe reaction to the design leaks into the evaluation of the idea. You walk away thinking you validated the concept when you validated the aesthetics.
On the questiondefine first what this prototype has to answer. Without that you ask for the most complete one possible and waste the test.
Thanks for the feedback. It helps sharpen the next lesson.