Connecting live research: Dovetail, Hotjar, and tickets as your corpus
Without real data connected, AI synthesizes UX insight from the average training user, not your real user. This lesson connects AI to the research repository (Dovetail), session recordings (Hotjar), support tickets, and interview transcripts, and teaches the anchored synthesis principle: every insight points back to a real source, or it doesn't count.
You ask a generic AI "why do users abandon sign-up at the second step?" and it returns a well-written answer, full of plausible reasons: long form, lack of trust, sensitive data requested too early. It sounds right because it's what usually happens in similar products. Except none of those sentences points to a real session in your product, to a real support ticket, to an interview someone actually conducted. It's a beautiful, generic insight, one that might even happen to be right, but that nobody can defend in a meeting when asked "where's the source for that?".
You ask a generic AI why delinquency rose this quarter, and it returns plausible reasoning about interest rates and committed income. It sounds well written. Except no sentence points to a real statement, to your specific client's actual payment history. It's a textbook hypothesis, not a conclusion that survives the auditor's question "where did that number come from?".
You ask a generic AI why that type of contract tends to generate litigation, and it returns coherent legal reasoning, full of plausible case law. It sounds solid. Except no sentence points to a real case at your firm, to that specific client's history. It's generic doctrine, not a thesis that survives the judge's question "cite the exact precedent".
You ask a generic AI why the campaign didn't convert, and it returns nice reasoning about tired creative and saturated audience. It sounds plausible. Except no sentence points to real data from your campaign, to the actual comment the client left on the ad. It's marketing textbook theory, not a conclusion that holds up to the director's question "show me the data".
You ask a generic AI why turnover rose on that team, and it returns plausible reasoning about workload and lack of recognition. It sounds well written. Except no sentence points to a real exit interview, to an actual comment from that specific team's climate survey. It's generic HR theory, not a conclusion that survives the director's question "who said that?".
You ask a generic AI why the new feature has low adoption, and it returns plausible reasons: confusing onboarding, unperceived value. It sounds right. Except no sentence points to a real support ticket, to an actual recorded session of a user getting stuck right there. It's an experienced PM's hunch, not a traceable insight that survives the team's question "where's the evidence?".
You ask a generic AI why leads aren't closing, and it returns plausible reasoning about price and timing. It sounds well constructed. Except no sentence points to a real recorded call, to an actual rejection email from that specific prospect. It's an experienced salesperson's intuition, not a conclusion that survives the manager's question "show me the transcript".
You ask a generic AI why the night shift's SLA got worse, and it returns plausible reasoning about reduced staffing and fatigue. It sounds correct. Except no sentence points to a real record from that shift, to the actual ticket that got stuck. It's an operations-manual hypothesis, not a root cause that survives the manager's question "which ticket proves that?".
You ask a generic AI why a control failed the audit, and it returns plausible reasoning about process failure. It sounds well written. Except no sentence points to a real evidence record, to the actual log from that incident. It's a compliance-manual assumption, not a finding that survives the auditor's question "where's the evidence?".
You ask a generic AI why the bug only shows up in production, and it returns plausible hypotheses about environment and data contention. It sounds technical and correct. Except no sentence points to a real log, to the actual stack trace from that specific incident. It's a senior engineer's guess, not a diagnosis that survives the team's question "show me the log".
You ask a generic AI "why do users abandon onboarding at the second step?" and it returns a competent, well-structured answer: form too long, permission requested too early, no perceived value at the start. Each sentence sounds plausible because it's exactly what usually happens in onboarding for similar products, the pattern it learned from seeing thousands of cases. You almost put this straight into the research report. Except none of those sentences points to a real session recorded in your product's Hotjar, to a real support ticket, to a real quote from an interview with a user who actually abandoned right there. It's a UX-textbook insight, correct in general, and you have no way to defend it in a product meeting when someone asks "where's the evidence that this is really what's happening, in our case?". The right answer requires connecting AI to your real corpus before asking for the synthesis, not after.
You ask a generic AI why the expansion into the new market is slow, and it returns competent reasoning about cultural barriers and local competition. It sounds correct. Except no sentence points to a real interview with a customer in that market, to an actual field report. It's a strategy-book hypothesis, not an analysis that survives the board's question "what was the primary source?".
Whoa, let me show you a scene that repeats itself all the time on UX teams that already use AI, and that goes unnoticed precisely because the answer comes out so well written. You ask AI why users abandon a flow, it returns a competent paragraph, full of plausible reasons, and the team approves that as if it were research. Except it isn't research. It's a good generic guess, dressed up as a conclusion. This lesson exists to close that gap, and it borrows a beat other tracks in this course have already used: connecting AI to the real source of your own data, before asking it to think.
The core idea of this lesson. Without a real connection, AI synthesizes a UX insight from what it learned to be common in similar products, the average user of its training, not your real user. The fix is connecting AI to your live research sources: the research repository (Dovetail), session recordings and heatmaps (Hotjar), support tickets, interview transcripts. It's the same "connect to the data" beat other tracks in this course have already installed in finance and marketing, except here the data you connect is your real user's experience. And along with that comes a principle that organizes everything: anchored synthesis. Every insight AI generates needs to point back to a real source, a citation, a session, a specific ticket. If the insight doesn't point anywhere, it isn't an insight, it's a well-written opinion.
01The corpus you already have, scattered and disconnected
Here's the irony. Most UX teams already produce a huge corpus of real data every single day, just scattered. There's a session recording sitting in Hotjar, with the exact moment the user hesitates before clicking. There's an interview transcribed and tagged in Dovetail, with the participant's literal words. There's a support ticket in the help desk, telling you exactly where the person got stuck. There's an app store review, complaining about a specific button. That's the raw material of the truth about your user, and in practice it lives isolated, each source in its own corner, not talking to the others or to the AI you use to synthesize.
When you ask AI without connecting any of that, it has no choice but to fill the gap with what it learned elsewhere: the average pattern of the entire internet about onboarding, about checkout, about any common flow. It's not the AI's fault. It's the wrong question asked in the wrong order. First you connect the source, then you ask for the synthesis, never the other way around.
02Anchored synthesis: the insight that points somewhere
Now we reach the principle that holds this whole lesson up, and one you'll use in every research effort from here on. We call it anchored synthesis, and the rule is simple to state and hard to keep without discipline: every insight AI generates needs to point back to a real, specific source. Not "users tend to complain about long forms" floating in the air. Yes: "three participants from the March interview round (session 04, session 07, session 11) literally used the word 'tiring' when reaching the address field, and the Hotjar recording of session 12 shows an 8-second hesitation right there".
Notice the practical difference. The first type of sentence is unfalsifiable and useless at the same time: it sounds true, but nobody can verify or challenge it. The second type is verifiable: anyone on the team can open session 12 and see the hesitation with their own eyes. It's the difference between an insight you can defend in a meeting and one you just hope nobody questions.
For the board. Before accepting any AI-generated UX insight, ask one question only: "points where?". If the answer is a citation, a session, a ticket, an interview number, the insight stays. If the answer is silence or "it's what usually happens," the insight is a hypothesis, not a conclusion, and it needs to be marked as such before going into the report.
03Connecting for real: what changes in your workflow
In practice, connecting AI to live research means giving it real access to the sources before asking for any synthesis, the same way finance connects AI to the source spreadsheet and marketing connects AI to the product catalog. Three concrete connections do most of the work:
- Research repository (Dovetail or equivalent). This is where the interview already lives, transcribed and tagged. Connecting here means that, when you ask about a topic, AI fetches the literal quote before summarizing, instead of inventing a generic user statement.
- Session recording and heatmap (Hotjar or equivalent). This is real behavior, without the filter of memory or of the participant's politeness (what they do is more reliable than what they remember doing). Connecting here gives AI the data on where hesitation, the wrong click, and abandonment really happen.
- Support tickets and reviews. This is the spontaneous complaint, not prompted by a researcher's question. Connecting here brings the exact vocabulary the user uses when genuinely frustrated, without the filter of a formal interview.
The antidote here is the same as the context debt you already saw in lesson 1.1 of this course: an AI that reads your real corpus before acting works with your context, instead of guessing. Applied to research, reading first means fetching the citation, the session, the ticket, before writing any synthesis sentence. Acting without reading first is the exact recipe for a loose insight nobody can defend.
Why a large corpus without tags doesn't solve anything on its own
Connecting the source isn't just dumping everything into a repository and waiting for magic. A Dovetail full of interviews without consistent tagging is almost as useless as having no interviews at all, because AI (and humans) can't navigate to the right source when they need to. The work that holds up anchored synthesis is disciplined tagging: every interview excerpt marked with the theme it illustrates, every Hotjar session annotated with the exact moment of the problem, every ticket categorized by type of friction. This organizational work looks bureaucratic and boring, but it's what makes the difference between asking "summarize last quarter's UX themes" and getting a generic summary versus getting a summary where every theme comes with a direct link to session 12, to interview 07, to ticket #4821. The next lesson in this track (N.ux.3) picks up exactly this point: how to turn a hundred loose interviews into a traceable pattern, without losing the thread back to the original source. And later on (N.ux.7), you'll learn to audit the moment when even a well-connected corpus can generate a synthesis that invents a pain point nobody actually mentioned, because connecting the source reduces the risk, but doesn't eliminate the need to check on its own.
Now build
Your exercise in this lesson is to build the Corpus Map of your product: a simple list that becomes your starting point for any synthesis from here on. Take your real task or your own product and answer.
FRONT 1 · WHAT YOU ALREADY HAVE
- List the live research sources that already exist in your product today: interview repository, session recordings, support tickets, app store reviews, satisfaction surveys. Mark which ones are active and which are forgotten.
- For each active source, note: is it connected in some way to the AI you use to synthesize, or do you still copy and paste loose excerpts?
FRONT 2 · THE ANCHOR TEST
- Take the last UX insight someone on your team presented in a meeting (it can be yours). Ask this lesson's question: "points where?". Does it cite a session, an interview, a specific ticket, or is it a loose, plausible-sounding sentence?
- If it's loose, rewrite that same insight in a way that can only be stated if someone actually goes and checks the source (even if you don't have the source on hand right now).
FRONT 3 · THE NEXT CONNECTION
- Pick a warming source from front 1 (one that's stalled, disconnected) and write in one sentence the first practical step to connect it: export it, tag it, give AI access.
In the end, you'll have a three-column map: what already exists, what's already anchored, and what's missing to connect. This map is the literal foundation for the next lesson, where these loose sources become a pattern with traceable citations.
Practice
1. Why does an AI, without a connection to real research sources, tend to generate generic insights about why users abandon a flow?
2. What characterizes a UX insight built through 'anchored synthesis,' according to this lesson?
3. What question does the lesson recommend asking before accepting any AI-generated UX insight?
4. How does the concept of 'context debt' from lesson 1.1 apply to UX research, according to this lesson?
Fair? Let's close the message of this lesson together. You already have the real research corpus, just scattered: interviews in Dovetail, recorded sessions in Hotjar, complaints in support tickets. Connecting that corpus to AI before asking for any synthesis is what separates a defensible insight from a well-written guess. And the test you carry forward is a single question: where does this insight point? If it points to a citation, a session, a specific ticket, it's anchored synthesis and it holds up in a meeting. If it doesn't point anywhere, it's a loose hypothesis, and it needs to be treated as such until you go find the real source. In the next lesson, you'll learn to turn this connected corpus into a traceable pattern, theme by theme, without losing the thread of who said what.
Thanks for the feedback. It helps sharpen the next lesson.