Business: Product · Lesson N.prod.7

The audit: when AI makes up a product metric

AI writes well and calculates poorly, with the look of total certainty. This lesson gives you a five-item audit checklist for every product metric (retention, activation, adoption) and the principle that a person always signs off on the number.

Examples for

A PM asked AI to build the quarter's retention summary. Out came an impeccable text, with a D30 retention rate and a closing line that sounded like it was written by an experienced head of product. He almost pasted it into the board deck. Before sending it, he redid the calculation directly in Amplitude out of habit: the retention AI had calculated was 42%. In the real dashboard, it was 29%. AI had counted as retained a user who only opened the app once during the period. Thirteen points of retention made up, ready to become a roadmap decision.

Man, stop for a second on that "almost". In every one of these cases the wrong number almost went through, not because the person was careless, but because AI delivers the wrong answer with the exact same confidence as the right one. That's this lesson's uncomfortable point: an AI's numerical output isn't trustworthy by default. And in product, where a retention or adoption metric becomes a roadmap decision, a hiring call, and a board presentation, that's not a detail. It's the difference between professionalism and roulette.

The core idea of this lesson. AI is excellent at text and terrible as a source of numerical truth: it counts wrong, applies a different criterion than the one you asked for, rounds up, and delivers all of it with total confidence. That's why the metric AI calculates is a proposal, not a truth. Auditing isn't distrust of the tool, it's professional hygiene that protects your signature. This lesson closes with a five-item checklist and a non-negotiable principle: whoever signs off on the metric is always a person.

01Great at text, terrible at truth

It's worth understanding why this happens, because understanding it changes how you use the tool. The AI you use is, at its core, a machine for predicting the next word that sounds good. It's trained so the text is fluent, plausible, and confident. Nobody trained it to make the math add up. Those are two different goals, and it only chases the first one.

The side effect is cruel for product. The same skill that makes AI write an impeccable analysis paragraph makes it write an impeccable metric that's wrong. It has no internal sense of "this can't be right". When there's no clarity about the criterion (what counts as a "retained user", what counts as "activated"), it doesn't freeze: it picks a plausible criterion on its own, and sometimes that criterion isn't yours.

Notice the classic failure modes. It counts as success what was actually abandonment. It applies a different "activation" criterion than the one the team defined. It rounds up because the text "sounds better" with a round number. It adds up cohorts that shouldn't be added. All of it comes out with the same absolute confidence.

02The feeling paradox: you think you're faster

Here's the study that made me stop and think the most. In 2025, METR measured experienced professionals working with and without AI on real tasks. The expected result was a speed gain. What showed up was the opposite: with AI, they got about 19% slower, and the treacherous part is they THOUGHT they were faster.

Why does this matter in a lesson about metric audits? Because the feeling of being right works just like the feeling of being fast. When AI hands you a pretty, well-formatted metric, with a convincing explanation of why the feature worked, your brain registers "this is settled" and drops its guard. In product, that false confidence costs dearly: the wrong number that "seems right" is exactly what slips past your review and makes it to the all-hands, the board meeting, the decision to double investment in a feature that actually wasn't performing.

faster slower the feeling I gained time what was measured minus 19 percent METR 2025: the feeling went up, the result went down

03The product metric audit checklist

Here's the heart of the lesson. Auditing a metric isn't a talent, it's a procedure. Five questions, in order, before any AI-calculated metric goes out under your name.

The first: does the math hold up when redone separately, straight from the source? You open Amplitude, the dashboard, or the source spreadsheet and recalculate without looking at AI's answer. If the two match, move on. If you only checked by reading AI's explanation, you audited nothing, you were convinced.

The second: is the criterion used the same one the team defined? "Retained user", "activated", "adopter" are definitions your team fixed somewhere. If AI used a different criterion (counted a single session as retention, for example), the number is right for the wrong definition.

The third: are the premises explicit? Which period, which segment, what's in and what's out of the base. A hidden premise in the prose is where the error hides.

The fourth: does any signal or order of magnitude look strange? Take a step back: does this 42% retention make sense for your product, given your history? A jump that's too big rarely survives a common-sense look.

The fifth, and the most important: does whoever signs off know how to defend that number at the all-hands, line by line? If the answer is "no, AI calculated it", the metric isn't ready.

metric proposed by AI 1 · does the math hold up redone at the source? 2 · is the criterion used the team's? 3 · are the premises explicit? 4 · odd signal or magnitude? 5 · can whoever signs off defend it? defensible metric
Learn more: why the metric's criterion changes everything, even with correct math

Two teams can calculate "D30 retention" with mathematically correct formulas and land on very different numbers, just because the definition of "retained" changes: opened the app once, or completed a value-generating action. AI doesn't know which definition YOUR team uses, unless you say so explicitly. The most common mistake isn't wrong math, it's correct math applied to the wrong criterion. That's why checklist item two (the criterion is the same one the team defined) is just as important as item one (the math holds up).

04The RESPOND ruler: AI proposes, the human checks, the person signs off

The checklist has a principle behind it. AI proposes. It's fast and great for generating the metric's draft, the first version of the calculation. Use it freely here. But a proposal is a proposal: nothing it produces is true just because it produced it.

The human checks. That's the five-question checklist running. It's not optional, it's not "if there's time". It's the step that turns a machine's proposal into an audited number.

And the person signs off. When the metric goes out under your name, at an all-hands, in a board deck, the responsibility is yours, entirely. "AI calculated it" doesn't exist as an excuse. The machine won't go to the meeting to defend the number. You will.

05Auditing isn't distrust: it's protecting your signature

Auditing isn't distrust of the tool, it's hygiene. The cost of auditing is minutes. The cost of not auditing is a wrong metric at the all-hands, a roadmap decision made on top of fiction, credibility scratched that took years to build. AI gives you back time on the proposal; you reinvest a fraction of that time on the check and keep the net gain AND your protected signature.

Do it now

Do it yourself

Take a real product metric you would produce or did produce with AI's help (your real task works well). Build YOUR five-item audit checklist, adapted to your context:

  1. DOES THE MATH HOLD UP REDONE AT THE SOURCE? How, exactly, would you recalculate this metric independently, without looking at AI's answer?
  2. IS THE CRITERION THE SAME ONE THE TEAM DEFINED? Write the official definition of "retained", "activated", or "adopter" your team uses, and check whether AI used that same definition.
  3. ARE THE PREMISES EXPLICIT? List the period, the segment, and what's in or out of the base.
  4. ODD SIGNAL OR MAGNITUDE? What value, if it showed up, would make you stop right away and say "that can't be right"?
  5. CAN WHOEVER SIGNS OFF DEFEND IT? Write the name of the person who signs off on this metric and the one-line defense they'd give at the all-hands.

Keep this checklist. It applies to every metric you produce with AI from here on.

Practice

1. AI delivered a quarterly retention summary with a very convincing executive text. What's the correct read before you use the number?

2. Two teams calculate 'D30 retention' with mathematically correct formulas and land on very different numbers. What's the most likely explanation, according to this lesson's checklist?

3. An AI-calculated activation number went out under your name in a presentation and was wrong. Who answers for that number?

For the board

On its naturegreat at text, terrible as a source of numerical truth. A proposal is not the truth.
On the definitiontwo correct formulas giving different numbers is a criterion problem, not an arithmetic one. Check the definition the team agreed on.
On the signatureresponsibility does not migrate to the tool. Whoever signs answers.
What did you think of this page?
Would you recommend this page to someone on your team?