The audit: when AI makes up a product metric
AI writes well and calculates poorly, with the look of total certainty. This lesson gives you a five-item audit checklist for every product metric (retention, activation, adoption) and the principle that a person always signs off on the number.
A PM asked AI to build the quarter's retention summary. Out came an impeccable text, with a D30 retention rate and a closing line that sounded like it was written by an experienced head of product. He almost pasted it into the board deck. Before sending it, he redid the calculation directly in Amplitude out of habit: the retention AI had calculated was 42%. In the real dashboard, it was 29%. AI had counted as retained a user who only opened the app once during the period. Thirteen points of retention made up, ready to become a roadmap decision.
At close, AI spat out a beautiful projected cash flow. Redoing the math separately, the number didn't add up: AI had flipped an outflow's sign into an inflow.
For a due diligence, AI calculated the consolidated labor liability and wrote a confident paragraph about the risk. Redoing the math case by case, two of the cited cases didn't exist in the database.
AI built the campaign's ROI report with a verdict of "healthy campaign". Redoing the math separately, the LTV used a made-up average ticket, and the real ROI was half of what was reported.
AI built the climate survey summary and spat out an eNPS of 62, with a confident paragraph about "above-average engagement". Redoing the math from the raw responses, the real eNPS was 41.
A PM asked AI to build the quarter's activation and retention summary to present at the all-hands. Out came an impeccable text, with 42% D30 retention and a closing line that sounded like it was written by an experienced head of product: "the new feature significantly raised the team's retention". He almost walked on stage with that number. Before presenting, he redid the calculation directly in Amplitude out of habit: real retention was 29%, thirteen points lower. AI had counted as a retained user anyone who opened the app a single time during the period, even without ever coming back. The text sounded like senior product analysis; the number was fiction with a dashboard.
At month close, AI consolidated the pipeline forecast and wrote that the team would beat the target with room to spare. Checking deal by deal, AI had applied a hundred percent probability to opportunities still under negotiation.
AI built the quarter's SLA report and wrote that the operation met 98% of deadlines. Redoing the math from the raw logs, AI had left out reopened tickets, and the real SLA was 86%.
For the annual risk report, AI consolidated the number of data-handling incidents and wrote a reassuring paragraph. Checking each case's source, AI had counted two incidents from a different period.
After the incident, AI calculated the quarter's uptime and wrote "availability within the 99.9% SLA". Redoing the math separately, AI had ignored the partial-degradation window.
AI summarized the round of usability tests and spat out an 88% task-success rate. Reviewing the recorded sessions, AI had counted as success users who abandoned the flow midway.
For the quarterly board deck, AI calculated the company's market share and wrote a confident paragraph about "consolidated leadership". Checking the source, it had added up data from two different market segments.
Man, stop for a second on that "almost". In every one of these cases the wrong number almost went through, not because the person was careless, but because AI delivers the wrong answer with the exact same confidence as the right one. That's this lesson's uncomfortable point: an AI's numerical output isn't trustworthy by default. And in product, where a retention or adoption metric becomes a roadmap decision, a hiring call, and a board presentation, that's not a detail. It's the difference between professionalism and roulette.
The core idea of this lesson. AI is excellent at text and terrible as a source of numerical truth: it counts wrong, applies a different criterion than the one you asked for, rounds up, and delivers all of it with total confidence. That's why the metric AI calculates is a proposal, not a truth. Auditing isn't distrust of the tool, it's professional hygiene that protects your signature. This lesson closes with a five-item checklist and a non-negotiable principle: whoever signs off on the metric is always a person.
01Great at text, terrible at truth
It's worth understanding why this happens, because understanding it changes how you use the tool. The AI you use is, at its core, a machine for predicting the next word that sounds good. It's trained so the text is fluent, plausible, and confident. Nobody trained it to make the math add up. Those are two different goals, and it only chases the first one.
The side effect is cruel for product. The same skill that makes AI write an impeccable analysis paragraph makes it write an impeccable metric that's wrong. It has no internal sense of "this can't be right". When there's no clarity about the criterion (what counts as a "retained user", what counts as "activated"), it doesn't freeze: it picks a plausible criterion on its own, and sometimes that criterion isn't yours.
Notice the classic failure modes. It counts as success what was actually abandonment. It applies a different "activation" criterion than the one the team defined. It rounds up because the text "sounds better" with a round number. It adds up cohorts that shouldn't be added. All of it comes out with the same absolute confidence.
02The feeling paradox: you think you're faster
Here's the study that made me stop and think the most. In 2025, METR measured experienced professionals working with and without AI on real tasks. The expected result was a speed gain. What showed up was the opposite: with AI, they got about 19% slower, and the treacherous part is they THOUGHT they were faster.
Why does this matter in a lesson about metric audits? Because the feeling of being right works just like the feeling of being fast. When AI hands you a pretty, well-formatted metric, with a convincing explanation of why the feature worked, your brain registers "this is settled" and drops its guard. In product, that false confidence costs dearly: the wrong number that "seems right" is exactly what slips past your review and makes it to the all-hands, the board meeting, the decision to double investment in a feature that actually wasn't performing.
03The product metric audit checklist
Here's the heart of the lesson. Auditing a metric isn't a talent, it's a procedure. Five questions, in order, before any AI-calculated metric goes out under your name.
The first: does the math hold up when redone separately, straight from the source? You open Amplitude, the dashboard, or the source spreadsheet and recalculate without looking at AI's answer. If the two match, move on. If you only checked by reading AI's explanation, you audited nothing, you were convinced.
The second: is the criterion used the same one the team defined? "Retained user", "activated", "adopter" are definitions your team fixed somewhere. If AI used a different criterion (counted a single session as retention, for example), the number is right for the wrong definition.
The third: are the premises explicit? Which period, which segment, what's in and what's out of the base. A hidden premise in the prose is where the error hides.
The fourth: does any signal or order of magnitude look strange? Take a step back: does this 42% retention make sense for your product, given your history? A jump that's too big rarely survives a common-sense look.
The fifth, and the most important: does whoever signs off know how to defend that number at the all-hands, line by line? If the answer is "no, AI calculated it", the metric isn't ready.
Learn more: why the metric's criterion changes everything, even with correct math
Two teams can calculate "D30 retention" with mathematically correct formulas and land on very different numbers, just because the definition of "retained" changes: opened the app once, or completed a value-generating action. AI doesn't know which definition YOUR team uses, unless you say so explicitly. The most common mistake isn't wrong math, it's correct math applied to the wrong criterion. That's why checklist item two (the criterion is the same one the team defined) is just as important as item one (the math holds up).
04The RESPOND ruler: AI proposes, the human checks, the person signs off
The checklist has a principle behind it. AI proposes. It's fast and great for generating the metric's draft, the first version of the calculation. Use it freely here. But a proposal is a proposal: nothing it produces is true just because it produced it.
The human checks. That's the five-question checklist running. It's not optional, it's not "if there's time". It's the step that turns a machine's proposal into an audited number.
And the person signs off. When the metric goes out under your name, at an all-hands, in a board deck, the responsibility is yours, entirely. "AI calculated it" doesn't exist as an excuse. The machine won't go to the meeting to defend the number. You will.
05Auditing isn't distrust: it's protecting your signature
Auditing isn't distrust of the tool, it's hygiene. The cost of auditing is minutes. The cost of not auditing is a wrong metric at the all-hands, a roadmap decision made on top of fiction, credibility scratched that took years to build. AI gives you back time on the proposal; you reinvest a fraction of that time on the check and keep the net gain AND your protected signature.
Do it now
Take a real product metric you would produce or did produce with AI's help (your real task works well). Build YOUR five-item audit checklist, adapted to your context:
- DOES THE MATH HOLD UP REDONE AT THE SOURCE? How, exactly, would you recalculate this metric independently, without looking at AI's answer?
- IS THE CRITERION THE SAME ONE THE TEAM DEFINED? Write the official definition of "retained", "activated", or "adopter" your team uses, and check whether AI used that same definition.
- ARE THE PREMISES EXPLICIT? List the period, the segment, and what's in or out of the base.
- ODD SIGNAL OR MAGNITUDE? What value, if it showed up, would make you stop right away and say "that can't be right"?
- CAN WHOEVER SIGNS OFF DEFEND IT? Write the name of the person who signs off on this metric and the one-line defense they'd give at the all-hands.
Keep this checklist. It applies to every metric you produce with AI from here on.
Practice
1. AI delivered a quarterly retention summary with a very convincing executive text. What's the correct read before you use the number?
2. Two teams calculate 'D30 retention' with mathematically correct formulas and land on very different numbers. What's the most likely explanation, according to this lesson's checklist?
3. An AI-calculated activation number went out under your name in a presentation and was wrong. Who answers for that number?
For the board
On its naturegreat at text, terrible as a source of numerical truth. A proposal is not the truth.
On the definitiontwo correct formulas giving different numbers is a criterion problem, not an arithmetic one. Check the definition the team agreed on.
On the signatureresponsibility does not migrate to the tool. Whoever signs answers.
Thanks for the feedback. It helps sharpen the next lesson.