The audit: when the dashboard lies with conviction
AI writes well and calculates badly, with the look of total certainty. This lesson gives you a five-item audit checklist for operational metrics and the principle that a person always signs off on the number.
Monday morning, and the logistics coordinator asks AI to build last week's SLA dashboard from the tracking system. An impeccable summary comes back: "98% SLA, healthy operation, no points of concern". He almost sends the screenshot to the director's group chat. Before sending, out of habit, he opens the TMS and redoes the math by hand, counting deliveries reopened due to complaints. The real SLA was 86%. AI had excluded from the count every ticket reopened after "resolved", because it only counted the first delivery attempt. Twelve points of healthy operation that didn't exist, ready to become executive peace of mind over a real problem.
The controller asked AI to calculate the month's cost per delivery from the logistics expense report, and it returned a clean value, with a confident line about "cost under control". Redoing the math separately, AI had divided the total cost by the number of orders shipped, not by the number of deliveries actually completed, and the real cost per delivery was 30% higher. The text sounded like mature financial analysis; the number was fiction with a spreadsheet under it.
For a transportation compliance audit, AI calculated the percentage of deliveries with all fiscal documentation correct and wrote a calm paragraph about "low risk". It almost went into the opinion without a check. Redoing the count case by case, two batches cited as compliant had an invoice with an unresolved discrepancy, and the real compliance percentage was much lower. The text sounded like a legal guarantee; the number was a bet dressed up as certainty.
AI built the quarter's delivery experience report, with post purchase NPS and a verdict of "logistics that delights the customer". It almost went straight into the brand presentation. Redoing the math from the raw survey responses, AI had counted non-respondents as favorable neutral, and the real NPS was much lower. The prose was convincing, the number was inflated, and next quarter's campaign nearly leaned on a customer perception that didn't exist.
AI summarized the operations team's absenteeism metric and spit out a 4% rate, with a confident paragraph about "stable headcount". It almost became a slide for the people committee. Redoing the math from the time clock system, also counting unjustified absences AI had left out, the real rate was 11%. Seven points of invented stability, ready to become a decision not to reinforce a shift that was already at its limit.
To prioritize improvements to the tracking app, AI calculated the rate of real time tracked deliveries and wrote a convincing recommendation on where to invest first. It almost went straight into the spec. Redoing the calculation separately, AI had counted as "tracked" any delivery with any status update at all, even without geolocation, and the real live tracking coverage was a third of what was reported. The chosen priority was the wrong one, and the confident prose gave no hint of the gap.
At month close, AI consolidated the punctuality metric that backs the contractual discount with the biggest client and wrote that the operation was entitled to the performance bonus. It almost went straight into the billing email. Checking delivery by delivery, AI had applied the wrong tolerance window, more generous than the contract's, and the real metric fell below the bonus threshold. The email almost went out charging for an amount the operation wasn't entitled to.
Friday afternoon, weekly close for the distribution center. AI pulls the WMS report and returns the week's fill rate: "94%, within target". The operations supervisor was about to paste that straight into the report going to the client on Monday. Redoing the calculation from the raw orders, including items substituted due to stock shortage that AI had counted as "fulfilled", the real fill rate was 79%. Fifteen points of difference between what was going to the client and what actually happened on the floor, and the only thing that stopped the error from going out was the supervisor's habit of checking before signing off.
For the quarterly risk report, AI consolidated the number of cargo transport security incidents and wrote a reassuring paragraph about the exposure level. It almost went into the committee report just like that. Checking the source of each case in the incident system, AI had counted two incidents that belonged to another unit and left out a relevant one classified under a different code. The total risk was made up, with the same confidence as a genuinely audited internal control.
After an order spike, AI calculated the quarter's routing system uptime from the logs and wrote a report with the line "availability within the 99.9% SLA". It almost went straight into the leadership status report. Redoing the math separately, AI had ignored the partial degradation window where the system responded slowly but didn't go down, and the real uptime of acceptable performance was 99.1%. The number didn't meet the SLA contracted with the technology provider.
AI summarized the driver satisfaction survey for the route app and spit out an 88% adoption rate, with a confident paragraph about the app "being ready to scale". It almost became a direct recommendation to the operations team. Reviewing the real usage data, AI had counted any driver who opened the app once as "adopted", even without using the suggested route. Real continuous-use adoption was 51%, and the pretty number nearly greenlit an expansion on top of an app most drivers still didn't really trust.
For the quarter's board deck, AI calculated the share of in-house versus outsourced operations in deliveries and wrote a confident paragraph about "insourcing advancing on plan". It almost went into the opening slide without review. Checking the source, AI had added up routes from two different distribution centers with different counting criteria, and the real progress was much smaller. The insourcing narrative the board would take to the investor rested on a number nobody had checked.
Let's stop for a second on that "almost". In every one of these cases the wrong metric almost went through, not because the coordinator was careless, but because AI delivers a wrong dashboard with the same confidence it delivers a correct one. That's this lesson's uncomfortable point: an AI's numeric output about an operational metric isn't reliable by default. And in operations, where a metric becomes a client report, a hiring decision, and a contract renewal, that's not a detail. It's the difference between real management and roulette with a pretty spreadsheet.
The core idea of this lesson. AI is excellent at text and terrible as a source of numeric truth about metrics: it invents a value when data is missing, cites a calculation it never did, swaps numerator for denominator, gets the order of magnitude wrong, and delivers all of it with total confidence. That's why the metric AI calculates is a proposal, not a truth. Auditing isn't distrusting the tool, it's professional hygiene that protects your signature. This lesson closes with a five-item checklist and a non-negotiable principle: a person always signs off on the metric.
01Great at text, terrible at truth
It's worth understanding why this happens, because understanding it changes how you use the tool. The AI you use is, at its core, a machine for predicting the next word that sounds right. It's trained so the text is fluent, plausible, and confident. Nobody trained it so the SLA math checks out. Those are two different objectives, and it only chases the first one.
The side effect is cruel for operations. The same skill that makes AI write an impeccable dashboard paragraph makes it write an impeccable metric that's wrong. It has no internal sense of "this can't be right for a distribution center this size". When a piece of data is missing, it doesn't stall: it fills in something plausible. In text, that's a strange detail you notice. In a metric, it's a number that looks exactly like all the others, and you don't notice.
Notice the four classic failure modes, because your checklist is going to aim exactly at them. It invents a value when data is missing from the source system. It cites a calculation it never did, describing an SLA or fill rate computation that looks like it happened but didn't. It swaps numerator for denominator, or counts "fulfilled" when it should count "fulfilled on time". And it gets the order of magnitude wrong, with one percentage point more or less that slips by unnoticed in the middle of the confident prose. All four come out with the same absolute confidence. Confidence isn't a sign of accuracy, it's just the tool's default tone.
02The paradox of feeling: you think the dashboard is right
Here's the study that should make you stop and think the most. In 2025, METR measured experienced professionals working with and without AI on real tasks. The expected result was a speed gain. What showed up was the opposite: with AI, they were about 19% slower. And the treacherous part: they thought they were faster. The feeling said "I saved time" while the stopwatch said "I lost time".
Why does this matter in a lesson about auditing metrics? Because the feeling of being right works exactly like the feeling of being fast. When AI hands you a pretty, well formatted dashboard, with a convincing closing sentence, your brain registers "this is settled" and lowers its guard. The fluency of the delivery generates a confidence that wasn't earned by any verification. You feel covered without being covered.
In loose text, that feeling costs little. In an operations metric, it costs dearly. The same false confidence that made the METR professional think they saved time makes you think the SLA is right. And the wrong metric that "seems right" is exactly what slips past your review and reaches the client, the board, the quality committee. The economic frame is direct: the feeling is free to produce and expensive to believe.
03The operations metric audit checklist
Here's the heart of the lesson. Auditing a metric isn't a talent, it's a procedure. Five questions, in order, before any metric calculated by AI goes out under your name. Think of them as a funnel: each question filters out one type of error, and what passes all five is a number you can defend.
The first: does the math check out when redone separately? This is the queen. You, or another tool, redo the SLA, fill rate, or cost per delivery calculation independently, without looking at AI's answer, and the two match. If you only checked by reading AI's explanation, you didn't audit anything, you got convinced. Redoing it separately is what catches the reopened delivery counted as success and the swapped denominator.
The second: does each number's source exist and match? Every metric comes from somewhere: the WMS, the TMS, the ticket system. You check that the source really exists and that the cited value matches it. This is what catches the calculation AI said it did but didn't.
The third: are the metric's assumptions explicit? Every operational metric carries assumptions: what counts as "on time", what counts as "reopened", what the cutoff period is, what goes into the base and what doesn't. If the assumptions are hidden inside the prose, the metric is hiding what it's made of from you. A hidden assumption is where the error hides.
The fourth: does any signal or order of magnitude look strange? This is the smell test. You take a step back and look at the number like a person: does a 98% SLA make sense for an operation you know had a rough month? An extra zero or an inverted metric rarely survives a common-sense look, as long as you stop to take that look.
The fifth, and the most important: can whoever signs off defend it? If you had to explain this metric in front of the client in a results meeting, line by line, could you? If the answer is "no, AI calculated it", the metric isn't ready. A number nobody can defend is an orphan number, and an orphan number doesn't go out.
04The RESPOND ruler: AI proposes, the human checks, the person signs
The checklist has a principle behind it, and that principle is what holds everything up. The RESPOND movement deals with exactly this: a decision machine with responsibility anchored to a human name. Applied to operations, it becomes a three-beat ruler you never collapse.
AI proposes. It's fast, tireless, and great at generating the metric's draft, the first version of the calculation, the sketch of the SLA report. Use it freely here, this is where it shines. But a proposal is a proposal: nothing it produces is true just because it produced it.
The human checks. This is the five-question checklist running. This step isn't optional, it's not "when there's time". It's the step that turns a machine's proposal into an audited metric. Skipping this step is the mistake that lets the wrong metric almost go through, like in the examples at the start.
And the person signs off. Here's the point of no return: a person always signs off on the metric. When the number goes out under your name, in a QBR with the client, in a report to the quality committee, in a results presentation, the responsibility is yours, entirely. "AI calculated it" isn't an excuse. It doesn't exist in the client meeting, it doesn't exist in the audit, it doesn't exist in the hard conversation after the wrong metric became a broken promise. The machine has no name, no badge, and it doesn't go to the meeting to defend the dashboard. You do.
05Auditing isn't distrust: it's protecting your signature
Let me clear up a common pushback before wrapping up. Some people feel that auditing AI's metric is not trusting the tool, that it's rework, that it wastes the speed gain. That frame is wrong, and it's expensive.
Auditing isn't distrust, it's hygiene. You don't check the SLA because you think AI is dumb, you check it because any number that's going out under your name deserves that care, whether it comes from AI, an intern, or you yourself at eleven at night. The experienced supervisor redoes the math not out of suspicion of the system, but because that's how you work with a metric that becomes a client report. AI is a calculator that sometimes makes up the result and never warns you. Auditing is the professional hygiene that requires.
And the economic frame closes the argument. The cost of auditing is minutes. The cost of not auditing is a wrong SLA in a QBR, a contract renewal negotiated on top of an inflated fill rate, a reputation for operational excellence that took years to build and collapses in one screenshot. Minutes against that isn't rework, it's the best cheap insurance there is. AI gives you back time on the proposal; you reinvest a fraction of that time on the check and keep the net gain, with your signature protected.
Do it now
Take a real operational metric you would produce or have produced with AI's help (your real task works well: SLA, fill rate, cost per delivery, cycle time). Your task isn't to redo the math right now, it's to build YOUR OWN five-item audit checklist, adapted to your context. For each item, write the question in your own words and the concrete check you would do:
- DOES THE MATH CHECK OUT REDONE SEPARATELY? How, exactly, would you redo this metric independently, without looking at AI's answer? (which system, which calculation, who does it)
- DOES THE SOURCE EXIST AND MATCH? What are the sources feeding this metric (WMS, TMS, ticket system, time clock), and how do you check that each value matches the real source?
- ARE THE ASSUMPTIONS EXPLICIT? List the three most important assumptions built into this metric (what counts as "on time", what counts as "reopened", which period). Are they written down or hidden?
- ODD SIGNAL OR MAGNITUDE? What's the smell test: what value, if it showed up, would make you stop right away and say "this can't be right for my operation"?
- CAN WHOEVER SIGNS OFF DEFEND IT? Write, in one sentence, the name of the person who signs off on this metric and the one-line defense they'd give in a meeting with the client. If you can't write that sentence, the metric isn't ready yet.
Keep this checklist. It's good for every metric you're going to produce with AI from here on.
Practice
1. AI delivered an SLA dashboard with an impeccable, very convincing executive summary. What's the correct read before you use that number?
2. The METR study (2025) showed experienced professionals were about 19% slower with AI, but thought they were faster. What does that teach about trusting a metric without checking it?
3. A metric calculated by AI went out under a supervisor's name in a client report and was wrong. Who's accountable for that number?
Fair enough? The lesson's takeaway is direct and economic: AI's metric is a cheap, fast proposal, and that's great, as long as you never confuse proposal with truth. The five questions take minutes; the wrong metric that gets through costs much more. Auditing isn't slowing AI down out of fear or distrusting the tool, it's the hygiene that protects your signature. AI proposes, you check, and you're the one who signs off. Always.
For the board
On the dashboardit lies with conviction. Numerical output is not trustworthy by default, however handsome the text around it.
On the feelingbelieving the number is right is what makes you drop your guard without having checked anything.
On the signaturethe AI proposes, the human checks, the person signs. Responsibility does not migrate to the tool.
Thanks for the feedback. It helps sharpen the next lesson.