Business: Operations · Lesson N.ops.7

The audit: when the dashboard lies with conviction

AI writes well and calculates badly, with the look of total certainty. This lesson gives you a five-item audit checklist for operational metrics and the principle that a person always signs off on the number.

Examples for

Monday morning, and the logistics coordinator asks AI to build last week's SLA dashboard from the tracking system. An impeccable summary comes back: "98% SLA, healthy operation, no points of concern". He almost sends the screenshot to the director's group chat. Before sending, out of habit, he opens the TMS and redoes the math by hand, counting deliveries reopened due to complaints. The real SLA was 86%. AI had excluded from the count every ticket reopened after "resolved", because it only counted the first delivery attempt. Twelve points of healthy operation that didn't exist, ready to become executive peace of mind over a real problem.

Let's stop for a second on that "almost". In every one of these cases the wrong metric almost went through, not because the coordinator was careless, but because AI delivers a wrong dashboard with the same confidence it delivers a correct one. That's this lesson's uncomfortable point: an AI's numeric output about an operational metric isn't reliable by default. And in operations, where a metric becomes a client report, a hiring decision, and a contract renewal, that's not a detail. It's the difference between real management and roulette with a pretty spreadsheet.

The core idea of this lesson. AI is excellent at text and terrible as a source of numeric truth about metrics: it invents a value when data is missing, cites a calculation it never did, swaps numerator for denominator, gets the order of magnitude wrong, and delivers all of it with total confidence. That's why the metric AI calculates is a proposal, not a truth. Auditing isn't distrusting the tool, it's professional hygiene that protects your signature. This lesson closes with a five-item checklist and a non-negotiable principle: a person always signs off on the metric.

01Great at text, terrible at truth

It's worth understanding why this happens, because understanding it changes how you use the tool. The AI you use is, at its core, a machine for predicting the next word that sounds right. It's trained so the text is fluent, plausible, and confident. Nobody trained it so the SLA math checks out. Those are two different objectives, and it only chases the first one.

The side effect is cruel for operations. The same skill that makes AI write an impeccable dashboard paragraph makes it write an impeccable metric that's wrong. It has no internal sense of "this can't be right for a distribution center this size". When a piece of data is missing, it doesn't stall: it fills in something plausible. In text, that's a strange detail you notice. In a metric, it's a number that looks exactly like all the others, and you don't notice.

Notice the four classic failure modes, because your checklist is going to aim exactly at them. It invents a value when data is missing from the source system. It cites a calculation it never did, describing an SLA or fill rate computation that looks like it happened but didn't. It swaps numerator for denominator, or counts "fulfilled" when it should count "fulfilled on time". And it gets the order of magnitude wrong, with one percentage point more or less that slips by unnoticed in the middle of the confident prose. All four come out with the same absolute confidence. Confidence isn't a sign of accuracy, it's just the tool's default tone.

02The paradox of feeling: you think the dashboard is right

Here's the study that should make you stop and think the most. In 2025, METR measured experienced professionals working with and without AI on real tasks. The expected result was a speed gain. What showed up was the opposite: with AI, they were about 19% slower. And the treacherous part: they thought they were faster. The feeling said "I saved time" while the stopwatch said "I lost time".

Why does this matter in a lesson about auditing metrics? Because the feeling of being right works exactly like the feeling of being fast. When AI hands you a pretty, well formatted dashboard, with a convincing closing sentence, your brain registers "this is settled" and lowers its guard. The fluency of the delivery generates a confidence that wasn't earned by any verification. You feel covered without being covered.

In loose text, that feeling costs little. In an operations metric, it costs dearly. The same false confidence that made the METR professional think they saved time makes you think the SLA is right. And the wrong metric that "seems right" is exactly what slips past your review and reaches the client, the board, the quality committee. The economic frame is direct: the feeling is free to produce and expensive to believe.

metric seems right metric checked the feeling seems settled the measured 12 points of gap METR 2025: the feeling of being right goes up, accuracy goes down

03The operations metric audit checklist

Here's the heart of the lesson. Auditing a metric isn't a talent, it's a procedure. Five questions, in order, before any metric calculated by AI goes out under your name. Think of them as a funnel: each question filters out one type of error, and what passes all five is a number you can defend.

The first: does the math check out when redone separately? This is the queen. You, or another tool, redo the SLA, fill rate, or cost per delivery calculation independently, without looking at AI's answer, and the two match. If you only checked by reading AI's explanation, you didn't audit anything, you got convinced. Redoing it separately is what catches the reopened delivery counted as success and the swapped denominator.

The second: does each number's source exist and match? Every metric comes from somewhere: the WMS, the TMS, the ticket system. You check that the source really exists and that the cited value matches it. This is what catches the calculation AI said it did but didn't.

The third: are the metric's assumptions explicit? Every operational metric carries assumptions: what counts as "on time", what counts as "reopened", what the cutoff period is, what goes into the base and what doesn't. If the assumptions are hidden inside the prose, the metric is hiding what it's made of from you. A hidden assumption is where the error hides.

The fourth: does any signal or order of magnitude look strange? This is the smell test. You take a step back and look at the number like a person: does a 98% SLA make sense for an operation you know had a rough month? An extra zero or an inverted metric rarely survives a common-sense look, as long as you stop to take that look.

The fifth, and the most important: can whoever signs off defend it? If you had to explain this metric in front of the client in a results meeting, line by line, could you? If the answer is "no, AI calculated it", the metric isn't ready. A number nobody can defend is an orphan number, and an orphan number doesn't go out.

metric proposed by AI 1 · does the math check out redone separately? 2 · does each number's source exist and match? 3 · are the assumptions explicit? 4 · odd signal or magnitude? 5 · can whoever signs off defend it? defensible metric

04The RESPOND ruler: AI proposes, the human checks, the person signs

The checklist has a principle behind it, and that principle is what holds everything up. The RESPOND movement deals with exactly this: a decision machine with responsibility anchored to a human name. Applied to operations, it becomes a three-beat ruler you never collapse.

AI proposes. It's fast, tireless, and great at generating the metric's draft, the first version of the calculation, the sketch of the SLA report. Use it freely here, this is where it shines. But a proposal is a proposal: nothing it produces is true just because it produced it.

The human checks. This is the five-question checklist running. This step isn't optional, it's not "when there's time". It's the step that turns a machine's proposal into an audited metric. Skipping this step is the mistake that lets the wrong metric almost go through, like in the examples at the start.

And the person signs off. Here's the point of no return: a person always signs off on the metric. When the number goes out under your name, in a QBR with the client, in a report to the quality committee, in a results presentation, the responsibility is yours, entirely. "AI calculated it" isn't an excuse. It doesn't exist in the client meeting, it doesn't exist in the audit, it doesn't exist in the hard conversation after the wrong metric became a broken promise. The machine has no name, no badge, and it doesn't go to the meeting to defend the dashboard. You do.

AI proposes the human checks the person signs a proposal isn't truth, truth is what you check and sign off on

05Auditing isn't distrust: it's protecting your signature

Let me clear up a common pushback before wrapping up. Some people feel that auditing AI's metric is not trusting the tool, that it's rework, that it wastes the speed gain. That frame is wrong, and it's expensive.

Auditing isn't distrust, it's hygiene. You don't check the SLA because you think AI is dumb, you check it because any number that's going out under your name deserves that care, whether it comes from AI, an intern, or you yourself at eleven at night. The experienced supervisor redoes the math not out of suspicion of the system, but because that's how you work with a metric that becomes a client report. AI is a calculator that sometimes makes up the result and never warns you. Auditing is the professional hygiene that requires.

And the economic frame closes the argument. The cost of auditing is minutes. The cost of not auditing is a wrong SLA in a QBR, a contract renewal negotiated on top of an inflated fill rate, a reputation for operational excellence that took years to build and collapses in one screenshot. Minutes against that isn't rework, it's the best cheap insurance there is. AI gives you back time on the proposal; you reinvest a fraction of that time on the check and keep the net gain, with your signature protected.

Do it now

Do it yourself

Take a real operational metric you would produce or have produced with AI's help (your real task works well: SLA, fill rate, cost per delivery, cycle time). Your task isn't to redo the math right now, it's to build YOUR OWN five-item audit checklist, adapted to your context. For each item, write the question in your own words and the concrete check you would do:

  1. DOES THE MATH CHECK OUT REDONE SEPARATELY? How, exactly, would you redo this metric independently, without looking at AI's answer? (which system, which calculation, who does it)
  1. DOES THE SOURCE EXIST AND MATCH? What are the sources feeding this metric (WMS, TMS, ticket system, time clock), and how do you check that each value matches the real source?
  1. ARE THE ASSUMPTIONS EXPLICIT? List the three most important assumptions built into this metric (what counts as "on time", what counts as "reopened", which period). Are they written down or hidden?
  1. ODD SIGNAL OR MAGNITUDE? What's the smell test: what value, if it showed up, would make you stop right away and say "this can't be right for my operation"?
  1. CAN WHOEVER SIGNS OFF DEFEND IT? Write, in one sentence, the name of the person who signs off on this metric and the one-line defense they'd give in a meeting with the client. If you can't write that sentence, the metric isn't ready yet.

Keep this checklist. It's good for every metric you're going to produce with AI from here on.

Practice

1. AI delivered an SLA dashboard with an impeccable, very convincing executive summary. What's the correct read before you use that number?

2. The METR study (2025) showed experienced professionals were about 19% slower with AI, but thought they were faster. What does that teach about trusting a metric without checking it?

3. A metric calculated by AI went out under a supervisor's name in a client report and was wrong. Who's accountable for that number?

Fair enough? The lesson's takeaway is direct and economic: AI's metric is a cheap, fast proposal, and that's great, as long as you never confuse proposal with truth. The five questions take minutes; the wrong metric that gets through costs much more. Auditing isn't slowing AI down out of fear or distrusting the tool, it's the hygiene that protects your signature. AI proposes, you check, and you're the one who signs off. Always.

For the board

On the dashboardit lies with conviction. Numerical output is not trustworthy by default, however handsome the text around it.
On the feelingbelieving the number is right is what makes you drop your guard without having checked anything.
On the signaturethe AI proposes, the human checks, the person signs. Responsibility does not migrate to the tool.
What did you think of this page?
Would you recommend this page to someone on your team?