The audit: why AI's number is not trustworthy
AI writes well and calculates poorly, with an air of total certainty. This lesson gives you a five-item audit checklist for finance and the principle that a person always signs off on the number.
An analyst asked AI to put together the month-end close summary. Out came flawless text, with margin, projection, and a closing line that sounded like it was written by a seasoned CFO. She pasted it into the email for the board. She almost sent it. Before clicking, out of habit she redid one calculation by hand: the margin AI computed was 34 percent. In the redone spreadsheet, it was 23. AI had added in revenue that hadn't come in yet. Eleven points of margin invented, ready to become a decision.
At close, AI spat out a beautiful projected cash flow, with every line justified in prose. It nearly became a board presentation. Redoing the math separately, the number didn't add up: AI had flipped the sign of an outflow into an inflow, and the projected cash was 180 thousand above real. The text gave no hint of the error.
For a due diligence, AI calculated a target company's consolidated labor liability and wrote a confident paragraph about the risk. It nearly went into the legal opinion. Checking the source of each number, two cited lawsuits didn't exist in the database and the total was underestimated by almost 40 percent. The text sounded like legal truth, the number was fiction in a robe.
AI put together the quarter's campaign ROI report, with CAC, LTV, and a verdict of "healthy campaign." It nearly went to the budget meeting. Redoing the math, the LTV used an average ticket AI had simply invented, and the real ROI was half of what was reported. The prose was convincing, the number was fiction, and next quarter's budget nearly came out of a table that didn't add up.
AI put together the climate survey summary and spat out an eNPS of 62, with a confident paragraph about "above-average engagement." It nearly became a slide for the people committee. Redoing the math from the raw responses, the real eNPS was 41: AI had counted promoters and neutrals together. Twenty-one points of climate invented, ready to become an action plan on top of a problem that didn't exist the way it was described.
To prioritize the roadmap, AI calculated the RICE score for each initiative and wrote a convincing recommendation of which feature to build first. It nearly entered the PRD. Redoing the calculation separately, one initiative's reach was inflated tenfold: AI confused total users with impacted users. The prioritized feature was the wrong one, and the text gave no hint of it.
At month close, AI consolidated the pipeline forecast and wrote that the team would beat the target with room to spare. It nearly went to the forecast meeting with the sales director. Checking deal by deal in the CRM, AI had applied 100 percent probability to opportunities still in negotiation. The forecast was 300 thousand above real, with the same trustworthy look as the correct ones sitting right next to it.
AI put together the quarter's SLA report and wrote that the operation met 98 percent of deadlines, with a verdict of "healthy performance." It nearly became a presentation for the client. Redoing the math from the logs, AI had left out reopened tickets, and the real SLA was 86 percent. Twelve points of efficiency that didn't exist, ready to support a renewal.
For the annual risk report, AI consolidated the number of data-handling incidents under LGPD and wrote a reassuring paragraph about the exposure level. It nearly went into the committee report. Checking the source of each case, AI had counted two incidents that belonged to another period and left out a relevant one. The total risk was cosmetically improved, with the same confidence as an audited internal control.
After the incident, AI calculated the quarter's uptime from the logs and wrote a post-mortem with the line "availability within the 99.9 percent SLA." It nearly went to the leadership status report. Redoing the math separately, AI had ignored the partial-degradation window and the real uptime was 99.2. The number missed the SLA, and the post-mortem's prose gave no hint of the gap.
AI summarized the usability testing round and spat out a task success rate of 88 percent, with a confident paragraph about the flow "being ready to ship." It nearly became a recommendation for the product team. Reviewing the sessions one by one, AI had counted users who abandoned the flow midway as successes. The real rate was 64 percent, and the pretty number nearly shipped a flow that still tripped up the user.
For the quarterly board deck, AI calculated the company's market share in the segment and wrote a confident paragraph about "consolidated leadership." It almost made the opening slide. Checking the source, AI had added together data from two different market cuts, and the real share was eight points lower. Eight points of invented leadership, ready to become the narrative the board would take to investors the following Monday.
Whoa, stop for a second on that "nearly." In every one of those cases the wrong number nearly got through, not because the person was careless, but because AI delivers wrong with the same confidence it delivers right. That's this lesson's uncomfortable point: an AI's numeric output is not trustworthy by default. And in finance, where a number becomes a decision, a contract, and money, that's not a detail. It's the difference between professionalism and roulette.
The core idea of this lesson. AI is excellent at text and terrible as a source of numeric truth: it invents values, cites a calculation it never made, flips signs, gets the order of magnitude wrong, and delivers all of it with total confidence. That's why AI's number is a proposal, not a truth. Auditing isn't distrusting the tool, it's professional hygiene that protects your signature. This lesson closes with a five-item checklist and a non-negotiable principle: the person who signs off on the number is always a person.
01Great at text, terrible at truth
It's worth understanding why this happens, because understanding changes how you use the tool. The AI you use is, at its core, a machine for predicting the next word that sounds right. It's trained so the text is fluent, plausible, and confident. Nobody trained it so the math adds up. These are two different goals, and it only chases the first one.
The side effect is cruel for finance. The same skill that makes AI write a flawless executive paragraph makes it write a flawless number that's wrong. It has no internal sense of "this can't be right." When a data point is missing, it doesn't stop: it fills in something plausible. This even has a technical name, hallucination, but the name hides the danger. In text, a hallucination is a strange detail you notice. In a number, it's a value with the same look as all the others, and you don't notice.
Notice the four classic failure modes, because your checklist is going to aim exactly at them. It invents a value when data is missing. It cites a calculation it never made, describing a computation that looks like it happened but didn't. It flips signs, turning an outflow into an inflow, an expense into revenue. And it gets the order of magnitude wrong, with an extra or missing zero that slips by unnoticed in the middle of the prose. All four come out with the same absolute confidence. Confidence isn't a sign of correctness; it's just the tool's default tone.
02The felt paradox: you think you're faster
Here's the study that most made me stop and think. In 2025, METR measured experienced professionals working with and without AI on real tasks. The result everyone expected was a speed gain. What showed up was the opposite: with AI, they were about 19 percent slower. And here's the treacherous part: they THOUGHT they were faster. The feeling said "I gained time" while the stopwatch said "I lost time."
Why does this matter in an auditing lesson? Because the feeling of correctness works just like the feeling of speed. When AI hands you a pretty, well-formatted number with a convincing explanation, your brain registers "this is settled" and lowers its guard. The fluency of the delivery generates confidence that wasn't earned by any verification. You feel covered without being covered.
In text, that feeling costs little: at most an odd sentence slips through. In finance, it costs a lot. The same false confidence that made METR's professional think they'd gained time makes you think the number is right. And the wrong number that "seems right" is exactly what gets past your review and reaches the board. The economic frame is direct: the feeling is cheap to produce and expensive to believe.
03Finance's audit checklist
Here's the heart of the lesson. Auditing a number isn't a talent, it's a procedure. Five questions, in order, before any AI number goes out under your name. Think of them as a funnel: each question filters out one type of error, and what passes through all five is a number you can defend.
The first: does the math add up when redone separately? This is the queen. You or another tool redo the calculation independently, without looking at AI's answer, and the two match. If you only checked by reading AI's explanation, you didn't audit anything, you got convinced. Redoing it separately is what catches the invented value and the flipped sign.
The second: does the source of each number exist and match? Every number enters from somewhere, a spreadsheet, a statement, a system. You check whether the source really exists and whether the cited value matches the source. That's what catches the calculation AI said it made but didn't, and the database it cited but invented.
The third: are the assumptions explicit? Every financial number carries assumptions, which ticket, which rate, which period, what's in and what's out. If the assumptions are hidden in the prose, the number is hiding from you what it's made of. Implicit assumptions are where the error hides.
The fourth: does any sign or order of magnitude look strange? This is the gut-check test. You take a step back and look at the number like a person: does this margin make sense? is this total the right size? is this value an inflow or outflow? An extra zero or a flipped sign rarely survives a common-sense look, as long as you stop to take that look.
The fifth, and the most important: can whoever signs off defend it? If you had to explain that number in front of the board, line by line, could you? If the answer is "no, AI did it," the number isn't ready. A number nobody can defend is an orphan number, and an orphan number doesn't ship.
04The RESPOND ruler: AI proposes, the human checks, the person signs
The checklist has a principle behind it, and that principle is what holds everything up. On the Three Movements map, the RESPOND movement deals with exactly this: a decision machine with responsibility anchored to a human name. Applied to finance, it becomes a three-beat ruler you never collapse.
AI proposes. It's fast, tireless, and great for generating the draft number, the first version of the calculation, the sketch of the report. Use it freely here, this is where it shines. But a proposal is a proposal: nothing it produces is true just because it produced it.
The human checks. This is the five-question checklist running. This step isn't optional, isn't "when there's time," isn't "if I have doubts." It's the step that turns a machine proposal into an audited number. Skipping this step is the mistake that lets the wrong number nearly get through, just like in the opening scenarios.
And the person signs. Here's the point of no return: the person who signs off on the number is always a person. When the number goes out under your name, in an email, in a legal opinion, in a presentation, the responsibility is entirely yours. And listen closely, because I repeat this everywhere: "AI calculated it" doesn't exist as an excuse. It doesn't exist on the board, doesn't exist in an audit, doesn't exist in the hard conversation after the wrong number became the wrong decision. The machine has no name, no ID, and it doesn't go to the meeting to defend the math. You do.
05Auditing isn't distrust: it's protecting your signature
Let me clear up a common pushback before wrapping up. Some people feel that auditing AI means not trusting the tool, that it's rework, that it wastes the productivity gain. That frame is wrong, and it's expensive.
Auditing isn't distrust, it's hygiene. You don't check the number because you think AI is dumb; you check it because any number that's going out under your name deserves that care, whether it came from AI, from an intern, or from you yourself at eleven at night. The experienced accountant redoes the math not because they distrust the calculator, but because that's how you work with money. AI is a calculator that sometimes invents the result and never warns you. The audit is the professional hygiene that requires.
And the economic frame closes the argument. The cost of auditing is minutes. The cost of not auditing is a wrong number on the board, a decision made on top of fiction, a reputation that took years to build getting scratched. Minutes against that isn't rework, it's the best cheap insurance there is. AI gives you back time on the proposal; you reinvest a fraction of that time on the check and you keep the net gain AND your signature protected. Auditing is what makes the productivity gain safe to use. Without it, the gain is just risk wearing an efficiency costume.
Do it now
Take a real financial number you would produce or have produced with AI's help (your real task works well). Your task isn't to redo the math right now, it's to build YOUR OWN five-item audit checklist, adapted to your context. For each item, write the question in your own words and the concrete check you'd perform:
- DOES THE MATH ADD UP WHEN REDONE SEPARATELY? How, exactly, would you redo this number independently, without looking at AI's answer? (which tool, which calculation, who does it)
- DOES THE SOURCE EXIST AND MATCH? What are the sources of the numbers that feed into this calculation, and how do you check that each value matches the real source?
- ARE THE ASSUMPTIONS EXPLICIT? List the three most important assumptions embedded in this number (rate, period, what's in and what's out). Are they written down or hidden?
- STRANGE SIGN OR MAGNITUDE? What's the gut-check test: what value, if it showed up, would make you stop cold and say "this can't be right"?
- CAN WHOEVER SIGNS DEFEND IT? Write, in one sentence, the name of the person who signs off on this number and the one-line defense they'd give the board. If you can't write that sentence, the number isn't ready yet.
Keep this checklist. It applies to every number you produce with AI from here on.
Practice
1. AI delivered a month-end close report with the month's margin, in flawless, very convincing executive prose. What is the correct read on this delivery before you use the number?
2. The METR study (2025) showed experienced professionals were about 19% slower with AI, but thought they were faster. What does this teach us about auditing numbers?
3. A number calculated by AI went out under your name in a legal opinion and was wrong. Who answers for that number?
Fair? This lesson's message is direct and economical: AI's number is a cheap, fast proposal, and that's great, as long as you never confuse proposal with truth. The five questions take minutes; the wrong number that gets through costs far more. Auditing isn't slowing AI down out of fear or distrusting the tool, it's the hygiene that protects your signature. AI proposes, you check, and you're the one who signs. Always.
For the board
On the nature of the toolgreat at text, terrible as a source of numerical truth. A proposal is not the truth.
On the feelingthe same feeling that made experienced professionals think they were faster makes you drop your guard without having checked anything.
On responsibilityit does not migrate to the tool under any circumstance. Whoever signs answers for the number.
Thanks for the feedback. It helps sharpen the next lesson.