Choreography: catching the anomaly before it blows up
AI sweeps through a volume of entries your eye could never cover and lights up the outliers as suspects. You are still the one who investigates, separates error from legitimate exception, and decides the action.
Wednesday morning, month-end expense close. Among eleven thousand entries, one is a travel reimbursement of four thousand two hundred reais, posted to a cost center with nobody traveling. Could be fraud. Could be an entry thrown into the wrong account in a hurry. Could be a perfectly legitimate exception you simply don't remember. The uncomfortable point is: you would never find that line by eye, because nobody reads eleven thousand lines. This lesson's question is: who sweeps that volume for you, and who decides what to do with what turns up?
Performance review cycle closed, eight hundred ratings to review. AI sweeps the whole batch and lights up two outliers: one manager gave the entire team top marks, another gave the worst score of the cycle to someone who always delivered well. Could be favoritism, could be a manager who doesn't calibrate, could be a real performance drop you haven't tracked closely. You would never find those two cases reading eight hundred reviews line by line; AI just gave you the short list to investigate. Before you flag either one as a problem in the HR system, you check the context, because treating the beep as a verdict is the fastest way to punish a manager who did nothing wrong.
End of quarter, time to look at product metrics on the dashboard. Among hundreds of events and funnels, AI lights up a feature whose activation dropped by half in a week, with no announced change. Could be a bug nobody reported, could be seasonality, could be another team's experiment touching the flow without telling you. You would never find that drop digging through every funnel by hand; AI just handed you where to dig, not the reason. Before you escalate this as an incident in the quarter's PRD, you check the cohort closely, because a real drop becomes a priority and a measurement error becomes just noise.
Friday, pipeline review before the forecast, six hundred opportunities in the CRM. AI lights up one worth two hundred thousand reais marked "likely to close," but with no logged interaction in forty days. Could be a dead deal the rep doesn't want to admit to, could be a negotiation running outside the CRM, could be a big account that's really going to close and is just poorly updated. You would never find that line reviewing six hundred opportunities one by one; AI just gave you the short list of where to look. Before you pull that value out of the forecast that goes to the sales director, you call the rep and confirm, because cutting a real deal from the number because of a stale CRM entry is expensive.
Start of the week, reading the SLA dashboard, thousands of processed orders. AI lights up an entire batch from one carrier that went out with a delivery time three times longer than the route average. Could be a new bottleneck at the distribution center, could be a registration error in the system, could be a legitimate exception, a region under construction that always runs late during that period. You would never find that batch reading thousands of orders line by line; AI just lit up where to investigate. Before you charge the carrier or change the process in Zendesk, you check the real reason, because acting on a legitimate exception as if it were a failure breaks a relationship that was working.
Quarterly internal audit, review of access and data handling under data-privacy law. AI lights up a log record of an employee exporting a database with personal data on a Sunday at dawn. Could be a leak, could be a scheduled technical routine nobody documented, could be legitimate access within a project you don't know about. You would never find that record reading thousands of logs line by line; AI just handed you the suspect. Before you take this to the risk committee as an incident, you investigate with the technical team, because opening a leak case over a normal scheduled routine costs credibility that compliance doesn't easily get back.
Monday morning, reviewing the weekend's deploys, hundreds of commits and production metrics. AI lights up a service that started consuming triple the memory after a Friday-night deploy. Could be a memory leak about to take everything down, could be a seasonal traffic spike unrelated to the code, could be an expected change the team didn't communicate. You would never find that service reading every dashboard of every service every morning; AI just pointed to where to look first. Before you open an incident or roll back, you check the log closely, because reverting a correct deploy because of a seasonal spike costs a night of work for nothing.
Journey analysis round after a launch, thousands of recorded sessions. AI lights up a signup screen in Figma with an abandonment rate far above every other screen in the flow. Could be a confusing field that trips up the user, could be a validation bug that only shows up on certain devices, could be expected behavior, people just taking a look with no intention of continuing. You would never find that screen watching thousands of sessions one by one; AI just gave you where to investigate. Before you redesign the field or open a ticket for the technical team, you review a few sessions closely, because tinkering with a field that isn't the real problem only delays fixing what actually trips up the user.
Quarterly review of the competitive dashboard, with two hundred market mentions monitored this week. Among them, AI lights up a small competitor launching a product nearly identical to the one you're about to announce to the board on Friday. Could be a real leak of your strategy, could be positioning coincidence, could be an independent move that only looks the same from a distance. You would never find that mention reading two hundred updates by hand; what AI did was light up the suspect, not conclude anything. Before you take this to the board as a sign of a leak, you investigate the source, and only after confirming does it become a real alert, not a false positive that would have caused internal panic for nothing.
Let me tell you about a job that has always been kind of impossible in controllership. You are responsible for making sure the numbers close and that nothing strange slips through. Except "nothing strange" lives inside thousands of lines nobody has time to read. So, in practice, you sample: you look at what's big, what stands out, what someone complained about. The rest passes in the dark. Think with me: what if you could sweep the rest of it too, and still only need to look at the short list of what seems off? That is exactly this lesson's choreography.
The core idea of this lesson. AI is not an auditor that concludes in your place. It is a metal detector: it passes over a giant field you would never cover by hand, and it beeps where something is buried. The beep is not the treasure or the trash, it is just the "dig here" signal. Who digs, who decides whether it's error, fraud, or legitimate exception, and who takes the action, is still you. AI gives you reach; the judgment is still yours.
01What AI does well: sweeping what the eye can't
Let's call it what it is. Controllership's bottleneck was never intelligence, it was volume. You know exactly what a suspicious entry looks like, but there is no way to scan every line, every month, in every account. So sampling is what's left, and sampling leaves gaps.
AI flips that. It doesn't get tired, doesn't skip a line, doesn't go "this looks fine." You point it at the whole database and ask it to light up what is out of pattern:
- Outlier expenses. A value far above the average for that category, that vendor, that month.
- Odd entries. A cost center that doesn't match, an account that doesn't usually receive that type of expense, a strange date (an entry on a holiday, on the weekend, at the last minute of the close).
- Duplicates. The same invoice, the same amount, the same vendor hitting twice a day or so apart.
- Pattern breaks. The vendor who always charges a round amount and this month came in broken, the recurring expense that disappeared, the one that showed up out of nowhere.
The real gain here is reach. You go from "I looked at the ten biggest" to "I went through all of them, and I have a list of forty to investigate." Fair?
How the sweep works, in practice:
02What stays yours: investigating, separating, and deciding
Here is the part AI doesn't do, and it's exactly the heart of controllership. AI lit up forty suspects. Now what? Now your work begins, and it's all judgment.
- Investigate the suspect. AI pointed, it didn't explain. That reimbursement in the wrong cost center: was it a misposted allocation, was someone taking advantage, was it an agreement you forgot? That means pulling the document, asking the person, looking at the history. AI doesn't do that for you.
- Separate error from legitimate exception. This is the fine point. Not everything out of pattern is wrong. The year-end bonus always breaks the curve, and that's correct. The new vendor always looks strange, and it's just new. Telling "this is wrong" apart from "this is right and just unusual" is reading context, and context is yours.
- Decide the action. Found an error? Reverse it, fix it, escalate it, have a conversation. AI doesn't confront anyone, doesn't block a payment, doesn't make a decision with consequences. Who signs off on the action, and answers for it, is a person. Always.
Notice the design: AI narrows your search field, it doesn't replace your trained eye. It turns "search everything" into "look at what matters." The thinking work stays entirely with you, just applied in the right place.
03The cost to name: the false positive
Now the part I can't let you forget, because it's where this choreography goes wrong if done poorly. AI lights up suspects, and not every suspect is guilty. This has a name: false positive. AI flags something as an anomaly and, when you investigate, it was nothing.
And here lies a human risk, not a technical one. If you treat AI's list as a verdict, you'll confront the wrong people, reverse a correct entry, create internal friction over a beep that was just the detector reacting to an old coin in someone's pocket. The cost of the false positive isn't the machine's, it's yours, in the form of wasted time and credibility spent accusing the innocent.
That's why this choreography's rule, which echoes the track's golden rule, is: the human filters before any action. AI lights it up, you investigate, and only after confirming does it become a confrontation, a reversal, or an escalation. No beep turns into an accusation without you digging first.
04The choreography, step by step
Put it all together and it becomes a simple system, one you run every month-end. Four steps, roles clearly separated.
- AI sweeps. You feed it the period's entry database and ask it to light up what is out of pattern, by anomaly category: outlier value, inconsistent cost center, duplicate, atypical date. Output: the short list of suspects, with the reason for each.
- You investigate. Take the list and dig into each line: document, context, history. Here you're no longer sweeping eleven thousand, you're looking at forty with real attention.
- You classify. Each suspect becomes one of three things: error (needs correcting), legitimate exception (it's fine, just unusual), or fraud/risk (needs escalating). The false positives die here.
- You decide and act. Only what's left confirmed becomes action: reversal, correction, process, conversation. And you close the loop by logging what was recurring, to tune next month's sweep.
Notice that AI shows up in one step only, at the beginning. It doesn't stand in for the controller, it's at the entry point, extending their reach. It's the metal detector before the digging, not the judge after it. Shall we?
05Where this fits in the rest of the course
This choreography doesn't live alone. It's a sensor inside a bigger system, and it's worth connecting two dots.
Back in lesson 6.1, you saw CI/CD run by agents: processes that have an automatic checking step before something moves forward. The anomaly sweep is exactly that applied to finance. Instead of checking code before deploy, you check entries before the close becomes truth. Same logic: an automatic verification layer up front, human decision behind it.
And on the "Harness Engineering" map, the sensors and observability piece is the idea of instrumenting your system so it warns you when something falls outside expectations. Here, AI sweeping entries is your controllership sensor: it observes the volume and triggers the alert. You're not trading control for AI, you're instrumenting control so it sees further. Fair?
Do it now
Grab an entry database you already have on hand, your real task or last month's close expenses. It doesn't need to be everything, take one period.
Ask AI to do a first sweep, being specific about what "out of pattern" means:
- List entries whose value is far above the average for their own category.
- Point out possible duplicates (same amount, vendor, and nearby date).
- Flag entries in unusual cost centers or accounts for that type of expense.
- For each suspect, state the reason for the alert in one line.
Now do your part, which is the part that matters: take the first five suspects and classify each one as error, legitimate exception, or false positive. Count how many were false positives. That number is your calibration: it shows you how much AI over-lights, and why the human filter is not optional.
Practice
1. In the AI controllership choreography, what is AI's correct role?
2. AI lit up an entry as an anomaly, but investigating you find it was a year-end bonus, perfectly legitimate. What is this?
3. Why does AI anomaly sweeping resemble the CI/CD run by agents from lesson 6.1?
For the board
On reachsampling leaves the rest passing in the dark. The AI's strength is scanning every line, not reading better than you.
On the false positivenot every suspect is guilty. It is the predictable cost of scanning, and the human filter exists so it does not become an unfair accusation.
On the parallelit is the same design as CI: automatic checking on the way in, human judgement on the way out. Only with entries instead of code.
Thanks for the feedback. It helps sharpen the next lesson.