Choreography: the exception that resolves itself
The choreography that takes operational exceptions in bulk, delay, missing item, invalid address, and carries them from triage to resolution, always with a human gate before any irreversible action.
Tuesday morning, logistics dashboard open: 340 deliveries for the day, and 22 of them are already born crooked, incomplete address, missing item on the packing list, recipient who won't answer the phone. Before, that used to become a whole morning of an analyst opening order after order, calling, deciding what to do with each one. Today AI scans the 340 in seconds, splits the 22 by problem type, and already proposes the action: resend the notification, reschedule delivery, escalate to the call center. The temptation is to let it fire everything off on its own and get the morning back. The problem is that inside those 22 there's an eighteen thousand dollar order AI wants to cancel for "invalid address", when actually the customer just moved to a different building in the condo. Before anything goes out to the customer or the system, someone needs to approve it. That's the gate that separates "the exception resolved itself, correctly" from "the exception resolved itself, and now it's a complaint".
The bank reconciler flags 340 unmatched entries this month, and AI classifies each one, duplicate, unforeseen fee, rounded exchange rate, and suggests the automatic write-off. Before reversing or reclassifying in bulk, someone in finance checks a sample of the largest amounts, because a reversal on the wrong account doesn't undo itself with one click.
340 supplier contracts go into the annual review, and AI flags 22 with an adjustment clause outside the company's standard, already proposing the standard corrective amendment for each. Before firing off the 22 amendments to suppliers, the legal team reviews a sample, because a wrong amendment sent to a strategic supplier doesn't get taken back once signed.
The email send ruler flags 22 campaigns with a broken link or off brand asset, and AI already proposes pausing the send and swapping the asset. Before pausing campaigns that are already running, someone checks a sample of the 22, because pausing a campaign that was actually fine costs reach that doesn't come back.
The time clock system flags 22 inconsistent punches this month, and AI classifies each one, forgot to clock in, timezone error, attempted fraud, proposing the standard adjustment for each class. Before applying an adjustment that touches payroll, HR checks the sample, because docking an employee for an AI classification error costs trust that an apology doesn't get back.
The billing funnel flags 22 subscriptions with a failed charge, and AI classifies by cause, expired card, exceeded limit, open dispute, proposing to suspend all 22 accounts. Before suspending accounts in bulk, someone on the team checks the sample, because suspending a paying customer's access by mistake turns into a real cancellation.
The CRM flags 22 proposals stalled for more than thirty days with no response, and AI already proposes firing off the standard re engagement email to all of them. Before sending to the 22 contacts, the salesperson checks the sample, because re engaging an account that's actually mid negotiation by phone, outside the CRM, sounds like someone who doesn't know what they're doing.
Friday afternoon, the operational support ticket queue hits 480 open tickets, and the on duty team can handle 60. AI reads the 480, groups them by root cause, connector down, duplicate billing, tracking not updating, and suggests the standard response and technical action for each group. In five minutes, what was a queue nobody could beat becomes 14 cause groups, each with an action ready to fire. The difference between this being a relief or a disaster is in what comes next: the "duplicate billing" group has 60 tickets and the proposed action is to auto refund. Refunding sixty amounts without looking at any of them is betting that AI classified all sixty correctly. One of them, it turns out, was a legitimate second purchase from the same customer on the same day. The human gate audits a sample of the group before letting the bulk refund go out, and that's what turns the disaster into just a tough day of work.
The access sweep flags 22 permissions granted outside the segregation of duties standard, and AI proposes automatically revoking all 22 accesses. Before revoking in bulk, compliance checks the sample, because revoking someone's access in the middle of a critical closing process stalls a process that should be running.
Monitoring flags 22 production error alerts, and AI groups them by service and already proposes the standard rollback for the groups with the highest incidence. Before running a rollback in production, the on duty team checks the sample, because reverting a deploy that was actually fine takes down a fix you just shipped.
The onboarding funnel flags 22 users stuck at the same step, and AI already proposes firing off the standard help email to all 22. Before sending in bulk, the team checks the sample, because sending help instructions to someone who was actually just exploring the product slowly sounds like the company admitting a problem that doesn't exist.
The partnership radar flags 22 agreements with a clause expiring without automatic renewal, and AI already proposes notifying the partner with the standard renewal draft. Before notifying strategic partners in bulk, someone on the team reviews the list, because a wrong renewal notice to a big account turns into a hard conversation that "it was a mistake" doesn't undo.
Friday afternoon, dashboard open, and there are 22 deliveries today with some kind of problem: a crooked address, a missing item, a customer who won't answer. You know exactly what to do with each one, except doing it with 22 at once, by hand, is the kind of task that eats your whole afternoon and still spills into Monday. AI solves this fast: it reads the 22, splits them by type, and already proposes the right action for each group. The temptation is to let it fire everything off and leave early. Except inside that group there might be a case it read wrong, and the action it proposes isn't reversible with a "my mistake". This lesson is about how to resolve exceptions in bulk without giving up the last step that's still yours.
The core idea of this lesson. AI scans a volume of exceptions you'd never review one by one, and proposes the right action for each class. That's a real time gain. But none of those actions go out and touch a customer, a supplier, or a system without passing through a human gate first. That gate's size isn't always the same: it changes according to how reversible the action is. A light action gets a light gate. An action that can't be undone gets a gate that doesn't let anything through without a look.
01The choreography has five steps, and the gate isn't optional
Resolving exceptions in bulk isn't pressing a button. It's a sequence with one exact point that can't be skipped. There are five steps: scan and classify, propose the standard resolution by class, pass through the human gate, execute only what was approved, and close the loop by learning from what was a real exception and what was a false alarm.
AI comes in strong on steps one, two, and four. It's fast at reading volume, grouping by cause, and firing off what's already been cleared. Step three is where the machine stops and waits for you, and that's not a delay, it's the system's correct design.
02Steps 1 and 2: AI scans the volume and already proposes the action
Here's the real gain. You point AI at the entire queue, the 340 orders of the day, the 480 open tickets, the 22 subscriptions with a failed charge, and it does two things the human eye can't do in time: it scans everything, without tiring and without skipping a line, and classifies each exception into a type that already implies the action. Invalid address calls for rescheduling. Missing item calls for resending with a notice. Duplicate billing calls for a refund. It doesn't make up the category out of nowhere, it applies the same criteria you would use, just across 340 cases at once.
The important point here is that it doesn't stop at just the diagnosis, like the anomaly detector you saw earlier in the track. It goes a step further and already proposes the resolution, ready to fire. That changes the risk. Pointing at a suspect is one thing. Preparing a real action, ready to go out, is another. And that's exactly why the next step exists.
03Step 3: the human gate, the heart of this lesson
Here the choreography changes hands. AI has already classified and already prepared the action. But none of those actions, canceling an order, notifying a customer, redirecting cargo, refunding a charge, revoking access, touch the real world before passing through a person. Notice that's different from "review afterward". Reviewing afterward means the customer already received the wrong cancellation by the time you notice. The human gate means the action hasn't existed yet outside the system until you say yes.
That gate's size isn't always the same, and that's where the judgment lives. Think of the action's reversibility as the ruler that decides the control's size. Resending a tracking notification is cheap to undo, if it's wrong, you send an apology and move on. Canceling an eighteen thousand dollar order from a big account, or revoking someone's access in the middle of a critical process, doesn't undo as easily. The more irreversible the action, the tighter the gate needs to be: you go from a sample based audit to a case by case review, one by one, before anything fires.
04Steps 4 and 5: executing what's approved and closing the loop by learning
After the gate, AI fires off only what's been cleared. Notice what changes: it's no longer deciding what to do, it's executing a decision that already passed through your judgment. It's the same principle you saw in the lesson on governing execution, the machine acts, the accountability stays with whoever approved it.
The fifth step is what turns this choreography into a system that gets better over time. After each round, you log what was a real exception and what was a false alarm, the same error versus legitimate exception classification you already trained in the anomalies lesson. That case of the address that had actually just moved to a different building in the condo? That becomes a new rule so AI doesn't get confused again. The system doesn't stand still, it learns from every gate decision you made, and the next scan comes back sharper, with less noise for you to review.
Do it now
Take a type of exception that repeats in your operation, your real task or another one (delivery delay, support ticket, failed charge, out of standard access). Design the choreography in five steps:
- SCAN: describe how AI would read the entire volume and classify each exception by type.
- PROPOSE: for each type, what would be the standard action AI would propose (resend, reschedule, escalate, refund, revoke)?
- GATE: mark, in one sentence, the exact point where a person needs to approve before the action goes out. Also say whether this gate is sample based or case by case, and why.
- EXECUTE: what fires off after approval, with no more intervention from you.
- LEARN: what kind of classification error, if it happens, should become a new rule for the next scan.
You've just designed a system that resolves exceptions at volume without taking you out of command at the moment it matters.
Practice
1. In the exception choreography, which step can never be outsourced to AI without a human gate first?
2. Why does the human gate change size depending on the action's reversibility?
For the board
On the gainthe AI reads the whole volume, sorts it by type and already proposes the right action for each group.
On the kill switchno action that cannot be easily undone goes out without passing a person first.
On the size of the brakeit follows reversibility. Resending a notification costs little; cancelling a large order does not undo the same way.
Thanks for the feedback. It helps sharpen the next lesson.