Choreography: due diligence and mass triage
The choreography of using AI to scan thousands of documents in a due diligence: the machine flags the points of attention and you investigate each one, paying the cost of the false positive to gain reach.
An acquisition lands on your desk: twelve hundred files between contracts, amendments, and guarantees, an eight-day deadline. At human pace your team manages thirty documents a day at the very top of its game, so either you read everything and blow the deadline, or you read a sample and pray the bomb isn't in what got left out. You upload the twelve hundred to AI and ask for one specific thing: don't conclude anything, just flag change-of-control clauses, open pending issues, and deadlines about to expire. It scans everything in minutes and hands back a list of forty flagged points, most of them a loose clause with no real effect, a false positive that only costs you the time to dismiss it. Two of them are serious: a guarantee expiring next week and an amendment that changes the deal's price. You didn't read the twelve hundred, you read the forty, with your full attention, and the due diligence comes in on time covering the whole volume, not the fraction your week could fit.
The committee wants your position on an eight-hundred-loan financing portfolio in three days, to decide the transfer price. Reading early-termination clauses, guarantees, and covenants one by one is weeks of work nobody has. You ask AI to scan the eight hundred and flag only what changes the portfolio's value: a violated covenant, insufficient collateral, an off-market rate. It hands back fifty flagged contracts; most is noise, a standard clause the system misformatted and picked up by mistake, and you dismiss it in minutes. Two flagged ones, though, show unprovisioned default that knocks seven percent off the portfolio's price. You'd rather have that noise than risk missing that contract among the other seven hundred and fifty nobody would have had time to open, and the position goes out in three days covering the entire portfolio.
A client is about to sign a big contract and asks you to review their entire vendor base beforehand, because the new contract has an exclusivity clause. There are four hundred old agreements scattered across folders. In one of them there could be a clause that conflicts with the new exclusivity and creates a lawsuit. Finding that needle by reading everything in detail doesn't fit the deadline.
The agency your company is about to absorb hands over six hundred client and vendor contracts, and you have one week to map every image-use right and exclusivity clause before signing the merger. Reading all six hundred with the same magnifying glass doesn't fit the week. You ask AI to scan everything and flag only two things: image assignment with no defined term, and exclusivity with an active client. It hands back thirty-two flagged contracts; most is exclusivity that already expired months ago, a false positive you dismiss by checking the date. One of them, though, locks the absorbed company into a two-year exclusivity deal with a direct competitor of your biggest client, and nobody had seen it before. You didn't read the six hundred, you read the thirty-two, and found the clause that would have locked up the operation after signing.
The senior opening went live and nine hundred resumes came in over five days, and your team reads about forty applications carefully per day at the very top of its game. The best candidates accept another offer before you finish reading the stack. You ask AI to scan the nine hundred and flag only what matches the role's three criteria: minimum time on the stack, stated seniority, and start-date availability. It flags one hundred ten resumes; a good chunk is a false positive, people who mentioned the keyword without the real experience, and you dismiss it in a quick phone pass. Four of them, though, are strong candidates buried on page twenty that nobody on your team would have opened at manual pace. The interview short list comes out on time, covering the entire stack, not just the first batch that came in.
A wave of twelve hundred support tickets, interview transcripts, and open-ended NPS answers landed in your lap, and the committee wants next quarter's prioritization in four days. Reading them one by one covers a fraction and prioritizes the roadmap based on what you had time to see, not what users feel most. You ask AI to scan the twelve hundred and flag only mentions of three pain points you suspect carry weight. It flags one hundred fifty reports; many are one-off complaints with no pattern, a false positive you dismiss by looking at the volume. But a cluster of thirty tickets, which nobody had grouped before, points to an onboarding problem no formal research had captured. The quarter's PRD goes into the committee backed by the entire volume of signal, not the sample your week could fit.
You inherited a pipeline of eight hundred stalled opportunities in the CRM, each with a history of emails, call notes, and old proposals, and the sales director wants to know by Friday which ones still have life before the quarterly forecast. Reviewing deal by deal takes weeks you don't have. You ask AI to scan the eight hundred and flag only a sign of an unresolved objection or a contact gone silent for more than thirty days. It flags one hundred twenty opportunities; most is a deal that's just genuinely slow, a false positive you dismiss with a look at the history. Twelve of them, though, show a decision-maker who changed companies two months ago without anyone updating the CRM, a phantom opportunity inflating your forecast. The number that goes to the director on Friday already comes out clean, without the twelve that would never close.
A big client demands an audit of your SLA compliance before renewing, and there are a thousand work orders from the last six months, each with a different deadline, deviation, and exception. Checking them one by one at human pace doesn't fit the deadline the contract gives you. You ask AI to scan the thousand and flag only deviations above twenty percent of the agreed deadline and exceptions with no recorded justification. It flags seventy orders; a good chunk is a one-day deviation on a holiday nobody disputes, a false positive you dismiss on the spot. But nine orders show a delay pattern on the same route, every Monday, that nobody had put together before seeing the list. You walk into the renewal meeting with the real bottleneck mapped, not with a sample of orders you happened to have time to look at.
The regulator asked for evidence and you have ten days to scan fifteen hundred records between contracts, data flows, and consent forms for processing outside policy. Reading everything in detail blows the deadline, and reviewing a sample and praying the non-compliance isn't in what got left out is a risk your position can't afford to take. You ask AI to flag only three signals: missing legal basis, consent with no collection date, and personal data flow with no recorded purpose. It flags ninety records; many are an old, already-discontinued form, a false positive you dismiss quickly. Six, though, show an active data flow involving a minor with no legal basis at all, a finding that goes straight to the risk committee. The evidence you deliver to the regulator covers the fifteen hundred, not the sample that would have fit your deadline.
The codebase of an acquired company is about to be integrated before go-live, and there are twelve hundred files between services, deploy scripts, and configs where a hardcoded credential or a vulnerable dependency could be hiding. Reviewing the whole repository by hand doesn't fit the window, and shipping without looking is inviting the incident into production. You run AI to scan the twelve hundred and flag only three signals: a plain-text key, a dependency with an open CVE, and a config pointing at the wrong environment. It flags fifty-five files; most is a false positive, a string that looks like a key but is just a test identifier, and you dismiss it in a quick read. Three files, though, have a hardcoded production database credential that nobody had seen in the original code review. Go-live happens with the entire repository scanned, not with the sample the sprint would have had time to review.
Before redesigning the checkout, you gather six hundred recorded sessions, survey responses, and usability test reports, and the design sprint starts in three days. Watching it all at human pace takes weeks the sprint doesn't have. You ask AI to scan the six hundred sessions and flag only long hesitation moments, repeated clicks on the same field, and explicit frustration comments. It flags eighty sessions; many are just someone reading calmly, a false positive you dismiss by watching thirty seconds. Twelve sessions, though, show the same pattern of repeated clicks on the zip code field, a finding no formal research had named before. The redesign prioritization enters the sprint backed by the real volume of sessions, not a sample of twenty your calendar could fit.
The board decides in one week whether the company enters a new market, and there are seven hundred dusty documents, research reports, committee minutes, feasibility studies shelved over the last five years, that might hold the reason a similar attempt didn't work before. Reading it all would take a month the decision doesn't have. You ask AI to scan the seven hundred and flag only mentions of this market and recorded reasons for dropping it. It hands back eighteen flagged documents; most is a passing mention, a false positive you dismiss in seconds. One of them is the minutes of a committee from three years ago that rejected the same entry for the same reason nobody remembered. You bring that to the board not as a guess, but as institutional memory recovered from a file nobody would have reopened on their own.
Let me provoke you with some math before any promises. You don't have the problem of reading a contract well; you know how to do that better than any machine. Your problem is volume: it's twelve hundred documents and eight days. The bottleneck was never your reading, it was your schedule. And that's exactly where AI comes in, in the volume you'd never scan in time on your own. It doesn't come to read better than you. It comes to read what you wouldn't have had time to open.
The core idea of this lesson. In a due diligence or mass triage, AI does one thing very well: it scans hundreds or thousands of documents and FLAGS the points of attention. It points, it doesn't conclude. The judgment stays yours: investigating each flagged point, understanding the legal impact, and prioritizing. The cost to name is the false positive and the false negative, and that's why the human filters while AI extends the reach. The machine is the metal detector; you're the one who digs.
01The bottleneck was never your reading, it was the volume
Take the acquisition scenario. Twelve hundred files, eight days. At human pace, your team opens a fraction and hopes what matters is in it. That's not triage, that's sampling with faith. The risk lives exactly in the documents nobody had time to open.
Notice where the squeeze is. It's not in the quality of your reading, which is high. It's in the quantity your week can hold. AI doesn't fix what you already do well; it attacks what your schedule can't hold. It scans all twelve hundred and hands you back a short list of what to look at first. What was a sample of thirty becomes a scan of twelve hundred with the points of attention marked.
It's the difference between buying blind and buying with the list of what to investigate in hand. Fair?
02The choreography: AI flags, you investigate
Think of a metal detector on the beach. It doesn't dig anything up and it doesn't tell you what's underneath. It beeps. It beeped, you mark the spot and dig. Sometimes it's a gold coin, sometimes it's a bottle cap. The detector covers the whole beach in an afternoon, which would take you weeks with a blind shovel. But you're the one who decides what's worth digging up.
AI in the triage is that detector. It scans the documents and FLAGS the points of attention: a change-of-control clause, an open pending issue, a critical deadline coming up, an inconsistency between two contracts that should match. It points to "look here." It doesn't conclude "this is a problem." The conclusion is yours.
And here's the part that stays entirely yours: investigating each flagged point, understanding the real legal impact, and prioritizing. AI doesn't know whether that change-of-control clause is fatal to the deal or irrelevant in the context of the operation. You know. It hands you the point; you hand back the judgment.
03The cost to name: false positive and false negative
Anti-hype time, because every tool has a cost and hiding the cost is selling an illusion. AI's scan errs in two ways, and both matter.
The false positive is when it flags a point that's nothing. It marked "change-of-control clause" and, when you open it, it's a generic mention with no effect. It cost you the time of investigating something harmless. The false negative is the opposite and more dangerous: it let a real problem slip through, didn't flag it, and it stays invisible in the pile. Like on the beach, the detector beeps at a bottle cap (false positive) and might not beep at a coin buried too deep (false negative).
That's why the choreography's design is deliberately asymmetric. You calibrate AI to flag generously, accepting false positives, because the false positive only costs the human's time filtering. The false negative costs the deal. It's cheaper to dismiss ten false alarms than to miss one real problem. AI extends the reach; the human filters the noise. It doesn't replace your judgment, it increases the size of the beach you can cover.
04The economic frame: compressing weeks into days
Now the part that pays the bill. AI's value here isn't "reading better," it's delivering the short list of what to look at first. And that list compresses weeks into days.
Run the numbers on the credit audit case. Eight hundred contracts, the committee wants the position in three days. At human pace, finding which ones carry a pending issue or a clause that changes the portfolio's value is weeks of work. With the scan, AI goes through all eight hundred and hands you back, say, sixty marked as points of attention. You don't investigate eight hundred; you investigate sixty, with your sharp judgment, and you still have time to spare. The reach went from a sample to the entire base, and your specialist time was spent where it pays off: judging the cases that matter.
This is the real gain, no hype: AI doesn't replace you, it changes the arithmetic of the operation. What didn't fit in eight days now fits. What was a sample becomes a scan. And your name, which signs the due diligence at the end, now sits on top of an analysis that covered the whole volume, not a bet on the fraction you had time to read.
This connects directly with lesson 6.1, on processes with checking: AI's scan is just the first pass, and it enters a flow where the human checks before it becomes a conclusion. And in the map The Harness Engineering, from the Machine Room, it's the sensor layer: AI is the sensor that detects signal in the volume; you're the one who reads the sensor and decides what to do with it.
05How to set up the triage without fooling yourself
Put it all together into a choreography you can run the next time a giant folder lands on your desk. First, you define what AI should flag: the list of signals that matter for this operation (change of control, pending issue, critical deadline, inconsistency between documents, the clause that conflicts with the new lock-in). You give the target; AI doesn't guess what's relevant to your business.
Second, AI scans the whole volume and hands back the short list with each flagged point and where it is. Third, and this doesn't leave your hands: you investigate point by point, dismiss the false positives, measure the real legal impact of each one that's left, and prioritize. Fourth, you close out knowing you covered the entire volume, not a sample.
The mistake that sinks people here is confusing the detector with the prospector. Whoever treats AI's beep as a ready-made conclusion signs off on what the machine marked without digging, and then the false positive turns into a wrong recommendation and the false negative stays buried. AI extends the reach; the judgment is yours. It covers the beach; you dig. Fair?
Do it now
Take a real folder of documents you'd need to triage (your real task works well) and design the choreography before throwing anything into AI. Answer in four blocks:
- THE VOLUME: how many documents are there, and how long would your team take to read them all carefully at human pace? Name the bottleneck: is it your reading or your schedule?
- THE SIGNALS: list 3 to 5 points of attention AI should FLAG in this specific operation (e.g.: change of control, open pending issue, critical deadline, a clause that conflicts with a new lock-in, an inconsistency between two contracts). This is the target only you can define.
- THE COST: for this folder, what's worse, a false positive (it flags what's nothing) or a false negative (it lets something through)? Decide whether you calibrate the scan to flag generously, and why.
- THE JUDGMENT: describe in one sentence what stays YOURS after the scan, meaning, what you investigate, measure, and prioritize, and who signs the due diligence at the end.
If you can't fill in block 2, AI will flag noise. The target comes from you; the scan comes from it.
Practice
1. In a due diligence with twelve hundred contracts and an eight-day deadline, what is the correct division of labor between AI and the lawyer?
2. Why, when calibrating the scan, do you usually accept more false positives than false negatives?
3. Which sentence best describes the economic gain of using AI in mass document triage?
Fair? The message closes like this: you're not outsourcing your judgment, you're stretching your reach. AI is the metal detector that covers the whole beach in an afternoon and beeps at the points; you're the prospector who digs, dismisses the bottle cap, and recognizes the gold. The false positive is the toll you pay on purpose so you don't risk the false negative. And in the end, when the due diligence goes out with your name on it, it covered the whole volume, not the fraction you had time for. That's compressing weeks into days without giving up the judgment.
For the board
On the bottleneckit was never your reading, it was your calendar. Twelve hundred documents in eight days is not a competence problem.
On calibrationit is cheaper to discard ten false alarms than to miss one real problem buried in the pile.
On the gainwhat used to be a sample becomes a scan of the whole volume. Your name signs an analysis that covered everything, not a fraction.
Thanks for the feedback. It helps sharpen the next lesson.