The audit: never decide about people without checking
AI writes a performance review or a termination recommendation with the same confidence as always, but no termination or promotion can ever come out of an AI score or opinion without an audit. A five-item checklist and the principle that whoever signs the decision is always a person.
A manager asked AI for a performance summary of an employee before deciding on a promotion, and the text that came back was flawless, with arguments that seemed to come from someone who'd closely followed the team all year. It almost became the basis for the decision with no further check. Checking the source of each praised item, two of them came from a project the employee hadn't even worked on, just listed as an observer in the meeting. The confident prose left no trace of the error.
AI calculated an entire team's variable bonus based on the target hit, and wrote a confident paragraph about "fair distribution per current policy". It almost went straight to payroll like that. Redoing the math separately, AI had applied the yardstick of an old policy, already revoked two quarters ago, and three people would have received a value well below the correct one, with the same confident tone as the correct values next to it.
In a labor due diligence, AI raised an entire team's overtime liability and wrote a confident paragraph about "controlled exposure". It almost went into the opinion with no additional check. Redoing it case by case, two of the timesheet records AI cited belonged to a system discontinued a year ago, and the real liability was almost 30% higher. The text sounded like a legal conclusion; the number was fiction wearing a badge.
AI put together the internal employer-branding campaign's climate report and wrote a verdict of "engagement rising among technical teams". It almost became the central argument for the next internal communications budget. Checking the response base, AI had counted silence (whoever didn't answer the survey) as neutral-positive, and real engagement was stagnant, not rising.
An employee is in the middle of a termination process, and AI drafts the opinion based on his review history in the system. Out comes a confident, polished text, with the phrase "recurring low performance and lack of cultural fit over the last two cycles". The manager reads it, agrees with the professional tone, and almost attaches the opinion as the official justification in the process. Before signing, she decides to reread the three original reviews AI cited, one by one, at the source. Two of the three "weak reviews" were from the period he was on medical leave, and the system had simply recorded that cycle's lowest score with no context note at all. AI read the number and treated health leave as ordinary poor performance. The opinion that almost justified the dismissal cited, with all the confidence in the world, exactly the kind of fact that legally becomes a presumption of discriminatory dismissal. The manager wasn't lucky to have noticed: she audited because it was a habit. Whoever doesn't have that habit signs the opinion, and the problem stops being AI's and becomes the company's.
AI evaluated an entire squad's performance from delivery metrics and wrote that the team "wasn't performing" that quarter. It almost became the justification for redistributing people. Looking closely at the source, for half the quarter the squad had been on loan to put out another team's fire, something AI never saw because it wasn't in any system it consulted.
AI ranked the quarter's salespeople by "adjusted performance" to decide who'd be promoted to coordinator, and wrote that one of them was the natural candidate. It almost became a closed decision. Checking the data, AI had compared his performance against a smaller territory target than his own, inflating his relative result compared to peers.
AI evaluated a shift's operators' productivity and wrote that one of them was "consistently below the team average". It almost became the basis for a formal warning. Checking the source, AI had compared the night-shift operator against the day shift's average, which has different volume and conditions, a comparison nobody had asked for and that made no sense.
For the annual conduct report, AI cross-referenced ethics-channel complaints with each cited person's review history and wrote a reassuring paragraph about "no concerning pattern". It almost closed the report like that. Checking case by case, AI had counted two distinct complaints against the same person as one, because their text was similar, masking a pattern that should already have escalated.
AI evaluated a team's commits and code reviews to suggest who deserved a technical promotion, and wrote a confident recommendation about who "contributed most". It almost became direct input for the career committee. Checking the source, AI counted lines of code generated by an automated tool as the personal contribution of whoever merely approved the PR, distorting the whole ranking.
AI analyzed exit interviews of designers who had resigned and wrote that the predominant reason was "lack of growth opportunity". It almost became HR's official conclusion about the team's turnover. Reviewing the original transcripts, half the interviews cited a specific manager as the real reason, a pattern AI diluted by generalizing the cause.
To decide whether to restructure a directorship, AI calculated a synergy index between the areas based on the org charts and wrote a confident merger recommendation. It almost went into the plan headed to the board. Checking the source, AI had added together two teams that had already reported to the same manager for months, artificially inflating the overlap that justified the merger.
Stop for a second on that "almost". In all these cases the wrong recommendation almost became a decision, not because the manager was careless, but because AI hands over a wrong opinion with the same confidence it hands over a correct one. In finance, a wrong number becomes a loss. Here, a wrong opinion becomes a person fired for a reason that never existed, or promoted for the wrong reason. It's the same trap from the whole module, except now the number has a first and last name.
The core idea of this lesson. AI is excellent at writing text and isn't trustworthy by default as the basis for deciding about a person: it can carry bias inherited from the history, strip a fact of its context (medical leave becoming "low performance"), invent a connection that doesn't exist, and hand all of that over with the same confidence as always. That's why AI's recommendation about people is a proposal, never a truth. This lesson closes with a five-item checklist and one non-negotiable principle: whoever signs the decision about a person is always someone named.
01Excellent at text, terrible as a basis for deciding about a person
It's worth understanding why this happens, because understanding it changes how you use the tool. The AI you use is, at its core, a machine for predicting the next word that sounds right. It was trained so the text is fluent, plausible, and confident. Nobody trained it so the reading about a person is fair. These are two different goals, and it only pursues the first one.
The side effect, here, is the most serious in the whole module. The same skill that makes AI write a flawless performance review makes it write, with the same quality of prose, a biased reading or a fact taken out of context. It has no internal sense of "this isn't fair to this person". If the history carries bias (a manager who always rated poorly whoever took leave, a team that logged health absence the same as disengagement absence), AI learns that pattern and repeats it with the look of neutral analysis. In a marketing text, a skewed reading is a slip. In a review of a person, it's a career, a labor lawsuit, a life.
02The paradox of feeling: you think you already audited it
In 2025, METR measured experienced professionals working with and without AI on real tasks, and the result was the opposite of expected: with AI, they were about 19% slower, but they thought they were faster. The feeling of a gain didn't match the measured result.
That same treacherous feeling shows up when deciding about a person. When AI's opinion arrives well written, organized, with a conclusion that sounds professional, your brain registers "this has already been analyzed" and lets its guard down. It's exactly in the most serious decision (firing, promoting, giving or denying a raise) that this false feeling of having already checked costs the most. Reading the opinion isn't auditing the opinion. They're different things, and that difference is the subject of the next section.
03The audit checklist for a decision about people
Here's the heart of the lesson. Five questions, in order, before any AI recommendation about a person becomes a termination, promotion, warning, or raise.
First: is the criterion used explicit, and is it the same one applied to everyone in the same situation? If the opinion says "low performance" without saying against which yardstick, and that yardstick isn't the same one used for colleagues, the criterion doesn't really exist, it only looks like it does.
Second: does the source of each claim exist and check out? Every sentence like "two weak reviews over the last cycles" needs to be checked at the source, one by one. This is the step that catches the review that was, in fact, from a period of health leave.
Third: is there calibration across groups? Look at the whole set: is AI's negative-recommendation rate disproportionate in some group (whoever took leave, whoever has a specific demographic profile, whoever's on a specific team)? A disproportionate pattern doesn't prove individual guilt, but it's a signal that the criterion may be carrying bias inherited from the history.
Fourth: would someone put this recommendation in front of the person and defend it eye to eye? If the answer is "I'd rather she never find out an AI generated this like this", the recommendation isn't ready.
Fifth, and the most important: is the final decision signed by a named person, with a name and responsibility? If the answer is "the system recommended it" with no human name behind it, the decision has no owner, and a decision with no owner can't go out.
04The RESPOND yardstick applied to termination and promotion
The checklist has a principle behind it, and it's the same one that holds up the whole module. In the Three Moves map, the RESPOND move deals with responsibility anchored to a human name. Applied to a decision about a person, it becomes a three-beat yardstick you never collapse.
AI proposes. It can generate the opinion, the score, the suggested action. Use it freely here, this is where it genuinely helps. But a proposal is a proposal: nothing it produces about a person is true just because it produced it.
The human checks. This is the five-question checklist running, with no shortcuts, especially when the decision is serious.
And the person signs. When the decision goes out (the termination letter, the promotion, the warning), the responsibility belongs to someone with a name, always. "The score recommended it" doesn't exist as an excuse, not in an HR conversation, not in a labor complaint. And the legal weight here is real, not just a nice-sounding principle: Súmula 443 do TST already recognizes a presumption of discriminatory dismissal when an employee with a stigma or serious illness is let go without clear justification, and LGPD, in article 20, guarantees the person the right to request review of an automated decision that affects their interests. An AI opinion that carries bias without an audit isn't just an error of judgment. It's real legal exposure, for the person and for the company.
05Auditing isn't distrust: it's protecting the person and the company
Let me clear up a resistance before wrapping up. Some people feel that auditing AI's opinion about an employee is bureaucracy, that it delays a decision that was already made anyway. That frame is wrong, and it's expensive on both sides.
Auditing isn't distrust of the tool, it's professional hygiene, just like you saw in finance. Except here, the cost of not auditing isn't a wrong spreadsheet, it's a person fired for a reason that never existed, or a labor lawsuit the company loses because the decision never had a defensible criterion behind it. The cost of auditing is minutes: rereading three reviews at the source, checking whether the criterion is the same one always used, asking whether you'd defend this eye to eye. The cost of not auditing can be a career, a settlement, a reputation. Minutes against that isn't rework, it's the cheapest insurance there is. AI proposes, you check, and whoever signs protects the person and the company at the same time.
Do it now
Take an AI recommendation about a person that you already have or are about to use (your real task works well): a performance opinion, a score from an evaluation system, a suggested action about someone on the team.
Run the five-item checklist, writing down the answer to each one:
- IS THE CRITERION EXPLICIT? What's the yardstick used, and is it the same one applied to everyone in the same situation?
- DOES THE SOURCE CHECK OUT? Pick the two strongest claims in the opinion and check each one against the original source, one by one. Did they check out?
- IS THERE CALIBRATION ACROSS GROUPS? Looking at the set of similar recommendations, is any group (whoever took leave, whoever's on a specific team) receiving a disproportionate share of negative recommendations?
- WOULD YOU DEFEND IT EYE TO EYE? If the person being evaluated read this opinion in front of you, would you defend it sentence by sentence?
- WHO SIGNS? Write the name of whoever signs this final decision, and the one-line defense that person would give if questioned, in a committee or a hearing.
If you got stuck on any of the five, the decision isn't ready to go out yet.
Practice
1. AI handed over a well-written termination opinion, citing the employee's past reviews. What's the correct reading before using this opinion?
2. The METR (2025) study showed professionals thinking they were faster with AI, when they were actually slower. What's the correct lesson from this for decisions about people?
3. A termination decision based on an AI score ended up being challenged in court, and the score was wrong. Who answers for that decision?
4. AI recommended terminating an employee citing 'recurring low performance', but two of the cited reviews were from a period he was on medical leave. What does this situation exemplify?
</content>
For the board
On what is at stakein finance the error becomes a loss. Here it becomes a person dismissed for a reason that never existed.
On the recommendationwell written and citing appraisals is not proof. It is a proposal until you open every appraisal it cites.
On responsibilityit does not migrate to the tool under any circumstance, including before an employment tribunal.
Thanks for the feedback. It helps sharpen the next lesson.