Choreography: bias-free screening
The choreography that keeps AI from reproducing the company's historical bias in resume screening: explicit criteria written before you run AI, never after, and a calibration audit on the result before any candidate advances a stage.
You ask AI to "pick the three best suppliers from this list, the way we always choose around here", without listing what counts as best. It learns from the historical purchasing pattern and repeats the same old preference, including the quirks nobody had noticed, like always passing over a small supplier for seeming "less reliable". The criterion was never written down, only inherited, and now it's automated.
You ask AI to "point out the customers with the best credit profile to release a higher limit", without defining the yardstick. It learns from who got a high limit in the past and repeats the pattern: neighborhood, type of job, a specific age bracket weigh more than they should. The credit criterion was never written out in full, only inherited from the old portfolio, which already had bias built in.
You ask AI to "prioritize the cases that deserve a settlement instead of going to trial", without listing what makes a case a priority. It learns from past settlements and reproduces the same tendency: cases from big clients always moved up the queue, cases from individuals were left behind. The prioritization criterion was never written down, only copied from the firm's habit.
You ask AI to "select the right creators for this campaign, based on what worked before", without saying what counts as right. It learns from past campaigns and repeats the pattern: one specific creator profile always came out on top, while regional niche creators were never considered, even with better engagement. The selection criterion never became a list, it just stayed in the taste of whoever approved it before.
You open AI and ask for something simple: "rank these 300 resumes for the senior developer opening, from best to worst fit". You didn't give any criteria, you figured it was obvious. AI hands back a clean ranking, with a score and a justification for each name. It looks professional, it looks objective. Except it didn't invent the criterion out of nowhere: it learned from the pattern of what "worked" in the company's recent hires, and those hires had already been biased for years. The ranking favors a traditional university, a stint at three well-known big companies, a specific type of name at the top of the list. Nobody wrote this rule down anywhere, it's hidden behind a score from 0 to 10 that looks like pure math. Two weeks later, someone on the diversity team notices: zero resumes from public universities outside the capital made it past the first round. You try to explain the criterion to the committee and realize there isn't one to defend, just a number AI generated on its own.
You ask AI to "prioritize the features users request the most", without defining what counts as a valid request. It learns from the ticket history and amplifies whoever already complained the loudest: the big-customer segment, with more of a voice, always on top, while the average user, who just quietly gives up, never shows up. The prioritization criterion was never written down, only inherited from whoever shouted loudest.
You ask AI to "point out which salespeople deserve the best account portfolio next quarter", without listing the merit criterion. It learns from who already had a good portfolio before and reinforces the pattern: whoever started with a big account keeps getting big accounts, and whoever started with a small account never gets out of there, even while performing better with what they had. The criterion was never written down, it just repeated the usual distribution.
You ask AI to "point out which drivers deserve the most profitable routes", without defining the merit criterion. It learns from the historical route-distribution pattern and reproduces it: whoever already had a good route keeps having one, whoever started in a hard region never leaves it, regardless of real performance review. The criterion was never written down, only inherited from habit.
You ask AI to "point out which areas deserve a priority audit this year, based on the history", without listing the real risk criterion. It learns from the pattern of past audits and repeats it: the same area that's always audited stays on top, while a new area, with growing real risk, never shows up on the list. The prioritization criterion was never written down, only copied from the previous cycle.
You ask AI to "point out which developers deserve a promotion to tech lead, looking at the team's history", without listing the criterion. It learns from who was promoted before and repeats the pattern: whoever worked on the "visible" projects rises, whoever kept the base stable quietly never shows up. The promotion criterion was never written down, only inherited from whoever got the most spotlight.
You ask AI to "point out which users to interview to validate the next prototype, based on who participated before", without defining the representativeness criterion. It learns from the past recruiting pattern and repeats it: always the same profile of engaged user with free time, never the user with a disability or the one who uses the product in a different context. The sampling criterion was never written down, only inherited from whoever was easiest to schedule.
You ask AI to "point out which business units deserve more investment next year, looking at the history", without defining the merit criteria. It learns from the previous funding pattern and repeats it: the units that already received more keep receiving even more, even when the real return was lower. The investment-prioritization criterion was never written down, only inherited from the usual distribution.
Notice something uncomfortable in the HR example above. Nobody decided, in writing, that a traditional university was worth more points. Nobody decided that a specific name should rise in the ranking. It simply happened, because AI needed some criterion to rank by and, in the absence of one from you, it used the only one it had on hand: the pattern of what "worked" before. And if what worked before was already skewed, the pretty numerical ranking just gave the same old bias a new set of clothes.
The core idea of this lesson. The right choreography is never "run AI and then audit the result if something looks off". It's the opposite: first you write out, in full, the criteria that actually matter for the opening and what can never count as a criterion (name, inferred gender, inferred age, address, photo). Then AI applies that yardstick across a volume of resumes your own eyes would never cover alone. And only after that comes the calibration audit: checking whether the cut's distribution matches the pool that entered the funnel, before any candidate advances a stage.
01The classic mistake: letting AI infer the criterion
Let's understand why this happens, because it isn't malice from the tool, it's the physics of how it works. When you don't write the criterion, AI doesn't end up with no criterion at all, it infers one from the pattern that already exists in the data it has: past hires, the resumes that "looked" good, what the manager approved before. The problem is that, if that historical pattern already carried bias, AI doesn't correct it, it learns it and amplifies it.
Notice the most common discriminatory proxies, because they hide in fields that look harmless. The name carries clues about gender and ethnic origin, even if nobody asked for that as a criterion. The zip code or address carries clues about race and social class, since housing is still deeply segregated. The university and graduation year carry clues about class and age, because attending a traditional university full-time is a privilege not every good professional had. None of these three fields is, on its own, a crime. The problem is using them without noticing they're weighing on the final score.
The most cited real-world case is Amazon's, revealed by Reuters in 2018: the company built a screening tool that learned from ten years of resumes received, mostly from men, in an industry historically dominated by men. The tool learned to penalize resumes containing the word "women's" (as in "captain of the women's chess club") and to downgrade graduates of women-only colleges. Nobody programmed this on purpose. It was the historical pattern, reflected back with the appearance of neutral math. Amazon scrapped the tool before using it in production, but the case became the canonical example of what happens when nobody defined the explicit criterion beforehand.
02Step 1: explicit criteria, written down, before you open AI
Here's the step that fixes the problem at the root, and it happens BEFORE any prompt. Before running any screening, you write, in a simple document, two lists.
The first list is what actually counts: the opening's real minimum competencies, not the manager's wish list inflated with "it would be nice if they also knew". Ask: what does this person need to know how to do in the first month to deliver the work? That's a criterion. "Graduated from a specific university" is not.
The second list is what can never count: name, inferred gender, inferred age (including graduation year, which is a direct proxy for age), address, photo, and any trait with no direct relationship to delivering on the role. You hand this list to AI as an explicit constraint: "ignore these fields, even if they appear on the resume".
03Step 2: AI applies the criteria at a volume your eyes can't cover
With both lists in hand, now AI's strength kicks in. Three hundred resumes, applying the same yardstick, without getting tired by resume two hundred and eighty, without giving more attention to the first resume of the day than to the last one. This is genuine mechanical work, and this is where AI beats any human: applying a fixed criterion, consistently, across a volume nobody would have the patience to repeat identically from resume one through three hundred.
Notice the difference in role compared to the previous step. In step one, you decided what matters. In step two, AI just executes the yardstick you designed, resume by resume, and hands you back the ranking with justification tied to the explicit criterion, not to a loose "looks like a good profile".
04Step 3: the calibration audit before advancing any candidate
Here's the step most people skip, and it's exactly what prevents the disaster from the opening example. After AI has applied the criterion and generated the cut, you look at the distribution of the result across whatever demographic groups are available, whenever the law and internal policy allow this kind of aggregate check. The question is simple: does this distribution match the pool that entered the funnel, or is there a disproportionate drop in some group?
A useful heuristic for knowing whether a deviation deserves investigation is the so-called four-fifths rule, used in selection analyses in the United States as a warning signal for unequal impact: if a group's pass rate falls below 80 percent of the rate of the best-performing group, that's a signal to investigate the criterion, not a Brazilian law to follow to the letter, but a borrowed thermometer that works well as an alarm. If, in the opening example, zero resumes from public universities made it past the first round, that number alone is already the alarm going off.
And here comes the most important point in this section: when calibration flags a deviation, you investigate the criterion, you don't discard the candidate or ignore the signal hoping it slips by unnoticed. Maybe the "university X" criterion was, unintentionally, working as a proxy for class. Maybe "years of experience" was penalizing whoever changed careers later. The audit doesn't tell you who to hire, it tells you where the criterion might be skewed, and it's up to you to fix the criterion and run it again, not simply accept the result because "AI calculated it".
05The golden rule of this choreography
Let's close the reasoning in one sentence worth sticking in your head: explicit criteria first, never let AI "decide what matters" on its own, calibration audit always before advancing anyone a stage. The order isn't a suggestion, it's what separates a screening process that scales your judgment from one that scales your historical bias wearing new mathematical clothes. AI gives you real reach, three hundred resumes read with the same yardstick. But the yardstick is yours, written beforehand, and checking that it was applied fairly is also yours, always afterward. Fair enough?
Do it now
Take a real opening you have posted right now, your real task or another recent one.
- Write the list of what ACTUALLY counts: the minimum competencies this person needs to have to deliver the work in the first month. Cut from the list any "it would be nice if" that isn't essential.
- Write the list of what can NEVER count: name, inferred gender, inferred age (including graduation year), address, photo, and any other trait with no direct relationship to delivering on the role.
- Run the screening with AI using ONLY these two lists as criteria, explicitly asking it to ignore the fields from list two even if they appear on the resume.
- After the cut, look at the distribution of the result across the groups you can identify in aggregate form (university origin, region, approximate age bracket). Does any drop look disproportionate?
- If something jumps out, go back to the criterion in list one before touching any candidate. The criterion is what gets investigated first, always.
You've just run the choreography that turns "AI ranked it" into "I defined the yardstick, AI applied it, and I checked that it was fair".
Practice
1. What's the essential difference between defining criteria BEFORE running AI and only auditing the result AFTER, if something looks strange?
2. On a resume, which of these fields tends to work as a discriminatory proxy, even if nobody defined it as a criterion on purpose?
3. The calibration audit showed that one group had a pass rate well below the four-fifths rule relative to the best-performing group. What should you do?
</content>
For the board
On the classic mistakewith no criterion of yours, the AI uses the only one at hand: the pattern of what worked before.
On the orderexplicit, written criteria before you open the AI. They are what exists to be defended later.
On the proxyinnocent-looking fields carry information about age and class, and weigh on the ranking without ever appearing as a criterion.
Thanks for the feedback. It helps sharpen the next lesson.