Business: HR · Lesson N.rh.3

Choreography: bias-free screening

The choreography that keeps AI from reproducing the company's historical bias in resume screening: explicit criteria written before you run AI, never after, and a calibration audit on the result before any candidate advances a stage.

Examples for

You ask AI to "pick the three best suppliers from this list, the way we always choose around here", without listing what counts as best. It learns from the historical purchasing pattern and repeats the same old preference, including the quirks nobody had noticed, like always passing over a small supplier for seeming "less reliable". The criterion was never written down, only inherited, and now it's automated.

Notice something uncomfortable in the HR example above. Nobody decided, in writing, that a traditional university was worth more points. Nobody decided that a specific name should rise in the ranking. It simply happened, because AI needed some criterion to rank by and, in the absence of one from you, it used the only one it had on hand: the pattern of what "worked" before. And if what worked before was already skewed, the pretty numerical ranking just gave the same old bias a new set of clothes.

The core idea of this lesson. The right choreography is never "run AI and then audit the result if something looks off". It's the opposite: first you write out, in full, the criteria that actually matter for the opening and what can never count as a criterion (name, inferred gender, inferred age, address, photo). Then AI applies that yardstick across a volume of resumes your own eyes would never cover alone. And only after that comes the calibration audit: checking whether the cut's distribution matches the pool that entered the funnel, before any candidate advances a stage.

01The classic mistake: letting AI infer the criterion

Let's understand why this happens, because it isn't malice from the tool, it's the physics of how it works. When you don't write the criterion, AI doesn't end up with no criterion at all, it infers one from the pattern that already exists in the data it has: past hires, the resumes that "looked" good, what the manager approved before. The problem is that, if that historical pattern already carried bias, AI doesn't correct it, it learns it and amplifies it.

Notice the most common discriminatory proxies, because they hide in fields that look harmless. The name carries clues about gender and ethnic origin, even if nobody asked for that as a criterion. The zip code or address carries clues about race and social class, since housing is still deeply segregated. The university and graduation year carry clues about class and age, because attending a traditional university full-time is a privilege not every good professional had. None of these three fields is, on its own, a crime. The problem is using them without noticing they're weighing on the final score.

The most cited real-world case is Amazon's, revealed by Reuters in 2018: the company built a screening tool that learned from ten years of resumes received, mostly from men, in an industry historically dominated by men. The tool learned to penalize resumes containing the word "women's" (as in "captain of the women's chess club") and to downgrade graduates of women-only colleges. Nobody programmed this on purpose. It was the historical pattern, reflected back with the appearance of neutral math. Amazon scrapped the tool before using it in production, but the case became the canonical example of what happens when nobody defined the explicit criterion beforehand.

02Step 1: explicit criteria, written down, before you open AI

Here's the step that fixes the problem at the root, and it happens BEFORE any prompt. Before running any screening, you write, in a simple document, two lists.

The first list is what actually counts: the opening's real minimum competencies, not the manager's wish list inflated with "it would be nice if they also knew". Ask: what does this person need to know how to do in the first month to deliver the work? That's a criterion. "Graduated from a specific university" is not.

The second list is what can never count: name, inferred gender, inferred age (including graduation year, which is a direct proxy for age), address, photo, and any trait with no direct relationship to delivering on the role. You hand this list to AI as an explicit constraint: "ignore these fields, even if they appear on the resume".

1 Criterion written first 2 AI applies large volume 3 Calibration audits distribution 4 Decision advances a stage steps 1 and 3 are yours, written down and defensible no candidate advances without passing through step 3

03Step 2: AI applies the criteria at a volume your eyes can't cover

With both lists in hand, now AI's strength kicks in. Three hundred resumes, applying the same yardstick, without getting tired by resume two hundred and eighty, without giving more attention to the first resume of the day than to the last one. This is genuine mechanical work, and this is where AI beats any human: applying a fixed criterion, consistently, across a volume nobody would have the patience to repeat identically from resume one through three hundred.

Notice the difference in role compared to the previous step. In step one, you decided what matters. In step two, AI just executes the yardstick you designed, resume by resume, and hands you back the ranking with justification tied to the explicit criterion, not to a loose "looks like a good profile".

04Step 3: the calibration audit before advancing any candidate

Here's the step most people skip, and it's exactly what prevents the disaster from the opening example. After AI has applied the criterion and generated the cut, you look at the distribution of the result across whatever demographic groups are available, whenever the law and internal policy allow this kind of aggregate check. The question is simple: does this distribution match the pool that entered the funnel, or is there a disproportionate drop in some group?

A useful heuristic for knowing whether a deviation deserves investigation is the so-called four-fifths rule, used in selection analyses in the United States as a warning signal for unequal impact: if a group's pass rate falls below 80 percent of the rate of the best-performing group, that's a signal to investigate the criterion, not a Brazilian law to follow to the letter, but a borrowed thermometer that works well as an alarm. If, in the opening example, zero resumes from public universities made it past the first round, that number alone is already the alarm going off.

And here comes the most important point in this section: when calibration flags a deviation, you investigate the criterion, you don't discard the candidate or ignore the signal hoping it slips by unnoticed. Maybe the "university X" criterion was, unintentionally, working as a proxy for class. Maybe "years of experience" was penalizing whoever changed careers later. The audit doesn't tell you who to hire, it tells you where the criterion might be skewed, and it's up to you to fix the criterion and run it again, not simply accept the result because "AI calculated it".

05The golden rule of this choreography

Let's close the reasoning in one sentence worth sticking in your head: explicit criteria first, never let AI "decide what matters" on its own, calibration audit always before advancing anyone a stage. The order isn't a suggestion, it's what separates a screening process that scales your judgment from one that scales your historical bias wearing new mathematical clothes. AI gives you real reach, three hundred resumes read with the same yardstick. But the yardstick is yours, written beforehand, and checking that it was applied fairly is also yours, always afterward. Fair enough?

Do it now

Do it yourself

Take a real opening you have posted right now, your real task or another recent one.

  1. Write the list of what ACTUALLY counts: the minimum competencies this person needs to have to deliver the work in the first month. Cut from the list any "it would be nice if" that isn't essential.
  1. Write the list of what can NEVER count: name, inferred gender, inferred age (including graduation year), address, photo, and any other trait with no direct relationship to delivering on the role.
  1. Run the screening with AI using ONLY these two lists as criteria, explicitly asking it to ignore the fields from list two even if they appear on the resume.
  1. After the cut, look at the distribution of the result across the groups you can identify in aggregate form (university origin, region, approximate age bracket). Does any drop look disproportionate?
  1. If something jumps out, go back to the criterion in list one before touching any candidate. The criterion is what gets investigated first, always.

You've just run the choreography that turns "AI ranked it" into "I defined the yardstick, AI applied it, and I checked that it was fair".

Practice

1. What's the essential difference between defining criteria BEFORE running AI and only auditing the result AFTER, if something looks strange?

2. On a resume, which of these fields tends to work as a discriminatory proxy, even if nobody defined it as a criterion on purpose?

3. The calibration audit showed that one group had a pass rate well below the four-fifths rule relative to the best-performing group. What should you do?

</content>

For the board

On the classic mistakewith no criterion of yours, the AI uses the only one at hand: the pattern of what worked before.
On the orderexplicit, written criteria before you open the AI. They are what exists to be defended later.
On the proxyinnocent-looking fields carry information about age and class, and weigh on the ranking without ever appearing as a criterion.
What did you think of this page?
Would you recommend this page to someone on your team?