Twelve out of forty. Then five out of forty. Same degree, same years of experience, same skills, line for line — the only thing that changed was the name at the top.
The case below is a composite — built to show a failure pattern I keep meeting in my work, not a transcript of one client's file — but every number and every mechanic in it is the kind I see when I go looking. Worth walking through slowly.
The Dashboard That Looked Fine
Take a mid-size professional-services firm running an AI resume screener over roughly 1,200 applications a month, cutting that pool to a shortlist of about 220 — an 18 percent pass-through rate. Every number on the recruiting dashboard says the system is working. Pass rate is steady. Time-to-shortlist has dropped since the tool went live. Recruiter satisfaction scores are up, because nobody on the hiring team is drowning in resumes anymore.
None of those metrics measure what I actually care about when I'm assessing a screening system: whether it treats otherwise-identical candidates the same way. A dashboard built around throughput and satisfaction has no way to show you that. You have to go looking for it on purpose, with a test designed to isolate the one variable the dashboard can't see.
The Matched-Pair Test
The test itself is simple to describe and exacting to run. Build 40 matched pairs of resumes — identical stated degree, identical years of experience, identical skills — and vary exactly one thing: the name at the top, chosen to signal a different national or regional origin. Run all 80 resumes through the live system, not a sandbox, in the same week, against the same open requisitions.
In this case: Group A was shortlisted 12 times out of 40 — 30 percent. Group B was shortlisted 5 times out of 40 — 12.5 percent. That's a disparity ratio of 0.42.
I use the four-fifths rule as my working bar — the idea, borrowed from US Equal Employment Opportunity Commission adverse-impact guidance, that a selection rate for one group falling below 80 percent of the rate for another group is worth investigating. It's a useful bright line, not a law that travels automatically across jurisdictions. In the MENA region specifically, there is no equivalent codified statistical threshold I can point to, and I say so plainly to clients rather than borrow authority the number doesn't have. What it gives you is a defensible, pre-agreed trigger for "look closer" — which is exactly what a ratio of 0.42 deserves.
What the Gap Was Actually Measuring
The instinct is to assume the model is reading the name directly. In this case it wasn't. A blind human review of the resumes sitting in the borderline-score band — the ones close to the cutoff either way — traced the gap to a phrasing pattern common in CVs written by non-native English speakers. The model had learned to score that pattern down, on its own, independent of the candidate's actual English-proficiency test result, which was already sitting elsewhere in the same application.
Nobody wrote a rule that said "penalize this phrasing." Nobody coded a proxy for regional origin into the model. It learned the association from training data and applied it quietly, underneath a dashboard that had no reason to flag it, because the dashboard was never built to ask that question.
The Fix, and the Retest
The fix was narrow, once the cause was located: drop the phrasing feature's independent weight and let the real, already-present proficiency score carry that signal instead. Re-run the same 40 pairs. Result: 30 percent versus 26 percent — a disparity ratio of 0.87, inside the bar set before anyone looked at the data.
The retest matters, but the paperwork around it matters just as much. The evidence file for a test like this holds four things: a methodology memo, written and dated before the results existed; the raw pre-remediation shortlist log, timestamped; the 0.8 threshold decision, signed off before anyone saw a single outcome; and the retest results themselves. That ordering — bar set first, results looked at second — is what separates an audit from a story you tell yourself afterward to explain a number you don't like. Skip that order and you don't have evidence. You have a rationalization with a spreadsheet attached.

The Method You Can Run on Monday
I call this the Matched-Pair Audit. It has four steps, and none of them require access to the model's internals:
- Build. Construct pairs identical on every stated qualification, differing only in the one attribute you suspect the system is reading.
- Run blind. Push them through the live system, not a sandbox, under real conditions.
- Set the bar first. Fix your pass/fail threshold before you look at any result, so you cannot rationalize a bad number afterward.
- Document either way. Record the test, the threshold, and the outcome regardless of what it shows. A clean result is also evidence.
Hiring is the easiest place to picture this, because everyone understands a resume. But the method has nothing specifically to do with resumes. Any system that scores, screens, or triages people or cases can be tested this way — a credit-scoring model, a claims-triage system, a fraud-flagging engine, a support-ticket router. The pairs change. The four steps don't.
I've written before about the disclosure test for deciding when an AI decision owes the person it affects an explanation, and the theme repeats: the risk in these systems is rarely dramatic. It's a quiet, defensible-looking number that nobody thought to test against a matched pair. The EU AI Act's obligations for high-risk AI systems, including requirements around data governance and bias examination in Article 10, point at the same underlying discipline — testing before deploying, and being able to show your work. You don't need a regulation to require it of yourself first.
Run the test before you need to explain a number you can't defend. That's the whole argument.
