10 to the 23 AI logo

Domain Atlas / Hiring & employment screening AI

Case fileUnited States (large employer; internal experimental recruiting tool, ~2014–2017)giant deployment

A resume screener that learned the past's bias

Explore this deployment in the PAN Lab ↗

An internal team built an experimental recruiting engine — roughly 500 models scoring resumes one to five stars per role and location — trained on ten years of the company's own hiring decisions, a period whose hires were predominantly male. The models learned that history: they penalized the word 'women's' and downgraded graduates of women's colleges, reading gender proxies as negative signal. The team patched the identified terms but concluded that term-level fixes could not guarantee neutrality against unknown proxies, because the model had learned the pattern rather than the words, and the company scrapped the project around 2017; per the company, recruiters saw the tool's recommendations but it was never used as a sole ranking.[]

What happened

An internal team built an experimental recruiting engine — roughly 500 models scoring resumes one to five stars per role and location — trained on ten years of the company's own hiring decisions. That training period's hires were predominantly male, and the models learned the history rather than the merit: they penalized the word "women's" and downgraded graduates of women's colleges, reading gender proxies as negative signal. This is the domain's defining mechanism in its clearest form. A screener trained on past hiring decisions does not learn who will succeed; it learns who was hired, and reproduces the selection function embedded in that history as if it were a prediction.

What makes the case a reference is what the team did next, and where it hit a wall. They found the bias, and they patched the identified terms. Then they reached the correct and important conclusion: term-level fixes could not guarantee neutrality against unknown proxies, because the model had learned the pattern, not the words. Remove "women's" and the model can still infer the same signal from a dozen correlates no one has named. That is the documented ceiling on the patch lever — it addresses the proxies you can see, and the learned correlation routes around them. The lock-in this exposes is measured elsewhere in the field: screeners trained on past hires raise hire rates but replicate historical selection, and only a screener that values exploring candidates the history under-selected breaks the loop.

The company scrapped the project and disbanded the team around 2017. Per the company, recruiters saw the tool's recommendations but it was never used as a sole ranking. The honest reading is that abandonment was itself a governance outcome — taken because a fix could not be guaranteed, and taken before any external harm was documented, which distinguishes this arc from the domain's adjudicated cases. The record is investigative, reported through interviews with team members rather than a company publication, and that provenance is part of the honest structure: the clearest account of an organization detecting its own model's bias comes from reporting, not disclosure.

The sociotechnical reading

This is the case that names the domain's core mechanism and the ceiling on its most common fix. The mechanism: a screener trained on an organization's past hiring decisions imports the past's selection function. The model was not malfunctioning when it penalized "women's" — it was doing exactly what training on ten years of majority-male hires asks a model to do, which is to reproduce that history as prediction. The governable insight is that "trained on our own hiring data" is not a neutral technical detail; it is a decision to encode the past's biases as the future's filter, and the harder that history's selection was skewed, the more faithfully the model reproduces the skew.

The second insight is the ceiling on the patch. The team did the responsible thing — found the bias, removed the offending terms — and then recognized what patching cannot do: a model that learned a pattern can re-derive it from proxies no one has enumerated, so removing named terms does not remove the learned correlation. The lever that actually breaks the loop is not a patch but a design change — a screener that values exploring the candidates the history under-selected rather than only exploiting the pattern of who was hired before — and that is a different and harder thing than scrubbing words. The third insight is that abandonment is a legitimate governance outcome. The team could not guarantee neutrality, so the project was scrapped; that is not a failure of governance but an instance of it, and notably it happened before external harm was documented, which is the honest end of the ladder this domain will otherwise show shielded or litigated. The map's instruction is that training data is a selection function, term-patching has a ceiling, and "we could not make it fair, so we stopped" is sometimes the correct answer. The honest boundary throughout: no candidate outcome is modeled on the Lab diagram. Applicants are boundary-only; scores, patches, and the shutdown decision are institutional signals, and the gender-proxy finding and the abandonment live in the case file, never on any network.

The concepts used in this reading are defined in the Field Guide; the governance responses live in the Practice Library. The model organization for this case can be stress-tested in the PAN Lab.

Grounding sources for this case

The same sources that ground this model organization in the PAN library: evaluations, government documents, investigative reporting, and advocacy documentation, each labeled by tier.

dastin2018GroundingInvestigativeSave

Dastin, J. (2018, October 10). Amazon scraps secret AI recruiting tool that showed bias against women. Reuters. https://www.euronews.com/business/2018/10/10/amazon-scraps-secret-ai-recruiting-tool-that-showed-bias-against-women

https://www.euronews.com/business/2018/10/10/amazon-scraps-secret-ai-recruiting-tool-that-showed-bias-against-women

Appears in: PAN framework development

Grounds: domain grounding: hiring and employment screening (resume screening, interview scoring, ATS); model org: amazon_resume_engine

li2020aGroundingPeer-reviewedSave

Li, D., Raymond, L.R., & Bergman, P. (2020). Hiring as Exploration. NBER Working Paper 27736. https://doi.org/10.3386/w27736 https://www.nber.org/papers/w27736

doi.org/10.3386/w27736

Appears in: PAN framework development

Grounds: domain grounding: hiring and employment screening (resume screening, interview scoring, ATS)

Seeing your organization in this case file?

The histories here are documented after the harm. Mapping a live deployment's pathways and pressures, before the incident report, is engagement work: intake, diagnosis, prescription, and monitoring, with every limitation stated.

Sources & Evidence

Claims made on this page and what supports them. The full registry lives in Evidence.

EmpiricalAn internal team built an experimental recruiting engine — roughly 500 models scoring resumes one to five star…

An internal team built an experimental recruiting engine — roughly 500 models scoring resumes one to five stars per role and location — trained on ten years of the company's own hiring decisions, a period whose hires were predominantly male. The models learned that history: they penalized the word 'women's' and downgraded graduates of women's colleges, reading gender proxies as negative signal. The team patched the identified terms but concluded that term-level fixes could not guarantee neutrality against unknown proxies, because the model had learned the pattern rather than the words, and the company scrapped the project around 2017; per the company, recruiters saw the tool's recommendations but it was never used as a sole ranking.

dastin2018GroundingInvestigativeSave

Dastin, J. (2018, October 10). Amazon scraps secret AI recruiting tool that showed bias against women. Reuters. https://www.euronews.com/business/2018/10/10/amazon-scraps-secret-ai-recruiting-tool-that-showed-bias-against-women

https://www.euronews.com/business/2018/10/10/amazon-scraps-secret-ai-recruiting-tool-that-showed-bias-against-women

Appears in: PAN framework development

Grounds: domain grounding: hiring and employment screening (resume screening, interview scoring, ATS); model org: amazon_resume_engine

EmpiricalTraining a screener on an organization's past hiring decisions imports the past's selection function: research…

Training a screener on an organization's past hiring decisions imports the past's selection function: research on hiring as exploration finds that models trained on prior hires raise hire rates but replicate historical selection, and that a screener which values exploration rather than only exploitation breaks that lock-in loop. Two governance lessons follow — the patch lever has a documented ceiling, since removing named proxies does not remove a learned correlation, and abandonment can itself be a governance outcome, taken here before any external harm was documented rather than after an adjudication.

li2020aGroundingPeer-reviewedSave

Li, D., Raymond, L.R., & Bergman, P. (2020). Hiring as Exploration. NBER Working Paper 27736. https://doi.org/10.3386/w27736 https://www.nber.org/papers/w27736

doi.org/10.3386/w27736

Appears in: PAN framework development

Grounds: domain grounding: hiring and employment screening (resume screening, interview scoring, ATS)