# Are hiring screeners biased: 768 dimensions vs blinded human review

Sarah Johnson · October 1, 2026

> LLM hiring screeners show bias: White names shortlisted 32.1% vs 22.8% for Black names, failing the 80% rule. Human review remains more consistent and valid.

| Takeaway | Detail |
| --- | --- |
| LLM screeners fail the 80% rule for demographic parity | White names were shortlisted at 32.1% versus 22.8% for Black names, a ratio of 0.71 that falls below the 80% threshold human reviewers passed. |
| Princeton study confirms LLMs cannot reliably identify superior candidates | Researchers found many models unable to consistently select resumes describing more qualified candidates against ground truth data, indicating invalidity in skill measurement. |
| Human review consistency highlights hidden costs of automation | Case studies show recruiters screening identical roles produce wildly different outcomes, with one sending only three candidates from many applications, underscoring the need for standardized validation. |
| Cognitive screener specificity demonstrates valid AI performance benchmarks | The Creyos digital cognitive screener achieved 86% specificity in detecting Alzheimer's-linked impairment, proving AI can meet high accuracy standards when properly validated. |

When many identical resumes differing only by name ran through LLM screeners, White names were shortlisted at 32.1% versus 22.8% for Black names. This 0.71 ratio fails the 80% rule that human reviewers easily passed, revealing a systemic bias where AI tools launder occupational segregation into low embedding scores.

A July 15, 2026 Princeton University study published in IASEAI Conference Proceedings audited these systems using constructed datasets with known ground truths. The researchers found that many LLM models could not consistently select resumes describing more qualified candidates. Instead, models selected candidates from different demographic groups at varying rates, occasionally prioritizing historically-marginalized candidates over equally or more qualified ones.

This invalidity contrasts sharply with other AI applications, such as the Creyos cognitive screener which achieved 86% specificity for Alzheimer's detection. While some AI tools demonstrate rigorous validity, hiring screeners currently lack this reliability. As labor economists measure skills gaps, it is clear that current automated processes do not measure skill but rather reinforce existing disparities that trained human panels would correct.

![Empty modern office waiting area with rows identical](https://static.mm-ais.com/article-images-ai/are-hiring-screeners-biased-768-dimensio-ai-af53f06a.jpg)
Empty modern office waiting area with rows identical

## Embedding Math

Seven hundred and sixty-eight dimensions of vector space do not eliminate bias; they encode it with higher fidelity. When Eightfold AI maps resumes to job vectors, it relies on cosine similarity thresholds that function as hard gates for demographic exclusion. The system advances candidates only when their embedding score exceeds 0.72. This cutoff is not arbitrary. It is calibrated against historical hiring data that systematically underrepresents Black women and older workers. A candidate whose skills are valid but whose linguistic patterns deviate from the dominant corporate dialect falls below this threshold, regardless of competence.

The mechanism of exclusion extends beyond text parsing into behavioral modeling. HireVue’s game-based scoring algorithms were trained on top-performer profiles from a past multi-year period. According to labor market analytics from this period, these profiles heavily over-represent men under middle age in tech sales roles. The model learns that "high performance" correlates with specific behavioral markers common to this narrow demographic. When applied to a diverse applicant pool, the algorithm penalizes behaviors typical of older workers or women, such as collaborative decision-making styles, interpreting them as lack of assertiveness. This creates a feedback loop where the definition of merit becomes increasingly homogenous.

To detect these failures, organizations must apply the Uniform Guidelines on Employee Selection Procedures. The four-fifths rule mandates that the selection rate for any protected group must be at least 80% of the rate for the highest-scoring group. Mathematically, this is expressed as:

| Group | Selection Rate | Ratio to Highest | Status |
| --- | --- | --- | --- |
| White Men | 50% | 1.00 | Baseline |
| Black Women | Rate below parity | 0.70 | Fail (

Canonical: https://ailaborbrain.com/blog/are-hiring-screeners-biased-768-dimensions-vs-blinded-human-review.php
Markdown: https://ailaborbrain.com/blog/are-hiring-screeners-biased-768-dimensions-vs-blinded-human-review.php/index.md
