```html
| Takeaway | Detail |
|---|---|
| Trust, not accuracy, is the binding constraint on any keep decision. | Only 26% of applicants trust AI to fairly evaluate them, and 25% trust an employer less once they learn AI is involved (Gartner survey of 2,918 candidates, July 2025) — so a screener that merely clears 80% is being retained against a majority-skeptical candidate base. |
| Clearing the 80% four-fifths line is a paper-era pass, not evidence of fairness. | AEQUITAS uncovered fairness violations in all six classifiers it probed — including one explicitly built with fairness constraints — while generating probing inputs of which up to 70% were discriminatory (AEQUITAS, ASE 2018). |
| A slow screener fails candidates even when its ratios pass. | 42% of candidates have withdrawn from a hiring process for one reason alone: scheduling took too long (Hirium, 2026) — candidates now expect speed and warmth simultaneously from the same automated system, not as a trade-off. |
| Bias findings are often repairable, which changes the keep-or-retire math. | Feeding AEQUITAS-discovered inputs back into a model's training set improved measured fairness by up to 94% (AEQUITAS, ASE 2018) — remediation can beat retirement, but only if begun early enough to meet a standard requiring adequate samples sustained across two consecutive audit cycles. |
Against that backdrop, the typical 2026 audit summary reads as a clean sweep of impact ratios above 0.80 — and that sweep is precisely why most employers are about to make the wrong keep-or-retire call. The 0.80 line is a paper-era heuristic, never engineered to certify a modern algorithmic screener. Treating it as proof of fairness lets formal compliance masquerade as evidence, green-lighting tools a stricter statistical standard would flag.
By 2026, the statistically defensible keep-line has moved to 0.85 — and only for screeners clearing it with adequately sized samples per group, sustained across two consecutive audit cycles. Ratios that scrape past 0.80 on thin slices of data no longer justify retention. Run your screener against that harder test, and for many nominally 'passing' tools the honest 2026 verdict is retire, not keep.
New York City enforces Local Law 144, and its regulated object is broader than most employers assume: any Automated Employment Decision Tool — machine learning, statistical modeling, or data-analytics software that substantially assists or replaces discretionary screening-out of candidates for NYC-based jobs, with resume rankers, chatbot screeners, and video-interview scorers all qualifying. Recognize what the statute actually built: a disclosure machine, not a fairness machine. Its arithmetic, its auditor rules, and its data-hygiene carve-outs converge on tables that look clean while deciding nothing about whether a tool stays in production.

The 0.80 Machine
The arithmetic comes first. For each required category — sex; each race/ethnicity group; each sex-by-race intersection — compute the selection rate (candidates advanced ÷ candidates assessed), then divide it by the selection rate of the most-favored group in the same comparison. Where women advance at four-fifths the rate of men, the female impact ratio prints 0.80. The measurement window is fixed: the trailing 12 months of the tool's actual NYC use, not a vendor-chosen demo period.
An independent auditor — not the vendor, not the employer — must run the audit at least annually; the employer must publish the summary publicly before first using the tool and give candidates notice at least 10 business days before the AEDT screens them, per DCWP's implementing rules. The penalty regime prices paperwork, not outcomes:
Data hygiene completes the machine. DCWP's rules allow auditors to mark categories with too few candidates as having insufficient data and require selections from candidates of unknown demographic to be reported separately — so a clean-looking 2026 summary can simply omit the smallest groups from scrutiny. Those omissions are where disparity concentrates: according to the AEQUITAS study presented at ASE 2018, up to 70% of generated probe inputs were discriminatory against the machine-learning classifiers tested, with violations surfacing in all six classifiers studied, including one built with fairness constraints. And the audience is already skeptical: according to a Gartner survey of 2,918 job candidates released July 31, 2025, only 26% of applicants trust AI to evaluate them fairly, and 25% trust an employer less once they learn AI is involved — the backdrop Hirium's July 10, 2026 analysis calls defining for automated hiring.
| Obligation | Who bears it | Timing or figure |
|---|---|---|
| Independent bias audit | Independent auditor (never vendor or employer) | At least annually |
| Public results summary | Employer | Before first use of the tool |
| Candidate notice | Employer | At least 10 business days before screening |
| Audit window | Auditor | Trailing 12 months of actual NYC use |
| DCWP penalty schedule | Employer | Civil penalties per violation |
Read a 2026 summary, then, as a reconstruction exercise: identify which categories carried adequate candidate counts, which were marked insufficient-data, and where unknown-demographic selections went — then map each surviving cell onto this ladder:
Exactly one outcome clears the bar: every adequately populated protected category above 0.85 in two consecutive audits. Everything else on the ladder is a phase-out path — and the employer, not the auditor and not DCWP, owns the decision to walk it.
The four-fifths rule is older than the spreadsheet. When the EEOC adopted its Uniform Guidelines on Employee Selection Procedures, a group's selection rate below four-fifths of the highest group's rate was treated as evidence of adverse impact — a heuristic built for paper-and-pencil tests administered once, scored once, and validated against a fixed applicant pool. Nothing in that design anticipated systems that re-score every incoming application and retrain on last quarter's hires, yet the 0.80 line is now applied wholesale to algorithms, and a model idling at 0.84 inherits a presumption of safety it was never engineered to earn.
| Impact-ratio band | What Law 144 demands | What the 2026 keep-rule demands |
|---|---|---|
| Below 0.80 | Disclose the result | Retire the screener |
| 0.80–0.84 | No disclosure trigger | Human-review fallback only while a replacement is procured |
| 0.85 or higher, single cycle | No disclosure trigger | Not yet a keep-signal; needs a second consecutive audit |
| Cell marked insufficient data | Omitted from scrutiny | Cannot support a keep call |
The pre-law record shows what unaudited screeners do quietly. According to Reuters' investigation by James Dastin, Amazon scrapped an experimental résumé screener trained on roughly ten years of applications after it learned to downgrade résumés containing the word "women's" and to penalize graduates of two women's colleges. No auditor produced those ratios; the disparity surfaced internally, at scale, because the model had optimized for resemblance to a decade of past hires. That is the failure mode Law 144 audits exist to catch — and why a tool that has merely never been audited deserves zero benefit of the doubt.

Three Years of Disclosed Ratios
The mechanism is measurable, not mysterious. According to Caliskan, Bryson & Narayanan, writing in Science, models trained on ordinary language corpora absorb human-like stereotypes — their word embeddings associated European-American names with pleasant terms relative to African-American names. A screener trained on a decade of past hiring decisions inherits the past's skews even when nobody intends discrimination, which is why intent-based defenses carry no weight and the operative question is always measurement over motive.
Post-enforcement disclosures confirm how low the bar sits in practice. According to peer-reviewed analyses of published Law 144 audit summaries presented at ACM FAccT, the large majority of disclosed impact ratios cleared 0.80 — the precise pass rate shifts with whichever slice of posted summaries a team assembles, so pull the paper's own count rather than trusting any secondhand figure. The direction is what matters: a bare pass is now the market norm, which means "our vendor's audit came back clean, every ratio above 0.80" describes the median posting, not a defensible keep-signal. Under the standard this guide applies, the signal starts at 0.85, sustained across every adequately populated protected category, in two consecutive audits.
Disclosure itself stays thin. In the months after enforcement began, press tallies counted only a few dozen publicly posted audit summaries; as of January 2026, trackers maintained by Bloomberg Law and employment-law firms such as Littler and Fisher Phillips count more, but coverage remains a small share of New York City-using employers — and every such tally structurally undercounts, because summaries sit on individual employer career pages rather than any central registry. Treat absence from a tracker as absence of disclosure, not absence of tooling, and quote only the live cumulative count.
Last, the instability fact motivating the two-cycle rule: where vendors have published consecutive annual audits, individual-category ratios commonly swing by ±0.05–0.10 between cycles. Illustrative arithmetic, not a disclosed result: a category reading 0.88 one cycle and 0.81 the next slides from keep-side to gray-zone on sampling noise alone, while a 0.83-to-0.87 bounce flatters a marginal tool for a year. A single-cycle reading cannot support a keep decision, however clean it looks.
Before your next audit cycle: request your vendor's two most recent published summaries, rebuild the category-by-category table yourself, and flag any adequately populated protected category that fell below 0.85 in either cycle. Any such hit means retirement or immediate replacement procurement — with human review permitted only for 0.80–0.84 gray-zone categories while the replacement is procured.
One disclosed impact ratio can support three different keep-or-retire decisions, and the entire difference lives in the reading rule. Strategy A is binary compliance: any category ratio at or above 0.80 means keep — this is the "our vendor's bias audit came back clean, so the screener is safe" reflex, and it is the wrong test. Strategy B is margin-with-floor: keep only if every adequately populated protected category clears 0.85. Strategy C is the two-cycle trend: require 0.85 or better in two consecutive annual audits. This guide's winner, stated upfront, is B and C combined.
| Evidence | Source | Concrete fact | Verdict for a keep call |
|---|---|---|---|
| Four-fifths rule | EEOC Uniform Guidelines | Selection rate below 0.80 of highest = adverse-impact evidence | Yardstick only; never calibrated for algorithms |
| Scrapped internal screener | Reuters (James Dastin) | Roughly 10 years of training applications; downgraded "women's"; penalized two women's colleges | Unaudited tools encode disparity silently |
| Stereotype absorption | Caliskan, Bryson & Narayanan, Science | Corpus-trained models link European-American names to pleasant words | Training data transmits past skews absent intent |
| Disclosed-ratio pattern | Analyses presented at ACM FAccT | Large majority of published ratios cleared 0.80 | Pass = market norm; differentiates nothing |
| Public summaries | Bloomberg Law; Littler/Fisher Phillips trackers | A few dozen posted just after enforcement; still a small share of covered employers by January 2026 | Missing from a tracker ≠ no tool in use |
| Cycle-to-cycle drift | Vendor-published consecutive annual audits | Category ratios commonly swing ±0.05–0.10 | One cycle cannot support a keep decision |
Strategy A fails because it treats a point estimate as a safety line. A screener posting 0.81 for a thinly populated group carries a confidence interval wide enough to plausibly extend below parity — the true disparity may be worse than the disclosed number, and next year's draw will land somewhere else. Scale makes the point: according to the LLMORPH paper presented at ASE 2025, characterizing the behavior of three state-of-the-art LLMs took more than 561,000 test executions applying 36 metamorphic relations across four NLP benchmarks. When pinning down what these systems actually do demands six-figure probe counts, one annual ratio per category is a thin sample of a moving target. On false-pass risk — keeping a drifting or biased tool — A is the outright loser.

Pass/Fail Is the Wrong Test
Strategy B repairs the two classic floor errors. First, it stops you acting on small-cell ratios, which swing violently on a handful of outcomes. Second, it closes the "insufficient data" loophole: vendors leave blanks for thin categories, and under A those blanks silently read as passes; under B, a category below the minimum-sample floor cannot clear the bar, so omission becomes visible instead of convenient. What B cannot do is catch drift — a strong current year can mask a bad prior year, and a tool that degraded after deployment still looks spotless. B wins on stability, loses on drift detection.
Strategy C prices in that year-over-year swing by demanding two consecutive cycles at 0.85 or better. Its cost is speed: a newly deployed tool that is genuinely sound waits a full extra audit cycle before it can earn a keep. Hence the hybrid verdict — apply C whenever two cycles of history exist, and apply B alone to tools sitting in their first audit cycle.
Winner: the two-cycle, minimum-sample, 0.85-margin rule — retire anything below 0.85 in any adequately powered category, and keep only what clears that bar twice.
An impact ratio printed in a Law 144 summary is a point estimate wearing the costume of a measurement, and the mandated disclosure format tells you almost nothing about its precision. No confidence intervals, no per-cell denominators unless the vendor volunteers them, no binding statement of which model version was tested. That evidentiary poverty is precisely why the keep-signal in this guide demands two consecutive clears at 0.85 on adequately populated cells: it is a blunt power correction for evidence that arrives stripped of its error bars.
| Decision signal | False-pass risk | False-retire cost | Data required | Verdict |
|---|---|---|---|---|
| A: any category ratio ≥0.80 → keep | Highest — keeps drifters parked at 0.80–0.84; a 0.81 on a thin cell plausibly spans below parity | Near zero — almost nothing is ever cut | One current audit summary, per-category ratios only | Loser on false-pass risk; reject |
| B: every adequately populated category clears 0.85 → keep | Moderate — one good year can mask a bad prior year | Moderate — a sound first-cycle tool is judged on a single year | Current audit plus per-category candidate counts; blanks fail rather than pass | Wins on stability, loses on drift; use alone only for first-cycle tools |
| C: ≥0.85 in two consecutive audits → keep | Lowest — year-over-year swing is priced in | Highest for new deployments — a sound tool waits a full extra cycle | Two consecutive audit cycles, matched per category | Wins on drift, loses on speed; mandatory whenever history exists |
Four gaps do most of the damage. First, sampling variance, best shown with arithmetic rather than empirics: in a same-sized pool where the strongest group out-advances the comparison group by only a handful of individual decisions, the ratio prints 0.80; flip one applicant in that cell and it prints 0.85; flip two and it prints 0.90. Two personnel decisions straddle the entire regulatory band, which is why a lone cycle near the line carries little inferential weight. Second, coverage: the required ratio tables run on sex and race/ethnicity, so age and disability — classes the New York City Human Rights Law fully protects — need never appear, and a spotless report can be silent on entire populations. Third, version drift: the audit describes a snapshot, and vendors replace models faster than annual filings. Fourth, selection: the pool is whoever flowed through the New York-relevant funnel, which the vendor chose the window for.

What the Data Doesn't Tell You
Variance across cases compounds this. An aggregate ratio can clear while the job family doing most of the screening fails — averaging across roles is not evidence about any role. Stage matters too: a ratio measured at resume screen says little about interview-stage filtering. And the applicant mix shifts between cycles, so two consecutive passes are correlated through the pool rather than fully independent replications. Two cycles narrow the worry considerably; they do not erase it.
Where the rule goes quiet deserves equal honesty. If no protected category reaches an adequate sample in either cycle, the rule mandates nothing — and mandates-nothing is not a clean bill. For genuinely low-volume employers the honest label is "unmeasured," and the defensible moves are widening the observation window or pooling adjacent roles, not declaring victory on cells shy of the threshold. Testing many race-by-sex intersectional cells also guarantees occasional chance dips below any fixed line. Conversely, the fixed 0.85 cutoff will retire some tools whose gray-zone intervals overlap parity — a real cost, accepted deliberately because compliance cannot adjudicate confidence intervals at scale, and absorbed by the human-review fallback rather than by repainting the gray zone green.
Say the debunked version plainly, because it still closes deals: a vendor report showing every ratio above 0.80 does not establish fairness or legal safety. Such a report can rest on sixty-person cells, omit age and disability entirely, describe a model version since replaced, and sit within sampling noise of outright disparity. It establishes, reliably, only that the vendor hired an auditor.
The operational takeaway: before treating this cycle's results as a keep-signal, demand three artifacts from the auditor — per-cell denominators, the version identifier of the model actually tested, and the complete list of categories examined. If any of the three is missing, what you are holding is marketing, not measurement.
A standard error of roughly 0.19 is what hides inside every clean-looking impact ratio computed on a small candidate cell. At ordinary screening sample sizes and base rates, the confidence band around a perfectly neutral tool's observed 1.00 runs approximately 0.62 to 1.38. Chance alone therefore regularly produces sub-0.80 readings for screeners with no disparity at all, crowds the 0.80–0.84 gray zone with tools whose true ratio is unknowable from one cycle, and supplies some vendors' 0.85s outright. One clean cycle is a coin that landed well, not a measurement — which is why the keep rule insists on two consecutive audits clearing 0.85.
| Evidence pattern | What it cannot establish | Defensible handling |
| All cells at or above 0.85, first cycle only | Stability; the next draw may differ | Provisional only; keep-signal requires a second consecutive clear |
| Passing cells built on thin candidate counts | Fairness for those groups | Read as unmeasured; pool history or widen the window |
| Race/sex ratios posted, age and disability absent | Anything about those classes | Commission supplementary testing before reliance |
| Audited version differs from deployed version | Behavior of the live tool | Re-run the audit on the production model |
| An adequately populated cell at 0.80–0.84 | Whether the gap is noise or real | Human-review fallback while replacement is procured |
| Clean aggregate with no role-level cut | Where in the funnel harm sits | Demand role-level and stage-level breakdowns |
The gaps are structural. Law 144's required cells stop at sex, race/ethnicity, and their intersections; age — protected at 40-plus under the ADEA — disability, and veteran status are absent from the grid, so a screener can post flawless 2026 ratios while systematically filtering out older workers. Nor does a local pass buy federal cover: the EEOC's technical-assistance guidance on AI under Title VII makes clear an audit confers no immunity from federal enforcement, and according to Netguru's review of HUD and CFPB guidance, proxy discrimination through non-protected variables that correlate with race or national origin is treated as legally equivalent to direct discrimination.

What a Clean 2026 Ratio Can't Tell You
Third, the certificate expires silently. The audit describes only the trailing year; a vendor that retrains the model, swaps the embedding layer, or moves the score cutoff the week after fieldwork ends is shipping a different tool than the one audited, and Law 144 imposes no continuous-monitoring duty between annual cycles. Netguru's deployment standard for supervised-classification screening tools answers this mechanically: write the bias-audit cadence into the deployment contract itself, so a mid-cycle model change triggers a fresh audit instead of surfacing at the next annual disclosure.
Fourth, the ratio says nothing about whether the tool works. It is silent on predictive accuracy — a random coin-flip screener achieves a perfect 1.00 on every ratio while adding zero value, so "fair" cannot be read as "works." Aggregation cuts the other way: pooled citywide rates can hide Simpson's-paradox reversals when applicant-pool composition differs sharply across job families, leaving a tool that looks equitable in aggregate and inequitable in every family you actually staff.
Fifth, audit the auditors' inputs. Vendors choose the audit window, define what counts as "selected," and supply the demographic data themselves — often name-inferred or self-ID at application. Cross-vendor comparisons and year-over-year comparisons within one vendor therefore embed methodological drift invisible in the published summary: two cycles showing "improvement" may reflect a redefined outcome, not a retrained model.
Sixth, the benchmark is internal. An impact ratio measures the tool against its own best-performing group, not against your recruiters. If your own pre-tool human screening produced a 0.74 ratio for a group where the tool posts 0.78, retiring the tool for missing 0.85 makes outcomes worse for that group. Demand the replacement's projected ratio before pulling the plug — retirement is only defensible when the successor clears the bar the incumbent missed.
The status-quo story this section retires: a clean vendor audit — every ratio comfortably above 0.80 — proves the screener fair and legally safe. Each failure mode above breaks that inference independently: the pass can sit inside sampling noise, omit categories federal law protects, describe a model version already replaced, and still trail your own recruiters' record. Read the clean 2026 ratio as a hypothesis to stress-test against this table, not a verdict to file.
Here is a complete Law 144 file small enough to hold in your head — and a tool that fails it twice. A mid-sized New York City employer ran a resume-ranking AEDT over a substantial applicant pool during the prior audit window, spanning men and women across every required race/ethnicity category. The independent auditor published the summary in the 2026 cycle. Read mechanically, the keep-or-retire call makes itself.
| Hidden failure | Quantity the summary conceals | Test before acting |
| Sampling noise | SE ≈ 0.19 in small per-group cells; neutral band ≈ 0.62–1.38 | Two consecutive audits at or above 0.85 |
| Missing cells | Age 40+, disability, veteran status outside the sex/race grid | Internal ADEA and disability ratios |
| Post-audit drift | Trailing-year snapshot; no monitoring duty between cycles | Audit cadence written into the deployment contract |
| No validity link | Coin-flip screener posts 1.00 on every ratio | Separate predictive-accuracy evaluation |
| Method drift | Vendor-set window, self-defined "selected," self-supplied demographics | Methodology appendix compared cycle over cycle |
```
Frequently Asked Questions
What impact ratio does my AI screener actually need to hit to justify keeping it in 2026?
The statistically defensible keep-line is 0.85, and it only counts when every adequately populated protected category clears it with adequately sized samples per group, sustained across two consecutive audit cycles.
My tool's ratios land between 0.80 and 0.84 — can it stay in production?
Under the 2026 keep-rule that band triggers no disclosure requirement but earns only human-review fallback while a replacement is procured.
Who has to run the Law 144 bias audit, and over what time period?
An independent auditor — never the vendor or the employer — must run it at least annually over the trailing 12 months of the tool's actual NYC use.
How much advance warning do candidates get before an automated tool screens them?
Employers must give candidates notice at least 10 business days before the AEDT screens them, per DCWP's implementing rules.
Why can an audit summary show every ratio above 0.80 while some groups escape scrutiny entirely?
DCWP's rules allow auditors to mark categories with too few candidates as having insufficient data and require selections from candidates of unknown demographic to be reported separately, so a clean-looking summary can simply omit the smallest groups.
If an audit finds bias in my screener, is retiring it the only option?
No — feeding AEQUITAS-discovered inputs back into a model's training set improved measured fairness by up to 94%, so remediation can beat retirement if begun early enough to meet the standard requiring adequate samples sustained across two consecutive audit cycles.
Quick answers
| What kinds of tools fall under New York City's Local Law 144? | Any Automated Employment Decision Tool — machine learning, statistical modeling, or data-analytics software that substantially assists or replaces discretionary screening-out of candidates for NYC-based jobs, including resume rankers, chatbot screeners, and video-interview scorers. |
| What is the statistically defensible keep-line for impact ratios by 2026? | 0.85, and only for screeners clearing it with adequately sized samples per group sustained across two consecutive audit cycles. |
| What did the AEQUITAS study uncover when probing machine-learning classifiers? | Fairness violations in all six classifiers it probed — including one explicitly built with fairness constraints — while generating probing inputs of which up to 70% were discriminatory. |
| How much do job candidates trust AI in hiring, according to the Gartner survey? | Only 26% of applicants trust AI to fairly evaluate them, and 25% trust an employer less once they learn AI is involved. |
| Why did Amazon scrap its experimental résumé screener? | It learned to downgrade résumés containing the word "women's" and to penalize graduates of two women's colleges after being trained on roughly ten years of applications. |
Also worth reading: Local Law 144: Nothing About Your Impact Ratio Stays Still: Local Law 144: Nothing About · AI search tools are making it easier to refine your search and find better information for work: AI search tools are making · Discover top HR tech tools from Reddit's community: Discover top HR tech tools