# Which AI Hiring Tools Need a Bias Audit: NYC Local Law 144 Rules for 2026

Sarah Johnson · October 11, 2026

> NYC Local Law 144 bias audit rules for 2026: which AI hiring tools need an audit, how to classify Path A vs Path B, and what reviewers will test.

The comparison resolves to a single decision rule. Ask whether the tool's output materially shapes the hiring outcome for a covered position. If yes, Path B, documented before go-live. If the tool only narrows a pile that a human then reviews independently, Path A as an intake screen plus an internal check may be enough — but write down why you classified it that way, because the classification is the thing a reviewer will test. Compliance leaders need the resources to implement and maintain a program that holds up, and the audit path is where that resource gap shows up first. Before choosing, confirm the auditor's independence in writing, confirm the method is stated in the report, and confirm the audit is dated before the tool's first live use.

One practical check before you choose: confirm the auditor's independence in writing, confirm the method is stated in the report, and confirm the audit is dated before the tool's first live use. If any of those three is missing, you have a Path A or Path C record wearing a Path B label — and that is the record that fails.

![glass and steel Manhattan office tower dusk warm interior lights](https://static.mm-ais.com/article-images-ai/which-ai-hiring-tools-need-a-bias-audit-ai-4d538e1e.jpg)
glass and steel Manhattan office tower dusk warm interior lights

## Costs and Numbers That Matter

Budget per tool, not per vendor. A single vendor can ship several covered automated employment decision tools, and each one that substantially assists or replaces a human decision about a NYC-based candidate or employee carries its own audit obligation. Independent third-party bias audits for a single AEDT typically run in the five-figure range per cycle. If your vendor's platform bundles a resume ranker, a video-interview scorer, and a chat-based pre-screener, that is three audit line items, not one — even though one contract, one invoice, and one account manager sit behind all three. The practical check: pull your vendor list, then pull the actual decision points each product touches, and count the tools, not the logos.

The cost of being wrong runs the other way. An unaudited tool that is later found to discriminate exposes the employer to back-pay, injunctive relief, and reputational damage that dwarfs the audit fee. Compliance failures generally surface as penalties, lawsuits, and reputational damage, per BizBot's legal-compliance overview — and the remediation costs land on the employer, not the vendor. That asymmetry is the whole argument for treating the audit as a cost of deployment rather than a discretionary add-on. A five-figure audit is a known, bounded, budgetable number. Back-pay and injunctive relief are neither bounded nor scheduled.

Set your audit cadence to at least annual, and re-audit immediately after any material change. Material change is the operative phrase, and it is broader than a new vendor. It covers a model update pushed by the vendor, a change in the tool's role in your workflow, a new job category the tool now screens, and any shift in the population the tool scores. The mechanism is straightforward: an audit is a snapshot of a tool's impact ratios at a point in time, and a model update invalidates the snapshot. If you cannot say when the last audit ran and what changed since, you cannot demonstrate compliance.

Documentation is the second cost center, and it is cheaper than the first. Effective decision audit trails capture timestamped notes and meeting evidence to demonstrate governance, controls, and regulatory compliance, per workmate.com's guidance on audit trails. For each covered tool, that means a dated record of the audit, the auditor's identity and independence, the version of the tool audited, and the date the tool went live relative to the audit. The rule: the audit record must predate deployment, not follow it. Tie every audit record to a specific model version and date, and treat any subsequent retraining, weight change, feature addition, or vendor-side update as invalidating that record until re-audited.

Run the arithmetic before you sign the vendor contract. Multiply the number of covered tools by the per-cycle audit cost, then add internal staff time for documentation and re-audit triggers. That total is your compliance floor for the year. If the number is uncomfortable, the correct response is to reduce the number of covered tools in production — not to reduce the number of audits.

![Costs and Numbers That Matter — Which AI Hiring Tools Need a](https://static.mm-ais.com/article-images-ai/which-ai-hiring-tools-need-a-bias-audit-ai-04a176aa.jpg)

## What the Evidence Does NOT Establish

An audit is a snapshot, not a warranty. A clean result dated Monday does not survive a model update on Tuesday, and it says nothing about the version your team is actually running today. The practical consequence: tie every audit record to a specific model version and a specific date, and treat any subsequent retraining, weight change, feature addition, or vendor-side update as invalidating that record until re-audited. If your vendor will not tell you when the underlying model changes, that silence is itself the finding — you cannot maintain a defensible record against a moving target you cannot see.

The first edge case where the rule breaks is sourcing-only tooling. A tool that surfaces candidates but assigns no score, produces no ranking, and triggers no rejection generally falls outside the AEDT trigger, because nothing in its output displaces or shapes a human decision. The moment it starts filtering — a knockout question, a similarity threshold, a "top matches" ordering that recruiters work down — it flips into scope. The check is not what the vendor's contract calls the feature; it is whether a reasonable reviewer would treat the output as a screen. Run that test on each feature separately, because a single platform can be out of scope in one module and covered in another.

The second edge case cuts the other way, and it is the one that catches multi-market employers. A tool that scores only non-NYC candidates sits outside Local Law 144's reach. But if the same tool scores a mixed pool, or if a NYC-based candidate's data flows into the same model, the audit obligation attaches to that use. Segment by candidate location at the point of scoring, not at the point of sourcing, and document which positions are NYC-covered before the tool goes live.

What the evidence does not establish is equally important. The available compliance research documents a gap between deployment and independent audit; it does not establish that any particular tool is biased, that a given vendor is non-compliant, or that an audit will surface a defect. An audit is a disclosure and documentation exercise, not a certification of fairness. Treat a passing result as evidence that you ran the required process on a defined version — nothing more.

Two limits follow for your file. First, an audit cannot cure a tool that was never in scope, and it cannot retroactively cover a period before the record's date. Second, no audit substitutes for the underlying decision trail: timestamped notes and meeting evidence showing who reviewed what, and when, are what demonstrate governance when a decision is later questioned. Build the audit record and the decision trail together, before go-live, and version both.

![What the Evidence Does NOT Establish — Which AI Hiring Tools Need a](https://static.mm-ais.com/article-images-pixabay/which-ai-hiring-tools-need-a-bias-audit-1844b956.jpg)

## Auditing One Screener

Take one live requisition: a resume-ranking tool scores 1,200 applicants for a single NYC-based role. The vendor's contract says "advisory only," and the recruiter's workflow tells a different story — the bottom 60% of scored applicants are auto-archived before any human opens a file. That auto-archive is the dispositive act. A reasonable reviewer would treat the score as the decision, because for 720 of those 1,200 candidates, no human ever reviews anything else. The tool substantially assists or replaces the human screen, so the audit obligation attaches to the tool, not to the label on the contract.

Checkpoint 1 — scope. Confirm the tool issues a score that is used to reject, not merely to inform. The test is observational: pull the workflow configuration and the applicant-tracking-system rules, then trace what happens to a low-scoring record. Here, the observed answer is yes — auto-archive fires at the 40th percentile, and the archived records receive no human review. Document the percentile cutoff, the field the score writes to, and the rule that triggers archival. If the score gates a stage a candidate cannot pass without it, the tool is in scope regardless of vendor language.

Checkpoint 2 — audit status. Request the most recent independent bias audit for this specific tool, not the vendor's platform-wide marketing summary. The observed answer here: none in the past 12 months. That gap is the finding. Record the request date, the response, and the absence itself. An unanswered request is evidence; a verbal assurance is not.

Checkpoint 3 — remediation. With no audit on file, the fix is sequenced, not optional. First, suspend the auto-archive rule so no candidate is rejected on score alone while the tool is unaudited — route the bottom band to human review instead. Second, commission an independent audit covering the scoring logic and its outcomes by protected class. Third, before the tool goes live again in a decisioning role, re-run this same three-checkpoint review and file it.

| Checkpoint | Question | Observed | Artifact to file |
| --- | --- | --- | --- |
| 1 — Scope | Does the score reject? | Yes; auto-archive at 40th percentile | ATS rule + workflow trace |
| 2 — Audit status | Last independent audit? | None in past 12 months | Dated request and response |
| 3 — Remediation | What changes before relaunch? | Suspend auto-archive; commission audit | Signed remediation memo |

The order matters more than the paperwork. Scope first, because an out-of-scope tool needs no audit and an in-scope tool needs one before it screens anyone. Audit status second, because the missing document is the violation. Remediation third, because the corrective step must land before the tool resumes decisioning. Run all three on the single screener in front of you, file the artifacts, and repeat for the next tool — the mandate is per-tool, and so is the audit.

![Auditing One Screener — Which AI Hiring Tools Need a](https://static.mm-ais.com/article-images-pixabay/which-ai-hiring-tools-need-a-bias-audit-0e8e57b0.jpg)

## Worked Example: Run the Numbers

Consider a concrete case. A New York City employer runs a resume-screening tool for a mid-level analyst opening posted in early 2026. The tool ingests 400 applications and returns a ranked shortlist of 40, which a recruiter then reviews. The vendor calls it a "productivity aid." The recruiter treats the ranking as the first cut. Under the functional trigger, that ranking is the decision point — the tool substantially assists the human screen, so the audit obligation attaches to the tool, not the label.

Now run the numbers the way an auditor would. Inputs: 400 applicants, 40 shortlisted, one covered position, one tool version, one deployment date. Step one: confirm the tool is an AEDT by asking whether a reasonable reviewer would treat its output as dispositive. A ranked shortlist that eliminates 360 candidates before human review qualifies. Step two: identify the audit path. Step three: document the audit before go-live, not after. The illustration below shows the arithmetic; treat every figure as an illustration, not a benchmark.

| Step | Input | Calculation | Result |
| --- | --- | --- | --- |
| Applicants screened | 400 | — | 400 |
| Shortlist size | 40 | — | 40 |
| Screen-out rate | 400, 40 | (400 − 40) ÷ 400 | 90% |
| Human review share | 40, 400 | 40 ÷ 400 | 10% |
| Audit trigger | 90% screen-out | Dispositive screen? | Yes |

The winner in this example is the audit-before-go-live path. The break-even trigger is the point at which the tool's output stops being advisory and starts being dispositive: when the human reviewer's practical choice is to accept the ranking or restart the process, the tool has crossed from assistance into replacement. In this example, a 90% screen-out rate with only 10% of applicants reaching human eyes is the trigger. If the shortlist had been 200 of 400 — a 50% screen-out — a reasonable reviewer could still treat the ranking as one input among many, and the audit case weakens. Document the audit before the tool goes live: capture the inputs, the calculation, the trigger determination, and the reviewer's name attached to the go/no-go call, so the process is provable rather than merely defensible.

Document the audit before the tool goes live. A timestamped decision audit trail — notes, evidence, and the reviewer's name attached to the go/no-go call — is what converts a defensible process into a provable one. The mechanism is simple: capture the inputs, the calculation, the trigger determination, and the date, then store it where a regulator or plaintiff can retrieve it. Without that record, the employer cannot show the audit happened at all.

One caution on arithmetic: recompute every ratio yourself before you rely on it. A 40-of-400 shortlist is 10%, and the screen-out is 90%; those two figures must sum to 100%. If your own tool's numbers do not reconcile, the audit trail is already compromised. Run the numbers, label them as illustrations, and declare the winner and the break-even trigger in writing before deployment.

![Worked Example: Run the Numbers — Which AI Hiring Tools Need a](https://static.mm-ais.com/article-images-pixabay/which-ai-hiring-tools-need-a-bias-audit-9f60ff73.jpg)

## Decision Rules for 2026 Compliance

Local Law 144's reach has already been established: the "substantially assist or replace" trigger controls coverage, not vendor labels. What follows here is the operational layer — four if/then rules that turn that trigger into a 2026 compliance workflow you can actually run.

**Rule 1: Dispositive output, audit before go-live.** If a tool's output can cause a rejection, ranking, or score that a human reviewer treats as final, treat it as a covered automated employment decision tool and complete the audit before the tool touches a single NYC candidate. The check is simple: hand the tool's output to a reasonable reviewer and ask whether they would override it without strong cause. If the honest answer is no, the output is dispositive, and the audit clock starts at deployment — not at your convenience. As compliance practitioners note, building legal requirements into the decision early is far cheaper than retrofitting them after enforcement (BizBot, Legal Compliance in Company Management).

**Rule 2: Vendor audits are screening inputs, not compliance records.** If the only audit you can produce is one your vendor supplied, commission an independent audit for any tool that materially shapes outcomes. A vendor's own report tells you the vendor tested itself; it does not establish your compliance posture for a NYC-covered position. Use the vendor audit to shortlist tools, then commission an independent one before go-live and keep it as your record.

**Rule 3: Change triggers re-audit, immediately.** If the tool's training data, model version, or job-family scope changes, re-audit at once. A prior clean audit does not carry forward across a changed instrument — the bias profile of a model is a property of that model version and that data, not of the product name. Practical check: require vendors to give advance notice of model updates in the contract, and log every change with a timestamped entry so your audit trail shows which version each audit covered (Workmate, Decision Audit Trails).

**Rule 4: Document the decision, not just the result.** If you conclude a tool is *not* covered — because its output is advisory rather than dispositive — write that determination down: who reviewed the output, what a reviewer would plausibly override, and why the tool falls outside the trigger. Regulators and plaintiffs alike will ask why an unaudited tool was making ranking calls; an undocumented judgment is indistinguishable from no judgment.

Run all four rules as a standing checklist per tool: coverage determination, pre-deployment independent audit, change-triggered re-audit, and a written record at each step. Automation can help track these steps in real time rather than relying on manual spreadsheets (UMA Technology, Automated Compliance), but the rules themselves are judgment calls a human owner must sign. The 2026 question is not whether you use AI — it is whether your documentation would survive the question of who was really deciding.

## What to do next

| Step | Action | Why it matters |
| --- | --- | --- |
| 1 | For each AI hiring tool in your stack, write down the exact output it produces for a NYC-covered position — a screen, a ranking, or a score — and ask whether a reasonable reviewer would treat that output as dispositive. | NYC Local Law 144's trigger is the tool's substantial assistance or replacement of a human decision about a NYC-based candidate or employee, not whether AI is used at all. |
| 2 | Separate tools that materially shape the hiring decision from tools that merely support logistics, and place only the former on your audit list. | The 2026 compliance question is whether your tool's output materially shapes the hiring decision, not whether you use AI. |
| 3 | Apply the reasonable-reviewer test to each remaining tool: dispositive screen, ranking, or score for a NYC-covered position. | Audit every AI hiring tool whose output a reasonable reviewer would treat as a dispositive screen, ranking, or score for a NYC-covered position. |
| 4 | Complete the bias audit and document it before the tool goes live — not after — for every tool that clears the reasonable-reviewer test. | Documenting the audit before go-live is the canonical rule; retroactive documentation does not satisfy the mandate. |
| 5 | Re-check each tool's status whenever its output changes in a way that could shift it from supporting to dispositive for a NYC-covered position. | The trigger is the tool's substantial assistance or replacement of a human decision, so a change in output can move a tool onto the audit list. |
| 6 | Keep the audit record tied to the specific NYC-covered position and the tool's output, so the documentation shows what was assessed and when. | Assess whether the tool's output materially shapes the hiring decision for a NYC-based candidate or employee — and document that assessment before the tool goes live. |

## Frequently Asked Questions

**What is the deciding factor for whether a tool needs the full bias audit or just an internal check?**

Ask whether the tool's output materially shapes the hiring outcome for a covered position: if yes, Path B documented before go-live is required, but if it only narrows a pile that a human then reviews independently, Path A as an intake screen plus an internal check may be enough.

**Can I classify a screening tool as Path A without documenting anything?**

No — you must write down why you classified it that way, because the classification is the thing a reviewer will test.

**What three things must I verify before accepting an independent bias audit?**

Confirm the auditor's independence in writing, confirm the method is stated in the report, and confirm the audit is dated before the tool's first live use.

**What happens if an audit is missing one of those three elements — does it still count as valid Path B documentation?**

If any of those three is missing, you have a Path A or Path C record wearing a Path B label — and that is the record that fails.

**Should I budget for bias audits per vendor or per tool?**

Budget per tool, not per vendor, because a single vendor can ship several covered automated employment decision tools, and each one that substantially assists or replaces a human decision about a NYC-based candidate or employee carries its own audit obligation.

**Where will the resource gap in a compliance program show up first?**

The audit path is where that resource gap shows up first.

## Quick answers

| What is the single decision rule for determining which audit path an AI hiring tool requires? | Ask whether the tool's output materially shapes the hiring outcome for a covered position — if yes, Path B documented before go-live; if it only narrows a pile a human then reviews independently, Path A may be enough. |
| --- | --- |
| Why must a company write down why it classified a tool the way it did? | Because the classification is the thing a reviewer will test. |
| What three things should be confirmed before choosing an auditor? | Confirm the auditor's independence in writing, confirm the method is stated in the report, and confirm the audit is dated before the tool's first live use. |
| What is the consequence if any of those three confirmations is missing? | You have a Path A or Path C record wearing a Path B label — and that is the record that fails. |
| How should compliance leaders budget for bias audits? | Budget per tool, not per vendor, because each covered automated employment decision tool that substantially assists or replaces a human decision about a NYC-based candidate or employee carries its own audit obligation. |

Also worth reading: **NYC Local Law 144 Bias Audit Costs: What $1.5K Buys in 2026**: [NYC Local Law 144 Bias](https://ailaborbrain.com/blog/nyc-local-law-144-bias-audit-costs-what-15k-buys-in-2026.php) · **NYC Local Law 144: $500/Day Fines, Point-in-Time Audits**: [NYC Local Law 144: $500/Day](https://ailaborbrain.com/blog/nyc-local-law-144-500day-fines-point-in-time-audits.php) · **NYC Law 144 Impact Ratios: Keep or Retire Your AI Screener**: [NYC Law 144 Impact Ratios:](https://ailaborbrain.com/blog/nyc-law-144-impact-ratios-keep-or-retire-your-ai-screener.php)

### Related reading

- [San Francisco hiring law: 2026 $27,400 audit vs Hybrid Lite](https://ailaborbrain.com/blog/san-francisco-hiring-law-2026-27400-audit-vs-hybrid-lite.php)
- [Hiring bias audits: Threshold in 12 days vs retrain in 9 weeks](https://ailaborbrain.com/blog/hiring-bias-audits-threshold-in-12-days-vs-retrain-in-9-weeks.php)
- [Who Gets the Overtime Shift: 2026 Fair Workweek Gap — 24% Audit or Trust](https://ailaborbrain.com/blog/who-gets-the-overtime-shift-2026-fair-workweek-gap-24-audit-or-trust.php)
- [EU AI Act 2026: documenting algorithmic hiring-system controls](https://ailaborbrain.com/blog/eu-ai-act-2026-documenting-algorithmic-hiring-system-controls.php)
- [Are hiring screeners biased: 768 dimensions vs blinded human review](https://ailaborbrain.com/blog/are-hiring-screeners-biased-768-dimensions-vs-blinded-human-review.php)
- [Job Automation Analysis: Map Tasks, 57% Augmentation vs. 43% Automation](https://ailaborbrain.com/blog/job-automation-analysis-map-tasks-57-augmentation-vs-43-automation.php)

### Latest

- [Who Gets the Overtime Shift: 2026 Fair Workweek Gap — 24% Audit or Trust](https://ailaborbrain.com/blog/who-gets-the-overtime-shift-2026-fair-workweek-gap-24-audit-or-trust.php)
- [EU AI Act 2026: documenting algorithmic hiring-system controls](https://ailaborbrain.com/blog/eu-ai-act-2026-documenting-algorithmic-hiring-system-controls.php)
- [Are hiring screeners biased: 768 dimensions vs blinded human review](https://ailaborbrain.com/blog/are-hiring-screeners-biased-768-dimensions-vs-blinded-human-review.php)

Canonical: https://ailaborbrain.com/blog/which-ai-hiring-tools-need-a-bias-audit-nyc-local-law-144-rules-for-2026.php
Markdown: https://ailaborbrain.com/blog/which-ai-hiring-tools-need-a-bias-audit-nyc-local-law-144-rules-for-2026.php/index.md
