| Takeaway | Detail |
|---|---|
| The statutory fine is fixed and finite; the audit risk is open-ended. | Operating an automated employment decision tool without a completed bias audit draws fines of up to $500 per day, accruing for as long as the tool stays in service — a schedule that turns one unremediated resume screener into a six-figure annual exposure. |
| Bias audits fail on cell size long before they fail on intent. | Under the four-fifths rule, an impact ratio below the threshold signals adverse impact; ratios computed on intersectional cells too small to support them cannot survive recalculation by DCWP or a plaintiff's expert. |
| A clean-looking point estimate is not proof of compliance. | A reported impact ratio of 0.90 can span 1.00 once a confidence interval is drawn on a thin cell, which is why durable audits publish interval estimates and significance tests alongside every point estimate. |
| The cheapest audit converts directly into the most expensive remediation. | Reworking a collapsed audit routes through external consultants billing $150-$500 per hour, and reactive evidence assembly stretches preparation from hours into weeks or months compared with audit-ready organizations. |
At $500 per day — the fine Local Law 144 attaches to an automated employment decision tool that runs without a completed bias audit — a single unaudited resume screener accrues fresh exposure with every passing day for a year. The independent audit that eliminates that exposure costs a small fraction of the fine stream it prevents. On paper the trade is absurd, and the fine is still the smaller of the two risks.
The larger liability is the audit most buyers actually purchase. Cheap independent reviews share a structural flaw: impact ratios computed on intersectional cells too small to mean anything. Under the four-fifths rule, a ratio below the threshold signals adverse impact — and a ratio of 0.90 can look clean yet span 1.00 once a confidence interval is drawn on a thin cell. Those are the numbers that collapse under Department of Consumer and Worker Protection scrutiny, or under a disparate-impact claim.
Closing that gap is what this guide is for: how to buy an audit that survives contact with the regulator, not one that decorates a careers page. That means cell sizes large enough to support every reported ratio, confidence intervals instead of bare point estimates, methodology that withstands expert questioning, and evidence practices that keep the file honest between annual snapshots — because the system drifts daily while the audit checks a moment.

The Fine Machine
Run one unaudited resume screener against New York candidates for twelve months and the city bills you for every one of those days — because Local Law 144 meters non-compliance by the day. Each calendar day an Automated Employment Decision Tool touches a NYC candidate without a valid bias audit behind it is a distinct violation, priced at $500. Note the drafting choice: the Department of Consumer and Worker Protection (DCWP) wrote this as a flow penalty, not a stock penalty, so delay compounds linearly. The myth to retire here: a violation is not a one-time event you absorb and move past. It is a running tab.
The statutory core imposes three duties on any employer or employment agency using an AEDT to score, rank, or cull NYC candidates: complete an independent bias audit; publish a summary of the results on your website before the tool goes live; and give affected candidates at least 10 business days' advance notice plus access to an alternative selection process. Miss any one and the meter runs — but the audit duty is the one that gates the rest, because a lapsed audit invalidates the compliance story attached to everything downstream.
What trips the law is broader than most HR teams assume. An AEDT is software that substantially assists or replaces discretionary candidate decisions using machine learning, statistical modeling, or data analytics — squarely covering resume screeners, asynchronous video-interview scorers, chatbot qualifiers, and gamified assessments. Tools that merely aggregate, sort, or organize candidates without supplying independent decision-making inputs are exempt. The classification test: does the tool's output change who advances, or only how the pipeline is displayed? A chatbot that ranks applicants is in scope; one that schedules interviews is not.
Enforcement follows a fixed chain, and it rewards fast movers:
| Stage | What happens |
|---|---|
| Trigger | A candidate complaint or a DCWP inspection opens the file |
| Notice | DCWP issues a formal notice of violation |
| Cure window | You remediate during the cure period; penalties have not yet attached |
| Metering | Any violation persisting past cure is assessed as a civil penalty per day |
The metering rule: $500 per violation per day, with each day a fresh violation for as long as the tool stays in service. Enforcement went live after the city pushed back the original start date — more than two years of live operation by 2026, which is why "we didn't know" carries no residual value as a defense.
Now the clause this guide turns on. The audit must be performed by an independent auditor — or, where none is used, by at least two people, at least one of whom played no role in designing, building, or using the tool. Read it twice: nothing requires the auditor to sit outside your company, only that one pair of hands stayed uninvolved. A vendor can field two of its own employees — the modeler who built the screener plus a colleague who never touched it — and stamp the output "bias-audited." The economics are blunt: according to Zof AI, external consultants charge $150–$500/hour for audit-prep work alone, so an in-house audit protects the vendor's margin while transferring the entire per-day exposure onto you, the licensee.
| Posture | Cost signal | Your exposure | Verdict |
|---|---|---|---|
| No audit | Nothing spent upfront | $500/day accruing for as long as the tool runs | Lose |
| Vendor self-audit (two-person clause) | "Free" certificate bundled with the license | Full metering resumes if DCWP rejects the independence claim | Fragile |
| Listed third-party auditor | Prep support at $150–$500/hour (Zof AI) | The only posture that survives the independence test | Win |
The verdict column is the thesis in miniature: the cheap postures embed a fine stream nobody itemizes for you. Buy the independent audit from the city's published list and renew at month 11 so coverage never lapses past the one-year recency limit — and when a vendor hands you a certificate, read the signature block first. If both names draw a paycheck from the vendor, you are holding the loophole, not the audit.

The Compliance Ledger
Ninety-nine percent of Fortune 500 companies run applicant tracking software, and 88 percent of employers concede their filters likely reject qualified high-skill candidates. Joseph Fuller's team at Harvard Business School published both figures with Accenture in Hidden Workers: Untapped Talent (2021), and together they define the universe Local Law 144 now meters: automated screening is not an edge case in New York hiring — it is the default infrastructure. Any budget that treats the bias audit as an optional line item starts from the wrong denominator.
Start with what the audit actually computes, because it is less negotiable than vendors imply. LL144's impact-ratio reporting inherits the four-fifths rule from the EEOC's Uniform Guidelines on Employee Selection Procedures: a protected group's selection rate below the four-fifths share of the best-performing group's rate signals adverse impact. That yardstick predates the city's statute by decades, so the arithmetic inside your audit is fixed. Two buyers hiring different listed auditors are buying the same computation applied to different data — the open variables are tool coverage and record hygiene, not methodology. You are purchasing a known calculation, not an open-ended consulting engagement.
The line that tempts buyers to defer is enforcement — and the status-quo myth worth killing is "wait until they actually fine someone." Notices of violation issued by DCWP are obtainable through FOIL, and employment-law trackers at Littler and Fisher Phillips maintain running tallies of them. Across more than two years of records, observed penalties have lagged far behind the statute's theoretical daily exposure. Read that gap correctly: it is a sampling artifact, not a safety margin. Cure periods mean collected totals undercount accrued liability, and the denominator — how many AEDTs actually touch New York candidates — is published nowhere, so no honest violation rate exists. Price the audit off the accrual clock, never off observed fines.
Market structure settles the last line. DCWP's guidance pages host posted audit summaries, and law-firm compliance surveys track who files them; the posted population has consolidated decisively around external auditors, while the statute's internal-review fallback — permitted only where an independent auditor is unavailable — sits essentially unused. The reason is audience: an internal review may satisfy the letter of the fallback, but it clears neither procurement teams nor plaintiffs' counsel, so it buys nothing in the venues where these documents actually get read. Before shortlisting anyone, run the ten-minute check: pull the posted-summary list, count how many entries name an outside firm versus an internal title, and weight your shortlist toward auditors already appearing on it.
The cheapest document in your compliance folder is the certificate your screening vendor emailed you — and on this scorecard it earns a zero. Local Law 144's independence clause means an audit supplied by the AEDT maker counts only when an independent third party verifies it; standing alone, it fails the first column below no matter how polished the underlying statistics. That single row kills the most common shortcut buyers take in 2026: treating the tool maker's self-grading as a substitute for a listed auditor.
| Ledger line | Figure | Named source | Decision it forces |
| Regulated universe | 99% of Fortune 500 use applicant tracking software | Fuller et al., HBS & Accenture, Hidden Workers (2021) | Assume coverage; scope every screener |
| Admitted filter failures | 88% of employers say filters likely reject qualified high-skill candidates | Same HBS/Accenture report | Treat bias risk as baseline, not edge case |
| Adverse-impact trigger | Below the four-fifths share of the best-performing group's selection rate | EEOC Uniform Guidelines | Arithmetic is fixed; fix data instead |
| Compliance cost | Quote-based, set per audit cycle | Published vendor rate cards; procurement disclosures | Capped and budgetable |
| Non-compliance cost | Daily accrual for as long as the tool runs (see The Fine Machine) | LL144 penalty schedule via DCWP | Effectively uncapped |
| Observed collections | Lag far behind theoretical exposure across 2+ years of notices | Littler & Fisher Phillips tallies of FOIL-obtained DCWP notices | Never price off observed fines |
| Posted auditor type | External firms dominate; internal fallback dormant | DCWP posted-summary pages; law-firm compliance surveys | Buy external only |
The scorecard's five columns track what the city's rules actually scrutinize: who performed the audit, what cuts were reported, whether the math is inspectable, whether the auditor knows New York filings, and what the engagement costs across years. The column most buyers skip is sample handling. Race × gender cells thin out fast below a few hundred applicants per cell, and auditors differ sharply on whether they publish a minimum-cell threshold or quietly merge categories. Get that rule in writing before signing — it predicts how usable your published results will be.

Auditor Scorecard
Read the verdicts directly off the grid. Holistic AI takes the top slot for employers auditing multiple AEDTs — breadth across resume screeners, chatbots, and video scorers, plus methodology documents you can hand to counsel, make it the default for multi-tool stacks. Babble.ai is the value pick for the single-tool employer processing lower volumes of NYC applicants a year; its low-cost tier exists precisely for that volume band. Parity is the specialist call when your funnel runs on game-based or psychometric assessments. Credo AI fits enterprises auditing whole AI portfolios; Trusaic suits teams already running compensation-equity cycles.
The last column decides year two. The statute's recency rule — an audit is valid only if completed no more than one year before the tool is used — converts every one-time fee into a subscription whether you budgeted for it or not. Calendar the refresh at month 11 so the new engagement closes before the old report ages out, and when collecting quotes, ask whether the figure covers fieldwork only or fieldwork plus renewal. Rosters change, so verify each finalist's current status on the city's published list before shortlisting.
| Auditor | Independence policy | Race × gender cells | Methodology transparency | NYC-specific support | Indicative price band | Renewal economics |
|---|---|---|---|---|---|---|
| Holistic AI | Published independence attestation; audit work walled off from product sales | Full race × gender cross-tabs | Published LL144 methodology documents with impact-ratio formulas | Dedicated LL144 practice | Quote-based; typically bundled with platform subscription | Confirm whether renewal is re-quoted or locked in the original contract |
| Babble.ai | Independent pure-play auditor; no competing tool to sell | Standard race × gender matrix | Publishes its impact-ratio approach; request the sample-handling appendix | Built around LL144 filings | Lowest-cost tier of the five for single-tool audits | Flat annual renewal designed for repeat audits |
| Parity | Independent; behavioral-science lineage from pymetrics | Cells reported across assessment sub-scores | Psychometric validation methods published | Strongest on assessment-driven NYC funnels | Scoped per assessment instrument | Re-scoped whenever you add or change test items |
| Credo AI | Software vendor that also audits — demand the firewall in writing | Configurable cell-level dashboards | Methodology embedded in the platform's evidence trail | Multi-regulation coverage extending past NYC | Premium enterprise-suite pricing | Audit priced as a line item inside the portfolio license |
| Trusaic | Independent; sells no hiring tool of its own | Race × gender cells inherited from equity regressions | Regression-style documentation carried over from pay-equity practice | Adjacent experience with NYC equity reporting | Quote-based; often paired with pay-equity analytics | Annual engagement tied to your equity cycle |
| Vendor's own certificate | Zero unless independently verified — the statute's independence clause governs | Whatever the tool maker chooses to show | Self-graded; no external inspection of formulas or sample rules | Marketing-grade, not filing-grade | Free upfront | Costs nothing until an inspector asks for the independent report |
A bias audit is a photograph, not a video feed. The Department of Consumer and Worker Protection specifies what a Local Law 144 report must contain — the tool's description, impact ratios by race, ethnicity, and sex — but publishes no minimum sample size, no confidence-interval convention, and no threshold for how much data makes a ratio trustworthy. That silence is the gap this section occupies: the sections above tell you what to buy; none of them tell you how much the resulting document actually proves. The status-quo myth worth killing here is that a passing impact ratio is a durable property of the tool. It is a property of a tool-data-window triple, and it expires with all three.
Start with the limitations of the evidence. An impact ratio is a point estimate computed on whatever candidate pool the vendor logged during the audit window, and it carries sampling error the report rarely discloses. Intersectional cells — Black women, Asian men — shrink fastest, and a thin New York requisition can produce ratios that swing between quarters without the algorithm changing at all. There is also a provenance problem: the auditor computes from logs the vendor controls, so upstream data-hygiene failures never surface in the output. Much of the surrounding "audit readiness" literature is imported from adjacent fields, which compounds the confusion. According to Identity Confluence, its framework structures readiness through definition, pillars, and a checklist meant to move teams from audit-reactive to always audit-ready; according to the ITAM Roadmap published on Medium, that same roadmap targets a 50% reduction in device onboarding time. Those are throughput metrics from IT asset management. They measure whether paperwork moves quickly, not whether a selection model treats candidates equitably — and mistaking a completed checklist for statistical assurance is the quiet failure mode of otherwise diligent compliance teams.
| Your situation | Buy | Why it wins |
|---|---|---|
| Multiple AEDTs touch NYC candidates | Holistic AI | Breadth across tools plus published methodology documents |
| One tool, low NYC applicant volume | Babble.ai | Low-cost single-tool tier built for that volume |
| Funnel runs on game-based or psychometric tests | Parity | Behavioral-science depth from the pymetrics line |
| Enterprise auditing an entire AI portfolio | Credo AI | Governance-suite fit with a built-in evidence trail |
| Existing pay-equity program in place | Trusaic | Extends equity infrastructure into hiring audits |
| Offered the AEDT vendor's own certificate | Accept only with independent verification | Standalone certificates score zero on independence |

What the Data Doesn't Tell You
Variance across cases is the second caveat. Two auditors from the DCWP's published list, handed identical logs, can publish different ratios, because the law mandates the protected categories but leaves binning discretion — age bands, aggregation of small groups, handling of unknowns — to the analyst. A screener audited on its national applicant pool will show diluted rates relative to its New York slice, since the city's candidate mix is not the country's. The same tool can look clean in aggregate and marginal once stratified; both reports are technically honest.
So when does the rule break? Never in the direction of skipping the audit — but the annual cadence assumes a stable deployment, and four situations void that assumption. Retrain or materially re-weight the model mid-cycle and the year-old report describes a tool that no longer exists. Add a second AEDT — asynchronous video scoring, automated scheduling — and the audit covers only the first tool. Acquire a company and you inherit screeners whose audit history may not survive the transaction. Reuse a national audit for New York and the stratification problem above applies. Each case calls for a supplemental, off-cycle engagement scoped to the specific change; none of them changes who may perform the work, which remains a third-party firm on the city's published list.
The actionable upgrade: when you contract the auditor, require pre-registration of binning decisions and New York-stratified reporting in the engagement letter, plus a written trigger clause obligating a supplemental audit whenever the model changes. Read every report you receive as evidence about a specific tool, on specific data, over a specific window — never as a permanent verdict — and
Frequently Asked Questions
How much does NYC fine you per day for running an automated hiring tool without a completed bias audit?
$500 per violation per day, with each calendar day an AEDT touches a NYC candidate without a valid bias audit counting as a distinct violation for as long as the tool stays in service.
Does a chatbot we use in hiring count as an AEDT under Local Law 144?
A chatbot that ranks applicants is in scope because its output changes who advances, while one that only schedules interviews is exempt, since tools that merely aggregate, sort, or organize candidates without supplying independent decision-making inputs fall outside the law.
How much advance notice do I have to give candidates before an automated tool screens them?
At least 10 business days' advance notice, plus access to an alternative selection process.
Can our software vendor perform the bias audit itself instead of us hiring someone external?
Yes — the statute only requires at least two people, at least one of whom played no role in designing, building, or using the tool, so a vendor can field its own modeler plus an uninvolved colleague and stamp the output 'bias-audited,' leaving you holding the full $500/day exposure if DCWP rejects the independence claim.
What happens after DCWP opens a file on us — do fines start immediately?
No — after a candidate complaint or inspection, DCWP issues a formal notice of violation and you get a cure period during which penalties have not yet attached, with the $500-per-day metering beginning only if the violation persists past cure.
If our audit shows an impact ratio of 0.90, are we compliant?
Not necessarily — under the four-fifths rule a protected group's selection rate below four-fifths of the best-performing group's rate signals adverse impact, and a reported 0.90 can span 1.00 once a confidence interval is drawn on a thin cell, which is why durable audits publish interval estimates alongside every point estimate.
Quick answers
| What is the fine under NYC Local Law 144 for operating an automated employment decision tool without a completed bias audit? | Up to $500 per day, accruing for as long as the tool stays in service. |
| Under Local Law 144, what three duties does the statute impose on employers using an AEDT to score, rank, or cull NYC candidates? | Complete an independent bias audit, publish a summary of the results on your website before the tool goes live, and give affected candidates at least 10 business days' advance notice plus access to an alternative selection process. |
| How does the four-fifths rule relate to adverse impact in a bias audit? | An impact ratio below the threshold signals adverse impact, and ratios computed on intersectional cells too small to support them cannot survive recalculation by DCWP or a plaintiff's expert. |
| Who can perform the bias audit if no independent auditor is used? | At least two people, at least one of whom played no role in designing, building, or using the tool. |
| Why can an impact ratio of 0.90 still fail compliance even though it looks clean? | Because it can span 1.00 once a confidence interval is drawn on a thin cell, which is why durable audits publish interval estimates and significance tests alongside every point estimate. |
Also worth reading: NYC Law 144 Impact Ratios: Keep or Retire Your AI Screener: NYC Law 144 Impact Ratios: · Local Law 144: Nothing About Your Impact Ratio Stays Still: Local Law 144: Nothing About · Why people are choosing to start a career in artificial intelligence right now: Why people are choosing to