What the LL144 Bias Audit Actually Requires
New York City Local Law 144, also known as Local Law 144-21, created a specific compliance process for employers and employment agencies that use automated employment decision tools, or AEDTs, to assist with hiring or promotion decisions. The central requirement is not simply to test whether a model produces attractive accuracy figures. An employer must conduct a bias audit of the AEDT at least once annually, and again when the tool is substantially modified, using an independent auditor. The audit must examine the tool's selection rates and impact rates for candidates or employees of different sex, race and ethnicity, and intersectional categories including sex, race and ethnicity, and disability where required by the City's rules. The results must be summarized in a publicly available, dated report, with a notice explaining how the audit can be obtained and the data the employer used for its analysis.
Also worth reading: What Does an Automated Employment Decision Tool Audit Actually Involve in 2026? · What should an HR AI compliance audit strategy look like in 2026, and how do companies actually build one? · What EU AI Act HR Bias Mitigation Strategies Actually Work in 2026?
The law applies to employers and employment agencies hiring or promoting employees in New York City, not to every algorithm used in every workplace. A tool may qualify as an AEDT if its primary purpose is to assist or replace human decision-making in employment. The City has emphasized that the requirement applies even when a vendor supplies the software, because the employer remains responsible for the audit, notice, and publication process. As of September 24, 2026, organizations should still verify the current rules, enforcement positions, and any amendments rather than treating this description as a substitute for current legal advice. The practical purpose is transparency and accountability, not a guarantee that an employer will never make a discriminatory employment decision.
Why a “Bias Audit” Is Not Just an Accuracy Test
A conventional model-performance review may ask whether the system predicts whether a candidate will pass a test or remain employed. An LL144 bias audit asks a different question: does the system distribute favorable outcomes differently across protected groups? The City's rules use selection-rate and impact-rate comparisons that can reveal disparities even when overall predictive performance appears strong. Accuracy can therefore be high while the employment consequences remain unfair or unlawful. This distinction matters because an employer can improve average prediction quality without improving fairness for women, people of color, applicants with disabilities, or applicants belonging to intersectional groups.
The audit should be connected to the actual employment decision process. If a system screens résumés, ranks applicants, recommends interviews, or evaluates promotion candidates, the audit should reflect the relevant tool, data, thresholds, and decision rules in use. Auditing a different model, an earlier version, or a hypothetical dataset may create a technically polished report with little operational value. The auditor should also understand what counts as a selection, how rejected applicants are represented in the data, and whether a human decision-maker can override the system. A model that recommends but does not automatically decide may still fall within the law depending on how substantially it guides the decision.
Bias is also a systems issue. Historical hiring data can reflect unequal access to jobs, biased job descriptions, inconsistent performance ratings, or differences in how managers describe candidates. Removing a protected characteristic from the model does not necessarily remove the social pattern encoded in proxies such as graduation year, ZIP code, employment gaps, or specific schools. A credible audit therefore examines data quality, feature design, selection rules, outcomes, and the human process surrounding the tool. The report should not present a single statistical ratio as proof of compliance or discrimination.
The Step-by-Step Methodology
The first step is to create an accurate inventory of the AEDTs used for New York City hiring and promotion. The inventory should identify the vendor, product name, model version, intended use, decision point, protected populations affected, vendor documentation, and whether the tool has changed since the last audit. This stage prevents an organization from auditing a low-risk tool while overlooking a resume screen or promotion-ranking system that has greater impact. It also clarifies whether the organization is acting as an employer, an employment agency, or both for particular workflows.
Next, the employer should preserve and evaluate the data used for the audit. The rules require the audit report to include the data and methodology used, subject to permitted privacy protections and specified exceptions. The organization should determine whether it has applicant-level or employee-level outcome data, protected-group information, dates, job categories, selection decisions, and information about unavailable demographic data. Missing data is not a reason to invent labels or quietly exclude a group from the report. The auditor should document missingness, its possible effect on results, and how the employer handled records with no reported race, ethnicity, sex, or disability information.
The audit should then calculate the required selection and impact measures, identify statistically significant disparities where the applicable methodology calls for them, and test the results in meaningful employment subgroups. Results should be reported by job or role when sample sizes permit, while recognizing that small groups can produce unstable percentages. The employer should also consider the effects of using multiple thresholds, such as a stricter cutoff for one category of candidates. A reasonable methodology may include statistical tests, confidence intervals, practical-effect measures, sensitivity analysis, and review of qualitative evidence such as adverse-impact concerns. The exact approach should be documented and reproducible, not selected simply to produce a preferred conclusion.
Finally, the employer must summarize the results, explain limitations, publish the report, provide required notice, and establish a process for candidates or employees to request access. The final report should be written for a broad audience, including applicants, and should distinguish measured disparity from an allegation of unlawful discrimination. A public report that is technically complete but impossible to locate does not meet the practical goal of transparency. Organizations should coordinate the publication process with counsel and the auditor before disclosing confidential data.
The Required Outputs, Notices, and Public Report
The public report normally needs a clear statement of the employer or agency, the audited tool, the audit period, the data sources, the methodology, the tested categories, the results, and any limitations. The report should be dated and remain available at a stable location so that applicants can inspect it. Notice must inform candidates and employees about the employer's use of AEDTs and explain how to request access to the audit information. The access process should specify where to submit a request, what information is required to locate relevant records, how the request will be handled, and what contact details apply.
The report should not be mistaken for a legal finding. A statistically significant difference in selection rates may justify further investigation, remediation, or monitoring, but it does not automatically establish liability under every applicable discrimination statute. Conversely, a report with no statistically significant difference does not prove that the tool is lawful in every circumstance. The organization should explain the statistical uncertainty, sample-size constraints, data limitations, and possible confounding factors. It should also document corrective actions, such as revised thresholds, feature changes, expanded monitoring, human-review procedures, or suspension of the tool.
The scope and enforcement of the requirement are not entirely static. New York City's rules have been discussed alongside developments in Illinois, Connecticut, and other jurisdictions that impose disclosure or anti-discrimination obligations involving employment AI. Those developments may create overlapping duties, but they do not eliminate the need to apply the New York City requirements when the tool is used within the City's jurisdiction. Employers should distinguish among model governance, vendor due diligence, discrimination risk management, and LL144-specific public reporting. Combining them into one vague statement—“we use responsible AI”—usually makes compliance harder to demonstrate.
What an Auditor Should Examine Beyond the Math
An independent auditor should be genuinely independent of the tool's design and operation, but independence does not mean that the auditor lacks practical knowledge of employment compliance. The auditor should be able to inspect the actual system, understand how features are generated, test the relevant versions, and challenge unexplained exclusions. The employer should provide reasonable access to documentation, data dictionaries, model cards, validation results, vendor certifications, change logs, and decision logs. If a vendor refuses to provide enough information to evaluate bias, that refusal is itself a governance issue that should be documented.
The auditor should also examine the process around the algorithm. Human reviewers may treat a recommendation as a fact, may ignore contrary evidence, or may apply different standards to different groups. Training can unintentionally tell reviewers to prioritize similarity to past hires, which can reproduce historical exclusion. The audit should ask whether reviewers know the tool's limitations, whether overrides are recorded, whether candidates receive an opportunity to correct inaccurate information, and whether the employer monitors changes in outcomes after deployment. A model that is statistically tested but embedded in a poorly managed workflow remains a business and compliance risk.
Special care is warranted for disability-related information and intersectional analysis. Employers may have legitimate reasons for handling sensitive data, but the audit should not treat a small sample as evidence that no issue exists. Missing or inaccessible data can make a tool look neutral because the people most affected were never visible in the dataset. The methodology should describe how the organization handles missing demographics, small cohorts, multiple testing, and data collected under different legal or contractual constraints. The best reports are transparent about what was not measured instead of presenting silence as fairness.
Comparing Internal, Vendor, and Independent Approaches
| Feature | Internal analysis | Vendor-provided assessment | Independent LL144 audit |
|---|---|---|---|
| Independence | Lower unless a separate team is structurally separated | Depends on the vendor's relationship with the employer | Highest when the auditor has no role in building, selling, or operating the tool |
| Access to system details | May be limited by contracts and technical silos | Often strongest for the vendor's own product | Must be supported with employer-provided access and documentation |
| Cost and speed | Often lower cost; may be faster for routine inventory work | Can be efficient, but scope and independence may be limited | Usually more expensive and time-consuming; strongest defensibility |
| Fit for LL144 | Useful as preparation, not automatically sufficient | Useful as evidence if it meets the required independence and methodology standards | Best when the full City reporting and public-access process is required |
| Main risk | Self-review and weak challenge | Conflicted conclusions or limited customization | Access barriers, sampling limitations, or misunderstood scope |
Common Mistakes That Undermine an LL144 Audit
One common mistake is assuming that removing race, sex, or disability from the model proves fairness. The tool may still use proxies, historical patterns, or threshold rules that generate unequal outcomes. Another is auditing only successful hires. Selection-rate analysis needs a defensible population of applicants or candidates, not merely people who passed the system, and impact-rate analysis must use an appropriate reference group and outcome definition. Employers also sometimes publish a report without the required notice, omit the audit period, or fail to explain how an individual can request access.
Timing errors are frequent. Organizations may wait until the end of the calendar year, even though the requirement is generally tied to an annual cadence and additional review when the tool is substantially modified. They may audit the vendor's original model while operating a customized configuration with different thresholds or data sources. Others may treat synthetic data as a substitute for real employment data without explaining why it is appropriate. Synthetic data can help with testing or software development, but its assumptions may reproduce or conceal the original bias, so the methodology and limitations should be explicit.
The final mistake is treating a completed audit as the end of compliance. A tool can be changed, a data pipeline can deteriorate, and reviewers can develop new habits after the report is published. Monitoring, change control, incident response, vendor review, and periodic re-testing should be part of the control system. This is also where AI-powered labor-law compliance and HR regulatory management tools can help organizations maintain inventories, evidence requests, deadlines, and version histories, provided the software itself is evaluated rather than assumed to solve the underlying legal problem.
When to Act, and What It May Cost
An employer should act immediately if it uses an automated tool to screen, rank, recommend, or assist New York City hiring or promotion decisions and has not confirmed the applicable AEDT classification. The first 30 days can be spent identifying the vendor, mapping the decision flow, listing the relevant job populations, and determining what demographic and outcome data exist. The next phase should establish an audit plan, appoint an independent reviewer, and set publication and notice dates. Organizations should not wait for an enforcement notice to discover that their vendor contract prevents the required analysis.
There is no single universal LL144 audit price. A narrowly scoped, lower-complexity review may cost several thousand dollars, while a multi-model, multi-workforce assessment with extensive data work can cost tens of thousands or more. Costs increase with the number of systems, job categories, jurisdictions, historical data periods, customization, statistical subgroup analysis, and the amount of documentation required. Vendors may offer compliance packages, but buyers should ask whether the price includes the independent test, the public report, notice language, access-request support, and remediation tracking. A low-cost template that only summarizes vendor marketing claims is not equivalent to a defensible audit.
For smaller employers, a staged approach may be more realistic than treating every HR technology purchase as a full-scale project, but staged does not mean optional where the rule applies. Legal counsel should confirm the jurisdiction, the tool's status, the current reporting language, and any interaction with employment discrimination obligations. As of September 24, 2026, organizations should also check whether local or state rules create additional notice, testing, or recordkeeping duties. The defensible position is not that the algorithm is certified harmless, but that the employer can show what it used, who tested it, what was found, what was corrected, and how the public can inspect the evidence.