# What Should Employers Include in an AI Employment Audit in 2026?

ailaborbrain.com · September 27, 2026

> What an HR AI audit actually covers An HR AI audit is a documented review of how artificial intelligence affects employment decisions and workforce...

## What an HR AI audit actually covers

An HR AI audit is a documented review of how artificial intelligence affects employment decisions and workforce management. It should cover recruitment, résumé screening, interview scheduling, promotion, performance management, employee monitoring, termination, accommodation requests, compensation, and payroll. The audit also needs to test whether the system’s inputs, outputs, human reviewers, and data-retention practices comply with applicable discrimination, privacy, notice, recordkeeping, and labor rules. It is not merely a security scan or a policy acknowledgment. By September 28, 2026, an employer may already have duties under federal law and new state or local rules governing automated employment decision tools, even if no single federal statute is titled “the HR AI audit.”

**Also worth reading:** [What Is the 2026 Employment AI Compliance Checklist for US Employers?](https://ailaborbrain.com/knowledge/what_is_the_2026_employment_ai_compliance_checklist_for_us_employers.php) · [How do employers comply with automated employment decision tools regulations in 2026?](https://ailaborbrain.com/knowledge/how_do_employers_comply_with_automated_employment_decision_tools_regulations_in_2026.php) · [How Do You Build an AI Hiring Audit Checklist for Fair, Compliant Employment Decisions in 2026?](https://ailaborbrain.com/knowledge/how_do_you_build_an_ai_hiring_audit_checklist_for_fair_compliant_employment_decisions_in_2026.php)

The audit’s central question is whether the employer can explain not only what the technology does, but also why a person reasonably determined to use it and how risks were controlled. A defensible review therefore combines legal requirements with technical testing and operating evidence. A policy that says the company “uses AI fairly” is not enough without model versions, decision records, test results, reviewer instructions, vendor responsibilities, and incident procedures. The appropriate depth depends on the system’s role: an internal tool that suggests interview dates presents different risks from software that rejects applicants or ranks employees for termination.

## The legal and operational tests employers should apply

The first test is governing law. Employers should map each automated employment tool to the jurisdictions in which affected workers live or work, because remote applicants and employees can trigger obligations beyond the office where the tool is deployed. Title VII, the ADA, the ADEA, the Genetic Information Nondiscrimination Act, and other federal statutes continue to prohibit discriminatory employment practices regardless of whether a human or algorithm recommends the result. Newer state and local statutes may add disclosure, data-access, impact-assessment, bias-audit, or appeal duties. New York City Local Law 144, for example, requires covered employers and employment agencies to conduct an annual bias audit of a covered automated employment decision tool and publish a summary.

The second test is decision impact. The review should compare the tool’s rates of selection, rejection, access, error, and adverse impact across legally protected groups and relevant job categories. Statistical testing does not by itself prove unlawful discrimination, but large or persistent disparities require a documented, job-related explanation. The third test is human control: reviewers should receive meaningful information, have authority to change outcomes, and understand when not to rely on the output. The fourth is governance: vendors should provide documentation about training data, intended use, known limitations, monitoring capabilities, security, and contractual remedies. These four tests convert a broad compliance concern into evidence an employer can examine and improve.

## A practical audit method from system selection through deletion

Begin with an inventory dated no later than 30 days after the audit begins, assigning each tool an owner, purpose, user population, vendor, model version, data source, and decision impact. Classify systems into high-, medium-, and low-impact groups, giving systems that screen, rank, diagnose, discipline, compensate, or terminate the most attention. High-impact tools may receive quarterly testing, while a low-risk calendar reminder may be reviewed annually. These are governance recommendations rather than universal legal deadlines, but they create a reasonable review cycle when laws or model behavior change quickly.

Next, collect at least 90 days of operating data for an initial review, including more history where seasonal hiring, layoffs, or model updates make a short period misleading. Test selection rates, error rates, score distributions, rejection reasons, overrides, and the effects of missing or proxy data. The team should also examine whether the tool causes an adverse disability accommodation request to be delayed, flagged, or disclosed without proper authority. Compare results with the employer’s actual job qualifications, not merely with an abstract ideal. Finally, document retention and deletion practices, obtain vendor support, assign corrective actions, and set a date—normally 30 to 90 days after the audit—for management to verify completion.

## Required audit evidence and recordkeeping

A useful audit file contains the governing-law inventory, system diagrams, data-flow descriptions, vendor contracts, intended-use statements, model and change logs, test protocols, sample outputs, reviewer training, and records of human decisions. It should also contain notices actually provided to candidates or employees, accommodation procedures, complaint channels, appeal outcomes, security controls, and incident tickets. The record must show who approved the system and what evidence informed that approval. Copies of marketing claims should not substitute for operating evidence because accuracy claims developed in another setting may not hold in a specific employer’s workforce.

Retention periods must match the jurisdiction, record type, litigation hold, and vendor agreement; there is no single safe duration for every category of HR AI data. New York City rules, for example, include data-retention requirements tied to the tool’s use and the candidate’s interaction with the employer. Employers should restrict access to personnel data, log administrative changes, and avoid retaining sensitive applicant attributes longer than needed. Ordinary records such as applications may have different retention rules from biometric identifiers, inferred health information, or model-training data. The audit should identify the legal or business reason for each period and specify who may extend it through a legal hold.

| Feature | Annual compliance snapshot | Continuous AI controls program | External specialist review |
| --- | --- | --- | --- |
| Best use | Smaller workforce or limited tool use | Recruiting, workforce, HRIS, and vendor operations | High-impact or legally complex deployments |
| Frequency | Usually 12 months | Monthly exceptions and quarterly testing | Before launch and after material changes |
| Staffing | HR, legal, compliance, and system owner | Cross-functional team plus control owners | HR/legal leads working with independent testers |
| Cost | Often $3,000–$15,000 if performed internally | Often $10,000–$75,000+ annually | Often $20,000–$150,000+ per engagement |
| Limitation | May miss between-period changes | Requires mature ownership and reporting | Does not replace internal accountability |

The cost figures are planning ranges, not regulatory fees. They reflect typical labor for a US audit and vary sharply by system count, data volume, testing depth, and whether a law firm, economist, statistician, or security laboratory is engaged.

## How employers should test discrimination, explainability, and accessibility

Testing should begin with job-relatedness and adverse-impact analysis, followed by review of how the model treats proxies for race, sex, age, disability, religion, and other protected characteristics. Where sample size permits, test both the overall population and relevant stages, such as application, screening, interview, offer, pay, promotion, and termination. The team should compare error rates as well as selection rates, because a system can produce equal selection percentages while still making a materially different type of error for one group. Statistical confidence intervals matter; a very small sample may show an alarming percentage without producing a reliable conclusion.

Accessibility testing should cover screen-reader compatibility, captioning, alternative formats, language access, and the employer’s accommodation process. An algorithm may be internally consistent but inaccessible to candidates with disabilities, while automated scoring may disadvantage applicants who communicate differently. Explainability should be proportionate to the audience and use: an employee needs a clear basis for challenging a decision, a manager needs review instructions, and a regulator may need model and impact documentation. Employers should not claim that a system is “explainable” merely because the vendor can display a score. The explanation must connect the output to job duties and allow a qualified human to evaluate the underlying information.

Manual review by one recruiter is not a cure when the human simply accepts the ranking without time, authority, or independent data. A credible override process records disagreement, requires reconsideration, and measures whether protected groups receive different results when reviewers exercise judgment. Some strong systems also offer an applicant notice explaining the principal AI-related characteristics, purpose, and contact channel, subject to applicable trade-secret and legal limits. The audit should test whether notices are understandable at the point of application, not merely whether they appeared in a privacy policy buried after submission.

## Vendor, security, privacy, and employee-monitoring controls

Vendor review should address more than algorithmic accuracy. Contracts should identify permitted purposes, prohibit unauthorized training or reuse, require security updates, define breach notification, preserve relevant records, and allocate responsibility for discrimination, intellectual-property, confidentiality, and data-access claims. The employer should know whether support tickets, interview recordings, voice data, inferred attributes, or candidate profiles are used to improve the service. Data processing terms should reflect actual transfer and storage locations. Automated decisions should not move into functions, countries, or purposes that the original assessment did not examine.

Security and privacy controls include access restrictions, encryption where appropriate, logging, retention limits, incident response, and validated deletion. For employee monitoring, the audit should confirm necessity, proportionality, transparency, and compliance with location-specific consent, off-duty-conduct, biometric, and workplace-privacy requirements. AI can infer a worker’s pregnancy, disability, union activity, health status, or other sensitive trait from apparently ordinary data, so inactivity in the vendor’s feature menu does not prove the system is free of such risks. Test data must be protected as well. Public reports, recruitment files, and email threads containing real applicants should not be copied into an experimental model merely because they are convenient.

The audit should establish a clear escalation route for affected workers, including complaints, accommodation requests, correction requests, and contests of automated recommendations. Employers should investigate whether monitoring or automated scoring changes supervisors’ behavior and whether employees can discuss working conditions without disproportionate surveillance. A monitoring system can create a chilling effect even when it records no protected activity. Periodic reassessment is therefore necessary, particularly after new capabilities are activated. The system owner should document the business reason for each monitored behavior, the less intrusive alternatives considered, and the period during which monitoring is justified.

## Common mistakes that make an HR AI audit unreliable

The most frequent mistake is treating a policy attestation as an audit. A signed statement that no AI is used may omit shadow tools, vendor features embedded in the applicant tracking system, recruiter extensions, and managers’ informal applications of generated recommendations. Another mistake is reviewing only final rejection statistics when disparities arise earlier, such as in keyword matching or access to interviews. Employers also err by testing one model version and assuming later updates preserve its behavior. Model changes, data feeds, thresholds, and third-party integrations can alter results without obvious notice to HR.

Undocumented human override, “sample” testing without a defensible selection method, and reliance on vendor assurances are additional weaknesses. Employers sometimes test protected data that the production tool does not use while missing proxies or inaccessible data that it does use. Others publish a legally required bias-audit summary but never connect it to remediation, because they mistake a disclosure obligation for a control process. The most damaging mistake is treating disparity as proof of motive or dismissing it because final pay or selection percentages appear close. Statistical patterns are signals for investigation, not substitutes for legal analysis, job evidence, or an assessment of how decisions were made.

A weaker alternative is to perform an external review without assigning an internal owner. External specialists may improve testing, but the employer remains responsible for workforce decisions and cannot outsource accountability to a purchased report. Another alternative is to launch first and audit later. Early use can create thousands of consequential records, expose applicant data, and establish employment decisions that are difficult to unwind. Organizations should complete a minimum viable review before deployment and expand it as risk increases. The review does not need to predict every future legal development, but it should establish a documented baseline before people are affected.

## When employers should act, escalate, or involve counsel

Action should begin before procurement, contract signature, or upload of worker data whenever AI will affect selection, evaluation, access to opportunity, or other employment terms. Existing tools should be inventoried immediately if they are used in hiring, performance management, promotion, discipline, compensation, or termination, especially where state or local disclosure and audit laws took effect in 2023 through 2026. A legal review is particularly important for high-volume recruiting, workers in multiple jurisdictions, tools used for people with disabilities, or deployments connected to medical or biometric information. The review should identify a responsible executive, not simply place the issue in an already overloaded HR analyst’s queue.

Escalate when testing shows a material disparity without a job-related explanation, the tool prevents meaningful human review, or an accommodation consistently causes an adverse result. Immediate investigation is also appropriate after unexplained model changes, unauthorized data sharing, security incidents, repeated false decisions, or evidence that candidates cannot discover or challenge the tool’s role. Authorities may request notices, audit reports, underlying data, contracts, and decision records, so the evidence should be retrievable rather than scattered across inboxes. A litigation hold or regulator inquiry changes ordinary retention practices. Counsel should coordinate privilege and preservation without using legal labels to conceal routine operating records.

A lightweight initial review can take two to four weeks for one or two low-risk systems, while a multi-system review involving statistical, legal, technical, and vendor work may take eight to sixteen weeks. Urgent incidents should not wait for a full annual report; contain the risk, pause or limit the affected function, preserve records, and evaluate affected workers. Employers must balance the benefits of automation against the risk that poor data or opaque process worsens an existing hiring bottleneck. There is no universal threshold based only on a headcount, but the number of people affected, duration of use, financial or career consequences, and reversibility of decisions provide practical risk measures.

## The minimum 2026 audit standard

The definitive minimum is a dated, risk-based review that names every relevant system and explains the law, data, testing, human oversight, notice, vendor, retention, security, and remediation applied to it. Management should approve the scope, certify that known limitations were disclosed, and receive metrics showing selection, error, override, complaint, and incident trends. The audit should include lawful-use statements that match actual configuration, documentation of material model changes, and accessible procedures for correction and accommodation. Required public summaries or notices must be prepared under the rules applicable to the specific tool and employer; copying another company’s template is not enough.

The audit is complete only when corrective actions have owners and due dates and management has tested whether the changes worked. Review the result at least annually and after significant legal, organizational, vendor, or technical changes. In New York City, an annual bias audit is a legal requirement only where Local Law 144 applies; that requirement does not transform the annual cycle into a safe-harbor for employers outside the city or for violations of other laws. Illinois’s employment AI framework, Connecticut’s amendments, state consumer and privacy laws, and emerging rules elsewhere may create distinct duties. As of September 28, 2026, the safest approach is jurisdiction-by-jurisdiction analysis rather than assuming one national checklist can satisfy every obligation.

A strong HR AI audit does not claim that technology is inherently fair or unfair. It creates traceable evidence showing which decisions were automated, how those decisions were tested, where people intervened, and what happened when outcomes appeared questionable. That record helps the employer correct problems, respond to workers, negotiate with vendors, and demonstrate responsible governance. It also reveals when a tool adds little value or presents risk disproportionate to its benefit. The objective is not maximal automation; it is dependable employment decision-making that employers can explain and workers can challenge.

## Quick answers

### Is a bias audit required for every employer using HR AI?

No single federal rule requires all employers to conduct an AI bias audit. Requirements vary by jurisdiction and tool, with New York City Local Law 144 imposing specific annual audit and publication duties for covered automated employment decision tools. Federal discrimination, accommodation, and privacy duties can still apply to other employers.

### How often should an employer review an AI hiring system?

A formal review at least once each year is a common governance practice, and high-impact systems may need quarterly or event-driven testing. Reviews should also follow model, data, vendor, law, or decision-process changes. Frequency should reflect the tool’s consequences, workforce exposure, and applicable law.

### Can an employer rely on a vendor’s AI audit?

A vendor report can provide evidence, but it does not replace the employer’s review of actual use, data, workforce, notices, and decision outcomes. Contracts should preserve relevant documentation and require notice of material changes. The employer remains responsible for employment decisions made with the tool.

### What should an employer do after finding an adverse-impact disparity?

Preserve the evidence, determine whether the result is statistically reliable, and compare the tool with documented job requirements. The employer should assess selection and error rates, examine human review, and pause or limit the affected use when risk cannot be promptly controlled. Counsel or qualified specialists may be needed before remediation.

### What do HR AI audits usually cost?

An internal review of a small number of lower-risk systems may cost roughly $3,000–$15,000 in labor, while broader controls programs may run from $10,000 to $75,000 or more annually. Independent statistical, legal, security, or model testing can add substantially and may involve $20,000–$150,000 or more. These are planning estimates, not statutory fees.

Canonical: https://ailaborbrain.com/knowledge/what_should_employers_include_in_an_ai_employment_audit_in_2026.php
Markdown: https://ailaborbrain.com/knowledge/what_should_employers_include_in_an_ai_employment_audit_in_2026.php/index.md
