If you screen candidates for a New York City role with an AI tool, the bias audit is your duty, and your vendor's audit counts only if your own applicant data went into it. That is the rule in the city's official FAQ: an employer can rely on a multi-employer audit only if it gave its own historical data to the auditor, or if this is the first time it is using the tool. Most buyers assume the vendor's audit transfers to them. Once you have used the tool and kept your data to yourself, it doesn't.
So the purchase decision is which independent auditor runs the numbers on your data. For an employer that has already used a screening tool, our pick is DCI Consulting Group, because its AI tool review service computes adverse impact on your own local decisions and adds criterion validation and job-relatedness review, which is the evidence a disparate-impact claim actually turns on. Use BABL AI if you want a fast, standards-based opinion, and insist on a direct engagement. Use Holistic AI if you run several tools and want one dashboard and a hosted summary. Warden AI loses as your Local Law 144 audit: it tests on its own consented datasets, which the city accepts only when your historical data is insufficient. It is a good signal when you vet a vendor.
| Auditor | What you get | Whose data it tests | Published price (checked Oct 10, 2026) | Pick it if | Skip it if |
|---|---|---|---|---|---|
| DCI Consulting Group | I/O psychology review, adverse impact at job-title level, criterion validation, NYC summary of results | Your local employment decisions | None; contact sales | You have used the tool for a year and want an audit that also holds up in litigation | You need a monthly release check on a product you build |
| BABL AI | ISAE 3000-style assurance opinion (Pass / Minor Remediations / Fail), public summary drafted | Whatever is submitted; attestation checks your testing, direct engagement measures it | None; contact sales | You have clean data and need an opinion in weeks | You want someone else to do the analysis but would sign up for an attestation |
| Holistic AI | Platform: subgroup metrics, monitoring dashboard, summary of results hosted for your site, internal report | Data you load into the platform | None; contact sales | You run several AEDTs across jurisdictions | You run one tool and have no one to own a platform |
| Warden AI | Continuous audits, public assurance pages, 10 protected classes | Warden's proprietary consented datasets | None; "Book Demo" | You are the vendor, or you want evidence before you sign | You need your own Local Law 144 audit after a year of use |
The workload for this comparison: one AI resume screener, already in use for a year, scoring about 40,000 applicants a year for roles tied to a New York City office, with self-reported race and sex on most applications. That is roughly the size of the pool in Pfizer's posted HiredScore audit, which counted 40,821 applicants by sex.
Who Owes What Under Local Law 144?
The hiring company or employment agency owes everything; the vendor owes nothing. The DCWP FAQ says it in one line: "The vendor that created the AEDT is not responsible for a bias audit of the tool."
An automated employment decision tool (AEDT) is software that uses machine learning, statistical modeling, data analytics or AI to substantially assist or replace discretionary hiring or promotion decisions. The law reaches you when the job is located at a New York City office, at least part time, or is fully remote but tied to one. Screening counts; so does the final decision.
Before you use the tool, three things must be true, per the city's AEDT page and FAQ:
- An independent auditor completed a bias audit within the past year. You can rely on an audit for one year from the date it was conducted.
- You published a summary of the results and the tool's distribution date (the date you started using it) on the employment section of your website, or linked to it.
- You told NYC-resident candidates you use the tool, what it assesses, and how to request an accommodation, at least 10 business days before use. A notice on your careers site works for applicants.
The summary must show the date of the audit, the source and explanation of the data, the count of people in an unknown category, and the number of applicants, selection or scoring rates, and impact ratios for every category. An impact ratio is a group's selection rate divided by the rate of the most-selected group; 0.80 is the conventional four-fifths line, though the law sets no pass mark. The FAQ says the law "does not require any specific actions based on the results of a bias audit," while federal, state and city anti-discrimination law still applies to what you do next.
Penalties under the statute are up to $500 for a first violation and $500 to $1,500 for each one after. Each day you use a tool without a valid audit is a separate violation, and each missing notice is its own violation.
Enforcement so far has been thin, and that is changing. A New York State Comptroller audit published December 2, 2025 covering July 2023 to June 2025 found DCWP received two AEDT complaints, reviewed 32 companies and found one instance of non-compliance, while the Comptroller's own review of the same companies found at least 17 potential issues. DLA Piper's summary reports that 75% of test calls to 311 about AEDTs never reached DCWP, and that the agency agreed to most of the recommendations, including interviews and demonstrations of AEDT tools where appropriate.
How Do the Four Auditors Compare on Your Data?
DCI does the most work on your own outcomes; BABL is the quickest way to an opinion; Holistic is a platform; Warden tests the product rather than your use of it. Every one of them is legal to hire. They differ on whose data they examine and how deep they go past the impact-ratio table.
DCI Consulting Group: The Pick for Employers
DCI is an industrial-organizational psychology and employment-compliance firm. Its AI-based tool review covers subgroup adverse impact on "local employment decisions" at the unit of analysis such as job title, input review for job relevance, comparison of the tool's output with I/O psychologists' judgments, criterion-related validation where feasible, and jurisdiction deliverables for New York City, California, Colorado, Illinois and the EU AI Act. For Local Law 144 it produces the bias audit and the summary of results. HireVue engaged DCI as its external bias auditor, so the firm knows the vendor side too.
Why it wins for an employer: the Local Law 144 table is the floor. If a screener does produce a disparity, the question in a lawsuit or a California proceeding is whether the tool is job-related and whether you tested and acted. California's rules make that testing relevant to the defense (more below). DCI's service is built around exactly that evidence.
Who should not pick DCI: an HR tech vendor that needs a bias check on every model release, and an employer that only wants the cheapest possible summary for one low-volume tool. DCI publishes no price; expect a consulting engagement.
BABL AI: fast, structured, and read the engagement type
BABL runs audits under ISAE 3000 assurance practice, its lead auditors are ForHumanity Certified Auditors, and its audits end in an opinion of Pass, Minor Remediations or Fail. A more recent BABL explainer puts most engagements at five to eight weeks from kickoff and names the distinction you need: an attestation verifies results the client already produced, while a direct engagement has BABL measure the system itself. BABL's own guidance agrees with the city: "Deployers who did not contribute their own data to your audit remain independently obligated."
The two BABL vendor reports discussed below are attestations. The Eightfold AI Interviewer report dated June 29, 2026 lists "Testing conducted by: Eightfold". The Harver Soft Skills Platform report dated July 17, 2025 lists "Testing conducted by: Harver". That is a legitimate model for a vendor. For your own audit, you want the auditor computing the ratios from your applicant file.
Who should not pick BABL: an employer without a data team that would end up on an attestation of numbers it produced itself.
Holistic AI: a platform for a portfolio of tools
Holistic AI sells a bias audit platform for AEDTs. The UK government's assurance case study describes it calculating metrics by race/ethnicity and sex, reporting findings in a dashboard for ongoing monitoring, and producing a summary of results that Holistic hosts for you to share, plus a fuller internal report. That suits a company with a sourcing tool, a screener and an assessment all touching New York roles.
Who should not pick Holistic: a company with one tool and no owner for another platform. You will be paying for monitoring capacity you never open. If you already run an AI governance platform such as Credo AI, check whether the AEDT already sits in that inventory before adding a second system of record.
Warden AI: the loser for this job
Warden is built for the vendor side. Its site lists ongoing audits mapped to Local Law 144 and the EU AI Act, "testing on proprietary consented datasets," and audit logs with dataset snapshots, with separate offers for vendors, staffing firms and enterprises. Greenhouse's bias audit statement says Warden tests Talent Matching across 10 protected classes on Warden's own dataset, runs disparate impact analysis plus tests of whether names or gendered words change the output, audits each new version before release, and publishes results on a public dashboard.
That is strong evidence when you are choosing a product. It is weak as your Local Law 144 audit. The FAQ says historical data "must be used," and allows test data only "if there is insufficient historical data available to conduct a statistically significant bias audit." An employer with a year of applicants through the tool is rarely in that position. Greenhouse's statement does not claim the audit satisfies an employer's obligation.
Who should not pick Warden: an employer looking for its annual audit. Who should: the vendor, or a buyer who wants to see a monthly record before signing.
What the Published Audits Actually Show
The audits are public by law, so read the ones from your shortlist before you read the marketing, and compare them on data source, thresholds and lowest ratio.
Pfizer on Workday's HiredScore
The summary uses Pfizer's own data: applicants to requisitions with NYC hires open from August 2022 to February 2025. It reports three cut points (grade A; A or B; A, B or C). The lowest ratio is Hispanic men at grade A, 0.822; Hispanic applicants overall are at 0.886, and Asian applicants are at 0.885 for A or B. Groups under 2% of applications were excluded, and 2,953 people had unknown sex or race. This is the shape you want: your data, several thresholds, the small-group rule applied in the open. The summary describes the audit as one "Pfizer conducted," so if you copy the format, make sure your file names the independent auditor.
Harver's pymetrics Assessment
The summary covers January to December 2024, 283,256 candidates by sex, a 50th-percentile threshold for "Recommend," and a lowest race ratio of 0.914 for Asian candidates. Groups such as Two or More Races (2,440 candidates) show "N/A," consistent with the 2% exclusion. The report is dated July 17, 2025. Under the one-year rule, it cannot support use after July 2026; check the date on whatever your vendor sends you.
Eightfold's AI Interviewer
BABL's June 2026 report passed the tool, with an emphasis of matter: race/ethnicity and intersectional results come from "synthetic data due to limited historical data," built from "LLMs acting as job candidates," with demographics signaled "via a candidate's name alone." Tests ran on the most common job descriptions in Eightfold's client base with default questionnaires and scoring. The auditor warns the results are subject to the limits of that setup. If you configure your own rubric, the published audit did not test your configuration.
Why the Answer Changes With Where the Role Sits
Of the four jurisdictions here, only New York requires a published audit. The other three change what you keep, what you disclose and what counts as a defense, so the audit you buy should satisfy the strictest place your roles sit.
- California: the Civil Rights Council's automated-decision system regulations took effect October 1, 2025. They apply FEHA to AI tools, define an employer's "agent" (which can include the vendor screening for you), and require you to keep ADS data for at least four years. Mayer Brown's analysis notes that anti-bias testing, its quality, scope and recency, and how you responded to the results are relevant to the defense.
- Illinois: HB 3773 took effect January 1, 2026. Foley's summary says it bars discriminatory AI use across recruiting, hiring, promotion and discipline, prohibits zip codes as a proxy for protected classes, and requires notice of AI use, with the Department of Human Rights to set the details by rule.
- Colorado: SB 26-189, signed May 14, 2026, replaced the 2024 AI Act with a narrower notice-based law effective January 1, 2027. Per JD Supra's summary, employers must give notice at the point of use, explain the tool's role within 30 days of an adverse outcome, and allow correction and human review; attorney general rules and an xAI lawsuit could still move the details.
Litigation sits under all three. In Mobley v. Workday, a federal court preliminarily certified a nationwide age-discrimination collective in May 2025 and later included applicants scored by HiredScore AI features. The vendor being sued does not take you out of the case.
What to Demand From the Vendor Before You Sign
The vendor owes the city nothing, so everything you need from it has to be in the contract. Put these in the order form or the DPA:
- Audit cooperation: the vendor supports an independent auditor of your choosing once a year at no extra charge, including model documentation and staff interviews.
- Your outcome data in a usable form: every score, grade or rank the tool produced for each applicant, with requisition ID and timestamp, exportable as a file, during the contract and for a set period after it ends. Without this, no auditor can run your numbers.
- Retention matched to the longest rule: four years for California roles, and longer if your litigation hold policy says so.
- Notice support: plain-language text describing what the tool assesses, for your careers page and job postings, updated whenever the vendor changes the model.
- Change notification: advance notice of model or feature changes that could alter scores, so you can decide whether a re-audit is needed.
- Indemnity for disparities caused by the tool's design, separate from your own configuration choices.
And a technical disclosure, in writing: every model in the chain (resume parser, ranker, video or voice scoring, any LLM interviewer); the features each one uses; whether the vendor tested proxies for protected classes such as zip code, names, graduation year and employment gaps; and what data its own audit ran on, pooled clients, synthetic candidates, or customers like you.
How to Run the Annual Audit as a Process
Treat the audit as a recurring control with an owner, a calendar date and a file. The steps:
- Inventory. List every tool that scores, ranks or screens candidates for NYC-tied roles. The FAQ excludes tools that only scan a resume bank or do outreach before someone applies.
- Confirm independence. The auditor must not work for you or the vendor, must not have been involved in using, developing or distributing the tool, and must have no direct or material indirect financial interest in either. DCWP keeps no approved list.
- Pull historical data first. Use self-reported demographics. The FAQ forbids imputed or inferred race or sex. Test data is the fallback only when your history is too thin, and the summary has to say why.
- Handle small groups openly. A category under 2% of the data may be excluded; every other category must be reported. DCWP sets no significance threshold, so ask the auditor to note where a ratio rests on a small count rather than drop it.
- Publish and notice. Post the summary and distribution date, then keep the 10-business-day notice running.
- Decide on the results. Write down what you did about any ratio under 0.80, even if the answer is a validation study and no change. California treats that response as part of the defense.
This Week: Find every screening, ranking or assessment tool touching NYC-tied roles, and pull the audit summary for each from the vendor and from your own careers site. Check the dates: anything older than a year is already out of window.
This Month: Ask each vendor, in writing, what data its audit used and whether yours was in it. Request quotes from DCI and one other auditor for a direct engagement on your own historical data, and confirm your ATS can export per-applicant scores with self-reported demographics.
Before Your Next Renewal: Add the six contract clauses above and the technical disclosure to the renewal paper. Set retention to four years for California roles and put the annual audit date on the compliance calendar.
The Bottom Line
Local Law 144 was written so that the company making the hiring decision carries the audit, and the published summaries now make it easy to see who did the work on their own data. Pfizer's posted HiredScore audit runs on its own NYC applicants. The vendor reports run on vendor-chosen data, sometimes synthetic. For a tool you have used for a year, buy the first kind, from an auditor who will compute the ratios from your file, and keep the vendor's report as diligence.
Continue Reading
- Colorado Gutted Its AI Law 46 Days Before Enforcement.
- Anthropic's Nov 12 Policy Pulls Internal Hiring AI Into Review
- EU AI Act Governance Tools: Buy Inventory, Not Policy Packs
- Human-in-the-Loop for AI Agents: Most Approval Gates Rubber-Stamp
- Agent Audit Logging: Platform Logs Miss the Decision You Must Prove
- Fetcher Shuts Down October 16 After Juicebox Buys Its Assets
