Awkwardly, because the function being asked to design the organisation's response to AI is itself heavily automated in its core task. Screening, drafting and first-line queries are the three things HR spends most of its hours on and the three things models do cheapest. The best available audit of resume screening found retrieval models favouring White-associated names in 85.1 per cent of comparisons and female-associated names in 11.1 per cent. The world's first law requiring bias audits of hiring tools produced published audit reports from 18 of 391 employers checked. And European law now classifies recruitment tools as high risk by name. HR is going to spend the next decade being held to a standard of evidence it has never previously been asked for.
The audit that measured the model rather than the vendor#
Wilson and Caliskan, presented at the AAAI/ACM Conference on AI, Ethics and Society in 2024, built a document retrieval framework that simulates candidate selection and ran a resume audit study through it. Over 500 publicly available resumes and 500 job descriptions across nine occupations, with 120 first names associated with male, female, Black and white candidates.
The Massive Text Embedding models tested favoured White-associated names in 85.1 per cent of cases and female-associated names in 11.1 per cent, with a minority of comparisons showing no statistically significant difference. Black male candidates were disadvantaged in up to 100 per cent of cases. The study also found document length and the corpus frequency of a name affecting selection, which means part of the bias is an artefact of how common a name is in the training data rather than anything about the candidate.
Two limits, both of which get dropped in the reporting. These are embedding models used for retrieval, not the chat assistants most people picture, and not the proprietary systems any particular vendor sells. What the study establishes is that the underlying representation layer, the thing almost every commercial tool is built on, carries the pattern before anyone adds a business rule on top. What it does not establish is the behaviour of a named product in a live pipeline, because nobody outside those vendors can test one.
Eighteen of three hundred and ninety-one#
New York City's Local Law 144, in force from July 2023, was the first law anywhere requiring commercial algorithmic hiring tools to be independently audited for race and gender bias each year, with the audit report posted publicly and a transparency notice posted with the job listing.
Wright, Muenster, Vecchione, Qu, Cai, Smith, Metcalf and Matias, with 155 student investigators acting as model job seekers, checked 391 employers. 18 posted an audit report, roughly 5 per cent. 13 posted a transparency notice, roughly 3 per cent. The paper names the resulting condition null compliance: a state in which non-compliance cannot be established, because the law's own design makes it impossible to tell whether an employer is using a covered tool at all.
The mechanism is the part HR should read twice. Local Law 144 requires an audit and says nothing about its results. It sets no discrimination threshold, not even the four-fifths convention that governs disparate impact in US employment practice, and it offers no guidance on remediation if an audit finds one. No federal body has established a safe harbour for audits conducted under a local law. So an employer who publishes an audit showing disparate impact has produced evidence for a different regulator with jurisdiction it would not otherwise have. The rational legal advice is to stay quiet, and 95 per cent of the employers checked did.
Any organisation designing an internal AI assurance process should take the lesson rather than the statute. An audit obligation with no threshold, no remediation duty and no protection for disclosure produces silence. That is a design failure rather than a compliance failure, and it will repeat inside companies exactly as it repeated in New York.
European law has already named the tools#
Annex III of Regulation (EU) 2024/1689 lists high-risk AI systems. Point 4 covers employment, workers' management and access to self-employment, in wording that leaves little room:
AI systems intended to be used for the recruitment or selection of natural persons, in particular to place targeted job advertisements, to analyse and filter job applications, and to evaluate candidates.
Point 4(b) extends to decisions on promotion and termination, task allocation based on individual behaviour or personal traits, and monitoring or evaluating performance. High-risk classification pulls in the requirements of Chapter III, including the risk management system of Article 9, the record-keeping of Article 12, the human oversight of Article 14 and, under Article 86, a right for an affected person to an explanation of an individual decision.
Set that against the New York finding and the shape of the next few years becomes visible. One regime asked for publication and got 5 per cent. The other asks for documentation, oversight design and explanation, and attaches them to the provider and the deployer rather than to a public posting. HR functions in scope will be producing the evidence whether or not anyone reads it, which is a materially different obligation from a website disclosure.
The function is exposed exactly where it advises#
The Cabinet Office has a published algorithmic transparency record for the verbal and numerical tests used to sift civil service applicants. The Department for Work and Pensions has published records for tools that read scanned citizen correspondence at volume. Whatever a given HR function believes about its AI position, systems that sort people are already in production and, in the UK public sector at least, some of them are documented in public.
The private sector equivalent is undocumented, and the tasks are the same ones HR owns: sifting applications, drafting job descriptions and offer letters, summarising interview notes, answering policy questions, writing performance narratives, running engagement analysis. Every one of those is a text task, and the text tasks are the ones that went first everywhere else. In the cross-government Copilot experiment, HR was among the professions where users flagged the sharpest reservations, with one participant warning that in grievance handling or performance evaluation an inaccurate output carries reputational risk.
Which produces the specific difficulty this page exists to name. HR is the function that would ordinarily run an organisation's response to capability loss, design the assessment that detects it and own the development that repairs it. It is also, on task composition, one of the functions most exposed to the substitution that causes it. A function running its own workload down through automation while advising the board on capability debt is in an uncomfortable position, and the discomfort is not a reason to look away from it.
Screening was never good, and that is the wrong defence#
A standard argument for automated sifting runs that human screening is biased too, so a model that is measurably biased is at least measurably so. Half of that is right. The unstructured interview has one of the weakest validity records in personnel selection, and the evidence for structured methods over intuition is old and strong.
The half that fails is about scale and correlation. A hundred human sifters produce a hundred partly independent errors. One retrieval model produces one error, applied identically to every applicant, at a volume no human process ever reached. The failure mode changes from noisy to systematic, and systematic failure is both more harmful and, on the New York evidence, less likely to be published. This research makes the same argument about population-level convergence at does AI make everyone think alike: individual-level improvement can conceal a collective cost that no individual measurement will show.
What nobody has established#
- No study has measured who actually got hired. Every finding above audits a model or a disclosure regime. Whether AI screening changes the composition of people who receive offers, at any real employer, is unmeasured, because the data sits with employers and vendors.
- The audited systems are not the deployed systems. Wilson and Caliskan tested open models in a simulated pipeline. Commercial products add filters, thresholds and business rules that could make the effect larger or smaller, and none of them can be independently tested.
- Nobody has measured HR capability under automation. There is no equivalent of the clinical deskilling finding for this function: no measurement of whether a practitioner who has drafted with a model for three years still writes a defensible dismissal letter without one.
- The four-fifths convention is not a legal threshold, and the paper that surfaced the null compliance problem says so explicitly. Any internal standard built on it is a convention, not a shield.
- The UK has no equivalent of Local Law 144. Employment screening sits under existing discrimination and data protection law, and no public register of employer bias audits exists.
Six things an HR function can do before it is asked#
- Write down every point in the people process where a model ranks, scores or filters a person. Most functions cannot currently produce this list. Any regulator or claimant asks for it first.
- Ask each vendor for the audit, and record the answer either way. A supplier who will not share one has told you something, and the refusal is worth minuting.
- Keep an unassisted sample. A proportion of applications sifted by a person, compared quarterly against the model's ranking on the same pool. It is the only method that will show drift before a tribunal does. See what is a capability audit.
- Decide who can override a rejection, and whether they ever have. Oversight that has never changed an outcome is not oversight. See why human in the loop is not a safeguard.
- Protect the drafting that builds judgement. Grievance findings, performance narratives and dismissal reasoning are where an HR practitioner learns what a defensible decision reads like. Automating the first draft of all of them is a decision about who is capable in 2036.
- Separate the efficiency case from the fairness case in every business case. They are argued with different evidence and a single AI adoption paper usually conflates them.
Key sources
- Wilson, K. and Caliskan, A. (2024). Gender, Race, and Intersectional Bias in Resume Screening via Language Model Retrieval. Proceedings of the 2024 AAAI/ACM Conference on AI, Ethics, and Society.
- Wright, L., Muenster, R. M., Vecchione, B., Qu, T., Cai, P., Smith, A., COMM/INFO 2450 Student Investigators, Metcalf, J. and Matias, J. N. (2024). Null Compliance: NYC Local Law 144 and the Challenges of Algorithm Accountability. FAccT '24.
- European Parliament and Council (2024). Regulation (EU) 2024/1689, Annex III: High-risk AI systems.
- European Parliament and Council (2024). Regulation (EU) 2024/1689, Article 14: Human Oversight.
- McDaniel, M. et al. (1994). The Validity of Employment Interviews: A Comprehensive Review and Meta-Analysis.
- Sackett, P. et al. (2023). Revisiting the design of selection systems in light of new findings regarding the validity of widely used predictors.
- Government Digital Service (2025). Microsoft 365 Copilot Experiment: Cross-Government Findings Report.
Related SuperSkills research#
On the function's own agenda, the CHRO guide to AI, AI workforce strategy and why reskilling programmes mostly fail. On assessment, assessing capability rather than output and the capability audit. On the pipeline, missing rungs, synthetic seniority and entry-level jobs. On the neighbouring sectors, the public sector and consulting.
About this research#
Rahim Hirji is the author of SuperSkills (Kogan Page, 2026), keynote speaker on AI and human capability, and founder of The SuperSkills Intelligence Company. The resume audit study, the FAccT compliance paper and the text of Annex III were each read at source. Litigation over AI hiring tools is active in the United States and is not summarised here, because the court orders could not be read at a primary source by this research and secondary legal commentary is not a substitute for the order itself.
Cite this
Hirji, R. (2026). How will AI change human resources? The SuperSkills Intelligence Company. Last reviewed 1 September 2026. thesuperskills.com/research/how-will-ai-change-human-resources
