In most jurisdictions, yes, with conditions attached that get stricter the closer the system comes to the rating itself. In the European Union the conditions are written into primary legislation, and a system used to evaluate the performance of employees is named as high risk in the text of the law rather than left to interpretation. The harder question sits underneath the legal one. A performance review is a manager's judgement about a person, recorded in a form that will later be used against or for them, and the thing that makes it a judgement is that somebody formed it.
The answer, in one line
In most jurisdictions yes, with conditions that tighten the closer the system gets to the rating.
Definition#
AI in performance reviews: any use of an AI system in the assessment of an employee, from drafting the write-up to scoring the rating. European law draws no line between those two: Annex III of Regulation (EU) 2024/1689 classifies as high risk any AI system intended to be used to monitor and evaluate the performance and behaviour of persons
in a work-related relationship. The obligations attach to the function, not to how impressive the tool is.
Europe already wrote this application into the statute#
Annex III point 4(b) covers AI systems intended to be used to make decisions affecting terms of work-related relationships, the promotion or termination of work-related contractual relationships, to allocate tasks based on individual behaviour or personal traits or characteristics or to monitor and evaluate the performance and behaviour of persons in such relationships
. Read at source on 23 September 2026. An appraisal system sits inside the last clause, and a system that feeds into promotion decisions sits inside the second.
Three duties fall on the employer as deployer, under Article 26. Paragraph 2 requires that human oversight be assigned to natural persons who have the necessary competence, training and authority
, along with the necessary support, so a named person with the standing to overrule the system rather than a committee that receives its output. Paragraph 7 requires that before putting such a system into service at the workplace, employers shall inform workers' representatives and the affected workers that they will be subject to the use of the high-risk AI system
. Paragraph 11 requires that the people subject to decisions made or assisted by an Annex III system be told they are.
On timing, the estate's position is narrower than most guidance. December 2027 is a longstop for Annex III rather than a start date: the Digital Omnibus ties the high-risk rules to the availability of standards and caps the delay at sixteen months for Annex III areas, so the obligations may begin sooner and a board planning to the outside date is planning to the wrong end of a range. Any advice saying the rules apply from December 2027 at the earliest
has the direction reversed.
Outside Europe the picture is patchier and the direction is the same. New York City's Local Law 144 has required an annual independent bias audit of an automated employment decision tool since 1 January 2023, with publication of the audit summary and notice to candidates, and its implementing rules reach promotion as well as hiring.
The first city to try it produced almost no disclosure at all#
Local Law 144 is the best natural experiment available on what one of these rules does once it exists, because it has been running for three years and two independent bodies have gone and looked.
Wright and colleagues organised 155 student investigators to record what 391 employers published, acting as model job seekers. They found 18 bias audit reports and 13 transparency notices. Their explanation is structural rather than a charge of indifference: because employers decide for themselves whether a tool is covered, non-compliance cannot be established from outside, so the observable state is neither compliance nor breach. They call it null compliance. They also record that the law requires an audit and says nothing about its results, sets no discrimination threshold, and offers no protection to an employer whose disclosure creates liability elsewhere.
Two years later the New York State Comptroller audited the enforcing department. The Department of Consumer and Worker Protection had reviewed 32 employers' and vendors' websites and identified one instance of likely non-compliance; the Comptroller's auditors reviewed the same companies and found 17 instances of potential non-compliance. The auditors also placed twelve test calls to the city's 311 service to file a complaint: three were correctly routed to the department, eight went to the New York State Department of Labor, and one was sent back to the employer. The department disagreed with some findings and agreed to adopt most of the thirteen recommendations.
For an employer the practical reading concerns detection rather than teeth. Nothing external will tell you whether your appraisal system is defensible. In the only place where anybody has looked twice, the second look found seventeen times what the enforcing body found, and the applicants had no working route to complain at all.
The health evidence comes from lorries and warehouses#
Software that instructs, monitors and evaluates workers has been measured against worker health in exactly one setting, and that setting is a depot. Hennum Nilsson and six colleagues surveyed 978 Swedish logistics workers between February and July 2024, 592 drivers and 378 warehouse staff, using an eleven-item exposure scale covering task allocation, surveillance and performance monitoring. Higher exposure was associated with psychological distress at a prevalence ratio of 2.12 (95% CI 1.49 to 3.02), occupational accidents at 1.92 (1.22 to 3.01), headaches at 1.68 (1.09 to 2.58) and musculoskeletal pain at 1.54 (1.23 to 1.92), in adjusted models. Associations were stronger among drivers.
Three limits belong with those numbers. The study is cross-sectional, so workers under strain may also perceive their management as more algorithmic. Everything is self-reported. And participants were recruited through social-media advertising, which makes this a self-selected sample rather than a representative one. The authors present it as evidence of association and nothing more.
What it establishes for a head of HR is a floor rather than a forecast. Where a system has been put between a manager and a worker at scale, somebody has now found it in the bodies of the people underneath it. Nobody has run the equivalent study on knowledge workers, and nobody should assume the result transfers.
The measurement a buyer wants does not exist#
No study anywhere establishes whether an AI-assisted performance review is fairer, less fair, more accurate or less accurate than one written unaided. There is no trial, no field experiment and no audit of outcomes. Vendors in this market sell on time saved and on consistency of tone, and both of those are real and neither is the thing being claimed when a system is described as reducing bias.
The adjacent literature runs in an unhelpful direction. The estate's homogenisation evidence shows that assistance makes different people produce more similar output, and sameness across reviews is the thing an appraisal process is built to prevent: a set of reviews that read consistently because a model wrote them is indistinguishable, from the outside, from a set of reviews that are consistent because the managers agreed. Consistency of prose has been mistaken for consistency of judgement before.
Writing the review was how managers learned to manage#
Sitting down once or twice a year to work out what somebody is actually good at, and to find words for the thing they keep getting wrong, is one of the few structured occasions on which a manager is forced to form a view. It is slow, most people dislike it, and the difficulty is load-bearing. Delegating the draft removes the occasion.
Two things follow from that, both of which this research covers elsewhere. The manager accumulates capability debt, because the practice of forming an assessment has been retired while the accountability for it has not. And the work that remains, reading the machine's draft and deciding whether it is true, is verification, which organisations reliably under-value even where it is harder than the writing. That is the verifier's discount, and in an appraisal it arrives with legal exposure attached, because the person who signed it owns it.
The failure mode to watch for is the one where a manager retains formal sign-off, exercises no real discretion over the content, and is nonetheless the person named when the review is challenged. The estate's term for that arrangement is the moral crumple zone, and a performance review is an unusually clean example, because the document is durable, contested and used in evidence.
Five tests before it goes near a rating#
- Separate the draft from the score. A tool that helps a manager phrase a view they have already formed is a different system from one that produces the view. European law reaches both, and only the second changes who is doing the assessing.
- Name the overseer. Article 26(2) asks for a person with competence, training and authority. If nobody can say who would overturn the system's output, and whether they are senior enough to do it without cost, the oversight is nominal.
- Tell people, in advance and in writing. Paragraph 7 makes this a duty towards workers' representatives and affected workers before the system goes live; paragraph 11 makes it a duty towards the individual. Both are cheap to do and hard to retrofit after a grievance.
- Ask the vendor what its evidence measures. Time saved and consistency of tone are measurable and are usually what has been measured. Fairness and accuracy have not been measured by anybody, so a claim about either is a claim without a study behind it.
- Keep one review a year written from scratch. That is how you find out whether the manager can still do it, using the same withdrawal test every measurement in this area has relied on.
What this page cannot tell you#
The legal position above is Europe and one American city. Employment law is national, and nothing here substitutes for advice on the jurisdiction where the employee sits. The Local Law 144 findings describe hiring tools more than appraisal tools, because that is where the disclosure obligations bite, so their transfer to annual reviews is an argument rather than a measurement. The Swedish logistics study is the only health evidence and it is about a different kind of work. And the central claim on this page, that drafting the review is where a manager builds the judgement the review depends on, follows from the withdrawal literature rather than from any study of managers. Nobody has tested it.
Key sources
- European Parliament and Council of the European Union (2024). Regulation (EU) 2024/1689, Annex III: High-Risk AI Systems Referred to in Article 6(2). Point 4(b) read at source 23 September 2026, in the EU AI Act Explorer rendering of the same official text. Graded entry.
- European Union (2024). Regulation (EU) 2024/1689, Article 26: Obligations of Deployers of High-Risk AI Systems. Paragraphs 2, 7 and 11 read at source 23 September 2026. Graded entry.
- European Commission, AI Act Service Desk (2026). Article 113: Entry into force and application, with the Digital Omnibus on AI amendments. Graded entry.
- Council of the City of New York; Department of Consumer and Worker Protection (2021). Local Law 144 of 2021, automated employment decision tools. In force 1 January 2023. Graded entry.
- Wright, L., Muenster, R. M., Vecchione, B., Qu, T., Cai, P., Smith, A., COMM/INFO 2450 Student Investigators, Metcalf, J. and Matias, J. N. (2024). Null Compliance: NYC Local Law 144 and the Challenges of Algorithm Accountability. FAccT '24, Rio de Janeiro. DOI 10.1145/3630106.3658998. Methods and Table 1 read at source 23 September 2026 in the authors' arXiv version, 2406.01399. Graded entry.
- Office of the New York State Comptroller (2025). Department of Consumer and Worker Protection: Enforcement of Local Law 144. Audit 2024-N-6, issued 2 December 2025, covering July 2023 to June 2025. Figures confirmed in the Comptroller's own release of the same date. Graded entry.
- Hennum Nilsson, K., Bodin, T., Strauss, P., Matilla-Santander, N., Badarin, K., Brulin, E. and Hakansta, C. (2025). Algorithmic management is associated with psychological distress, musculoskeletal pain, and occupational accidents: a cross-sectional study in logistics. International Archives of Occupational and Environmental Health, 98. PMID 41251732. Full text read at source 23 September 2026. Graded entry.
Related SuperSkills research#
On the principle underneath the legal question, is it ethical to let AI judge people and what meaningful human oversight requires. For the function as a whole, what CHROs should do about AI and how AI changes human resources. On bias, can AI be unbiased. On what the manager loses by delegating the drafting, the missed reps and capability debt. On who answers when the assessment is wrong, the moral crumple zone and the invisible work of oversight. On sameness of output, does AI make everyone think alike.
Evidence review · SS-2026-294 · Graded against the published rubric
Hirji, R. (2026). Can I use AI for performance reviews?. The SuperSkills evidence base, SS-2026-294. https://thesuperskills.com/research/can-i-use-ai-for-performance-reviews. Last reviewed 23 September 2026.
An evidence review by Rahim Hirji, not peer-reviewed research. For a material claim, cite the underlying study as well; every study here carries its own permanent link.
How citations and IDs work