← Research
Research · Question

Can I use AI for performance reviews?

The law already names this use. What nobody has measured is whether the review comes out better.

Last reviewed: 23 September 2026 · Next review due: 23 September 2027

What Annex III point 4(b) covers and what Article 26 requires of an employer, what New York learned in three years of trying to enforce the first law of this kind, the only health evidence there is and why it comes from logistics, and what a manager stops practising when the drafting goes to a machine.

Question this page answersQuestion this page partly answersAll 887 questions this research covers

In most jurisdictions, yes, with conditions attached that get stricter the closer the system comes to the rating itself. In the European Union the conditions are written into primary legislation, and a system used to evaluate the performance of employees is named as high risk in the text of the law rather than left to interpretation. The harder question sits underneath the legal one. A performance review is a manager's judgement about a person, recorded in a form that will later be used against or for them, and the thing that makes it a judgement is that somebody formed it.

The answer, in one line

In most jurisdictions yes, with conditions that tighten the closer the system gets to the rating.

Share as a card

Definition#

AI in performance reviews: any use of an AI system in the assessment of an employee, from drafting the write-up to scoring the rating. European law draws no line between those two: Annex III of Regulation (EU) 2024/1689 classifies as high risk any AI system intended to be used to monitor and evaluate the performance and behaviour of persons in a work-related relationship. The obligations attach to the function, not to how impressive the tool is.

Share this definition as a card

Europe already wrote this application into the statute#

Annex III point 4(b) covers AI systems intended to be used to make decisions affecting terms of work-related relationships, the promotion or termination of work-related contractual relationships, to allocate tasks based on individual behaviour or personal traits or characteristics or to monitor and evaluate the performance and behaviour of persons in such relationships. Read at source on 23 September 2026. An appraisal system sits inside the last clause, and a system that feeds into promotion decisions sits inside the second.

Three duties fall on the employer as deployer, under Article 26. Paragraph 2 requires that human oversight be assigned to natural persons who have the necessary competence, training and authority, along with the necessary support, so a named person with the standing to overrule the system rather than a committee that receives its output. Paragraph 7 requires that before putting such a system into service at the workplace, employers shall inform workers' representatives and the affected workers that they will be subject to the use of the high-risk AI system. Paragraph 11 requires that the people subject to decisions made or assisted by an Annex III system be told they are.

On timing, the estate's position is narrower than most guidance. December 2027 is a longstop for Annex III rather than a start date: the Digital Omnibus ties the high-risk rules to the availability of standards and caps the delay at sixteen months for Annex III areas, so the obligations may begin sooner and a board planning to the outside date is planning to the wrong end of a range. Any advice saying the rules apply from December 2027 at the earliest has the direction reversed.

Outside Europe the picture is patchier and the direction is the same. New York City's Local Law 144 has required an annual independent bias audit of an automated employment decision tool since 1 January 2023, with publication of the audit summary and notice to candidates, and its implementing rules reach promotion as well as hiring.

The first city to try it produced almost no disclosure at all#

Local Law 144 is the best natural experiment available on what one of these rules does once it exists, because it has been running for three years and two independent bodies have gone and looked.

Wright and colleagues organised 155 student investigators to record what 391 employers published, acting as model job seekers. They found 18 bias audit reports and 13 transparency notices. Their explanation is structural rather than a charge of indifference: because employers decide for themselves whether a tool is covered, non-compliance cannot be established from outside, so the observable state is neither compliance nor breach. They call it null compliance. They also record that the law requires an audit and says nothing about its results, sets no discrimination threshold, and offers no protection to an employer whose disclosure creates liability elsewhere.

Two years later the New York State Comptroller audited the enforcing department. The Department of Consumer and Worker Protection had reviewed 32 employers' and vendors' websites and identified one instance of likely non-compliance; the Comptroller's auditors reviewed the same companies and found 17 instances of potential non-compliance. The auditors also placed twelve test calls to the city's 311 service to file a complaint: three were correctly routed to the department, eight went to the New York State Department of Labor, and one was sent back to the employer. The department disagreed with some findings and agreed to adopt most of the thirteen recommendations.

For an employer the practical reading concerns detection rather than teeth. Nothing external will tell you whether your appraisal system is defensible. In the only place where anybody has looked twice, the second look found seventeen times what the enforcing body found, and the applicants had no working route to complain at all.

The health evidence comes from lorries and warehouses#

Software that instructs, monitors and evaluates workers has been measured against worker health in exactly one setting, and that setting is a depot. Hennum Nilsson and six colleagues surveyed 978 Swedish logistics workers between February and July 2024, 592 drivers and 378 warehouse staff, using an eleven-item exposure scale covering task allocation, surveillance and performance monitoring. Higher exposure was associated with psychological distress at a prevalence ratio of 2.12 (95% CI 1.49 to 3.02), occupational accidents at 1.92 (1.22 to 3.01), headaches at 1.68 (1.09 to 2.58) and musculoskeletal pain at 1.54 (1.23 to 1.92), in adjusted models. Associations were stronger among drivers.

Three limits belong with those numbers. The study is cross-sectional, so workers under strain may also perceive their management as more algorithmic. Everything is self-reported. And participants were recruited through social-media advertising, which makes this a self-selected sample rather than a representative one. The authors present it as evidence of association and nothing more.

What it establishes for a head of HR is a floor rather than a forecast. Where a system has been put between a manager and a worker at scale, somebody has now found it in the bodies of the people underneath it. Nobody has run the equivalent study on knowledge workers, and nobody should assume the result transfers.

The measurement a buyer wants does not exist#

No study anywhere establishes whether an AI-assisted performance review is fairer, less fair, more accurate or less accurate than one written unaided. There is no trial, no field experiment and no audit of outcomes. Vendors in this market sell on time saved and on consistency of tone, and both of those are real and neither is the thing being claimed when a system is described as reducing bias.

The adjacent literature runs in an unhelpful direction. The estate's homogenisation evidence shows that assistance makes different people produce more similar output, and sameness across reviews is the thing an appraisal process is built to prevent: a set of reviews that read consistently because a model wrote them is indistinguishable, from the outside, from a set of reviews that are consistent because the managers agreed. Consistency of prose has been mistaken for consistency of judgement before.

Writing the review was how managers learned to manage#

Sitting down once or twice a year to work out what somebody is actually good at, and to find words for the thing they keep getting wrong, is one of the few structured occasions on which a manager is forced to form a view. It is slow, most people dislike it, and the difficulty is load-bearing. Delegating the draft removes the occasion.

Two things follow from that, both of which this research covers elsewhere. The manager accumulates capability debt, because the practice of forming an assessment has been retired while the accountability for it has not. And the work that remains, reading the machine's draft and deciding whether it is true, is verification, which organisations reliably under-value even where it is harder than the writing. That is the verifier's discount, and in an appraisal it arrives with legal exposure attached, because the person who signed it owns it.

The failure mode to watch for is the one where a manager retains formal sign-off, exercises no real discretion over the content, and is nonetheless the person named when the review is challenged. The estate's term for that arrangement is the moral crumple zone, and a performance review is an unusually clean example, because the document is durable, contested and used in evidence.

Five tests before it goes near a rating#

What this page cannot tell you#

The legal position above is Europe and one American city. Employment law is national, and nothing here substitutes for advice on the jurisdiction where the employee sits. The Local Law 144 findings describe hiring tools more than appraisal tools, because that is where the disclosure obligations bite, so their transfer to annual reviews is an argument rather than a measurement. The Swedish logistics study is the only health evidence and it is about a different kind of work. And the central claim on this page, that drafting the review is where a manager builds the judgement the review depends on, follows from the withdrawal literature rather than from any study of managers. Nobody has tested it.

Key sources

On the principle underneath the legal question, is it ethical to let AI judge people and what meaningful human oversight requires. For the function as a whole, what CHROs should do about AI and how AI changes human resources. On bias, can AI be unbiased. On what the manager loses by delegating the drafting, the missed reps and capability debt. On who answers when the assessment is wrong, the moral crumple zone and the invisible work of oversight. On sameness of output, does AI make everyone think alike.

Evidence review · SS-2026-294 · Graded against the published rubric

Cite this page

Hirji, R. (2026). Can I use AI for performance reviews?. The SuperSkills evidence base, SS-2026-294. https://thesuperskills.com/research/can-i-use-ai-for-performance-reviews. Last reviewed 23 September 2026.

An evidence review by Rahim Hirji, not peer-reviewed research. For a material claim, cite the underlying study as well; every study here carries its own permanent link.

How citations and IDs work
Questions answered on this page

Can I use AI for performance reviews?

In most jurisdictions yes, with conditions that tighten the closer the system gets to the rating. In the European Union, Annex III point 4(b) of Regulation (EU) 2024/1689 classifies as high risk any AI system intended to monitor and evaluate the performance and behaviour of people in a work-related relationship, which brings documentation, human oversight and information duties with it. Employment law is national, so the position where your employee sits is the one that governs.

Is using AI to write a performance review legal?

European law does not distinguish between drafting the write-up and producing the score: Annex III reaches any AI system used to evaluate performance. Under Article 26 an employer deploying one must assign oversight to a natural person with the necessary competence, training and authority, must inform workers' representatives and the affected workers before putting it into service at the workplace, and must tell the individual that they are subject to it. None of that makes the use unlawful; it makes it a regulated activity with named obligations.

Does AI make performance reviews fairer?

Nobody has measured it. There is no trial, field experiment or outcome audit establishing whether an AI-assisted performance review is fairer, less fair, more accurate or less accurate than one written unaided. Vendors sell on time saved and consistency of tone, both of which are real and neither of which is fairness. The adjacent evidence on homogenisation points the other way, since assistance makes different people produce more similar output, and consistency of prose is easy to mistake for consistency of judgement.

Do I have to tell employees that AI was used in their review?

In the European Union, yes, on two separate counts for a high-risk system. Article 26(7) requires an employer, before putting the system into service at the workplace, to inform workers' representatives and the affected workers that they will be subject to it. Article 26(11) requires deployers of Annex III systems that make or assist decisions about people to inform those people that they are subject to the system. In New York City, Local Law 144 requires notice to candidates and employees for automated employment decision tools, together with an annual independent bias audit whose summary is published.

What is the risk of letting AI draft appraisals?

Two risks that are not legal ones. The manager stops practising the thing the review is for, which is forming and defending an assessment of a person, and that capability erodes while the accountability for it does not. And the manager can end up signing a document they did not really form a view on, which leaves them named when it is challenged. A durable, contested document used in evidence is an unusually clean example of a moral crumple zone.

In this hub

Organisations and leadership

What a leadership team actually has to decide, and what to measure.

Ask the evidence
What does the evidence actually show?What should our board be asking about this?Where does Rahim disagree with the consensus?
Bring this into your organisation

If this describes something happening in your teams, say so.

Keynotes, board sessions and advisory work, drawing on research across more than 200 organisations in 30 countries. Tell me the room, the date and the shift you need. A reply within 24 hours.

Start a conversation

Topics and audiences  ·  All research

Box of Amazing

Rahim’s free weekly letter on AI and human capability

If this was useful, the weekly letter is where the thinking happens first. Most of what ends up on this site starts there. Weekly essays on AI, capability and the future of work. Read by 25,000 people, every week since 2017. Free, and one click to stop.

Opens Substack to confirm. No pitch in it, unsubscribe in one click, and nobody follows up because you read something.

Running an event, or responsible for how AI arrives in your organisation? Keynotes  ·  Advisory for CEOs and boards  ·  Enquire