← Research
Research

Is it ethical to let AI judge people?

The argument is conducted about accuracy. The obligations that matter survive being correct.

Last reviewed: 10 September 2026

Four questions decide it: can the decision be explained to the person it lands on, can it be contested by somebody able to change it, does a named human carry it, and was the system checked on the population it is used on. Accuracy is the easiest to measure, which is how it came to stand in for the rest.

Questions this page answersAll 811 questions this research covers

Almost every version of this argument is conducted about accuracy. One side says the system is more consistent than the people it replaces, the other says it makes mistakes, and both proceed as though the ethics turn on who is right more often. They do not. A decision about a person carries obligations that survive being correct: that somebody can say why, that somebody can be argued with, and that somebody carries the consequence of being wrong. Accuracy is the easiest of the four to measure, which is how it came to stand in for the rest.

The answer, in one line

It turns on four obligations that are usually collapsed into one. Whether the decision can be explained to the person it lands on. Whether it can be contested by somebody with power to change it.

Share as a card

The short answer#

It depends on four things that are usually collapsed into one. Whether the decision can be explained to the person it lands on. Whether it can be contested by somebody with the power to change it. Whether a named human carries responsibility for it. And whether the system was checked on the population it is being used on. A system that fails those can be highly accurate and still indefensible, and a system that meets them is defensible even when it errs, because errors were expected and provided for.

Regulators have converged on roughly this list, which is a useful signal. It is arrived at from different legal traditions and lands in the same place.

Accuracy is real, and only the first test#

The case for algorithmic judgement is not frivolous. Human decisions about people are inconsistent in ways that are documented and unflattering, and a system applies the same rule to everyone. That is a genuine ethical gain, and dismissing it is as lazy as ignoring the losses.

The trouble is what happens when the accuracy claim goes unchecked. An external validation of a widely deployed sepsis prediction model found an area under the curve of 0.63 against the 0.76 to 0.83 cited by its developer. At the alerting threshold in clinical use, sensitivity was 33 per cent and positive predictive value 12 per cent. It failed to identify 1,709 of 2,552 patients with sepsis. This is one model at one health system, and the authors say so. It had nonetheless been implemented widely on the strength of the developer's own numbers.

So even the easy criterion is not being met, and the reason is structural: the party with the strongest interest in the accuracy figure is usually the only party who has measured it.

The first obligation: an explanation the person can use#

In February 2025 the Court of Justice of the European Union decided Dun and Bradstreet Austria and set out what an explanation of an automated decision has to contain. A controller must describe the procedure and the principles actually applied, so the person can understand which of their data was used and how it fed the outcome. Disclosing the algorithm does not satisfy this, and a blanket refusal on trade-secret grounds is not permitted.

Read that as an ethical standard rather than a legal one and it is demanding. The test is not whether an explanation exists somewhere in the vendor's documentation. It is whether the person affected can follow it well enough to know whether it was applied to them correctly. Most deployed systems would fail that test today, and most organisations have never asked the question in that form.

The second: somebody who can be argued with#

An explanation with no route to challenge is a courtesy. The estate has a name for what happens when the route exists on paper and not in practice, borrowed from the literature rather than coined here: the moral crumple zone, where a human is positioned close enough to the decision to absorb blame and given neither the authority nor the information to have changed it.

The practical test is the one this estate keeps returning to and almost no organisation can answer: not whether an override exists, but when it was last used, by whom, and what happened to that person afterwards. A control nobody has exercised has not been shown to work.

The third: the checking nobody does#

The evidence on whether anyone verifies these systems is the most deflating part of the picture, and it comes from the jurisdiction that tried hardest.

New York City's Local Law 144 requires an independent bias audit of an automated employment decision tool before it is used, publication of the summary, and notice to candidates. A field audit in which 155 investigators posed as job seekers across 391 employers found about 5 per cent had published an audit. The New York State Comptroller then examined the enforcing department: reviewing the same 32 companies, the department had found one instance of non-compliance and the state auditors found at least seventeen.

The ethical weight of this is easy to miss. An organisation deploying a system to judge people is not choosing between a checked machine and an unchecked human. On the available evidence it is usually choosing between two unchecked things, one of which operates at scale.

Where the line has actually been drawn#

Legislators have not answered the ethical question in the abstract. They have answered a narrower one, by naming the decisions where the obligations bite hardest. Annex III of the EU AI Act classifies as high risk, among others, systems determining access to education, evaluating learning outcomes where those outcomes steer a learner's path, and a list of employment decisions.

That is a workable starting point for an organisation with no view of its own: the decisions that shape what a person is permitted to become. Classification is not evidence that any particular system is unsafe, and the Annex says nothing about how well the resulting duties are met. What it offers is a defensible answer to which decisions deserve the full apparatus, from a body that had to write one down.

Separating the four obligations before the argument starts#

The question of principle is left open#

It does not answer the question as a matter of principle, nor try to. Whether it is ever right to let a machine decide something about a person is a moral question that evidence cannot close, and people who hold that some decisions require a human regardless of accuracy are making an argument this page does not defeat.

The four obligations also carry a cost that is rarely priced. Explanation, contest and named accountability slow decisions down and consume the labour the system was bought to save. An organisation that adopts the full apparatus honestly may find the business case disappears, and nobody in this literature has been willing to say so plainly.

Evidence review · SS-2026-212 · Graded against the published rubric

Cite this page

Hirji, R. (2026). Is it ethical to let AI judge people?. The SuperSkills evidence base, SS-2026-212. https://thesuperskills.com/research/is-it-ethical-to-let-ai-judge-people. Last reviewed 10 September 2026.

An evidence review by Rahim Hirji, not peer-reviewed research. For a material claim, cite the underlying study as well; every study here carries its own permanent link.

How citations and IDs work
Questions answered on this page

Is it ethical to let AI judge people?

It turns on four obligations that are usually collapsed into one. Whether the decision can be explained to the person it lands on. Whether it can be contested by somebody with power to change it. Whether a named human carries responsibility. And whether the system was checked on the population it is used on. A system failing those can be highly accurate and still indefensible; a system meeting them is defensible even when it errs, because error was expected and provided for.

Is AI more consistent than a human decision-maker?

Often, and that is a genuine ethical gain rather than a talking point. The difficulty is that the accuracy claim usually goes unchecked. An external validation of a widely deployed sepsis prediction model found an area under the curve of 0.63 against the 0.76 to 0.83 cited by its developer, with 33 per cent sensitivity and 12 per cent positive predictive value at the threshold in clinical use, failing to identify 1,709 of 2,552 patients with sepsis. It had been implemented widely on the developer's own figures.

What counts as a proper explanation of an automated decision?

The Court of Justice of the European Union decided this in Dun and Bradstreet Austria in February 2025. A controller must describe the procedure and principles actually applied, so the person can understand which of their data was used and how it produced the outcome. Disclosing the algorithm is not sufficient, and a blanket trade-secret refusal is not permitted. As a practical test: can the affected person follow it well enough to know whether it was applied to them correctly.

Which decisions about people are treated as high risk?

Annex III of the EU AI Act names them, including systems determining access or admission to education, evaluating learning outcomes where those outcomes steer a learner's path, and a list of employment decisions. The organising idea is decisions that shape what a person is permitted to become. Classification is not evidence that any particular system is unsafe, and the Annex says nothing about how well the resulting duties are met in practice.

Does anyone actually audit systems that judge people?

Rarely, on the best evidence available. Under New York City's Local Law 144, a field audit across 391 employers found roughly 5 per cent had published a required bias audit. The New York State Comptroller then reviewed the enforcing department and found it had identified one instance of non-compliance among 32 companies where the state auditors identified at least seventeen. An organisation is usually not choosing between a checked machine and an unchecked human, but between two unchecked things, one of which operates at scale.

In this hub

Judgement, oversight and accountability

Who decides, who checks, and who is answerable when the machine was involved.

Ask the evidence
What does the evidence actually show?What should our board be asking about this?Where does Rahim disagree with the consensus?
Bring this into your organisation

If this describes something happening in your teams, say so.

Keynotes, board sessions and advisory work, drawing on research across more than 200 organisations in 30 countries. Tell me the room, the date and the shift you need. A reply within 24 hours.

Start a conversation

Topics and audiences  ·  All research

Box of Amazing

Rahim’s free weekly letter on AI and human capability

If this was useful, the weekly letter is where the thinking happens first. Most of what ends up on this site starts there. Weekly essays on AI, capability and the future of work. Read by 25,000 people, every week since 2017. Free, and one click to stop.

Opens Substack to confirm. No pitch in it, unsubscribe in one click, and nobody follows up because you read something.

Running an event, or responsible for how AI arrives in your organisation? Keynotes  ·  Advisory for CEOs and boards  ·  Enquire