- Is it ethical to let AI judge people?
- What makes an automated decision explainable enough?
- Which decisions about people count as high risk?
Almost every version of this argument is conducted about accuracy. One side says the system is more consistent than the people it replaces, the other says it makes mistakes, and both proceed as though the ethics turn on who is right more often. They do not. A decision about a person carries obligations that survive being correct: that somebody can say why, that somebody can be argued with, and that somebody carries the consequence of being wrong. Accuracy is the easiest of the four to measure, which is how it came to stand in for the rest.
The answer, in one line
It turns on four obligations that are usually collapsed into one. Whether the decision can be explained to the person it lands on. Whether it can be contested by somebody with power to change it.
The short answer#
It depends on four things that are usually collapsed into one. Whether the decision can be explained to the person it lands on. Whether it can be contested by somebody with the power to change it. Whether a named human carries responsibility for it. And whether the system was checked on the population it is being used on. A system that fails those can be highly accurate and still indefensible, and a system that meets them is defensible even when it errs, because errors were expected and provided for.
Regulators have converged on roughly this list, which is a useful signal. It is arrived at from different legal traditions and lands in the same place.
Accuracy is real, and only the first test#
The case for algorithmic judgement is not frivolous. Human decisions about people are inconsistent in ways that are documented and unflattering, and a system applies the same rule to everyone. That is a genuine ethical gain, and dismissing it is as lazy as ignoring the losses.
The trouble is what happens when the accuracy claim goes unchecked. An external validation of a widely deployed sepsis prediction model found an area under the curve of 0.63 against the 0.76 to 0.83 cited by its developer. At the alerting threshold in clinical use, sensitivity was 33 per cent and positive predictive value 12 per cent. It failed to identify 1,709 of 2,552 patients with sepsis. This is one model at one health system, and the authors say so. It had nonetheless been implemented widely on the strength of the developer's own numbers.
So even the easy criterion is not being met, and the reason is structural: the party with the strongest interest in the accuracy figure is usually the only party who has measured it.
The first obligation: an explanation the person can use#
In February 2025 the Court of Justice of the European Union decided Dun and Bradstreet Austria and set out what an explanation of an automated decision has to contain. A controller must describe the procedure and the principles actually applied, so the person can understand which of their data was used and how it fed the outcome. Disclosing the algorithm does not satisfy this, and a blanket refusal on trade-secret grounds is not permitted.
Read that as an ethical standard rather than a legal one and it is demanding. The test is not whether an explanation exists somewhere in the vendor's documentation. It is whether the person affected can follow it well enough to know whether it was applied to them correctly. Most deployed systems would fail that test today, and most organisations have never asked the question in that form.
The second: somebody who can be argued with#
An explanation with no route to challenge is a courtesy. The estate has a name for what happens when the route exists on paper and not in practice, borrowed from the literature rather than coined here: the moral crumple zone, where a human is positioned close enough to the decision to absorb blame and given neither the authority nor the information to have changed it.
The practical test is the one this estate keeps returning to and almost no organisation can answer: not whether an override exists, but when it was last used, by whom, and what happened to that person afterwards. A control nobody has exercised has not been shown to work.
The third: the checking nobody does#
The evidence on whether anyone verifies these systems is the most deflating part of the picture, and it comes from the jurisdiction that tried hardest.
New York City's Local Law 144 requires an independent bias audit of an automated employment decision tool before it is used, publication of the summary, and notice to candidates. A field audit in which 155 investigators posed as job seekers across 391 employers found about 5 per cent had published an audit. The New York State Comptroller then examined the enforcing department: reviewing the same 32 companies, the department had found one instance of non-compliance and the state auditors found at least seventeen.
The ethical weight of this is easy to miss. An organisation deploying a system to judge people is not choosing between a checked machine and an unchecked human. On the available evidence it is usually choosing between two unchecked things, one of which operates at scale.
Where the line has actually been drawn#
Legislators have not answered the ethical question in the abstract. They have answered a narrower one, by naming the decisions where the obligations bite hardest. Annex III of the EU AI Act classifies as high risk, among others, systems determining access to education, evaluating learning outcomes where those outcomes steer a learner's path, and a list of employment decisions.
That is a workable starting point for an organisation with no view of its own: the decisions that shape what a person is permitted to become. Classification is not evidence that any particular system is unsafe, and the Annex says nothing about how well the resulting duties are met. What it offers is a defensible answer to which decisions deserve the full apparatus, from a body that had to write one down.
Separating the four obligations before the argument starts#
- Separate the four obligations before arguing. Explanation, contestability, responsibility, verification. Most disputes about AI ethics are two people defending different ones.
- Write down who is accountable, by name, before deployment. If that sentence cannot be written, the system has produced a crumple zone rather than a decision process.
- Do not accept the vendor's accuracy figure. The sepsis case is the standing example: widely implemented on developer numbers that independent validation did not reproduce.
- Test the explanation on someone it would land on. The legal standard is whether the affected person can understand which of their data was used and how. That is checkable in an afternoon and almost never checked.
The question of principle is left open#
It does not answer the question as a matter of principle, nor try to. Whether it is ever right to let a machine decide something about a person is a moral question that evidence cannot close, and people who hold that some decisions require a human regardless of accuracy are making an argument this page does not defeat.
The four obligations also carry a cost that is rarely priced. Explanation, contest and named accountability slow decisions down and consume the labour the system was bought to save. An organisation that adopts the full apparatus honestly may find the business case disappears, and nobody in this literature has been willing to say so plainly.
Evidence review · SS-2026-212 · Graded against the published rubric
Hirji, R. (2026). Is it ethical to let AI judge people?. The SuperSkills evidence base, SS-2026-212. https://thesuperskills.com/research/is-it-ethical-to-let-ai-judge-people. Last reviewed 10 September 2026.
An evidence review by Rahim Hirji, not peer-reviewed research. For a material claim, cite the underlying study as well; every study here carries its own permanent link.
How citations and IDs work