- Is AI decision-making biased?
- Can AI be unbiased, or does it reproduce human bias?
- Do bias audits actually happen?
The question is usually put as a hope. Machines have no prejudices, so a machine decision ought to be cleaner than a human one. The measurements say otherwise, and they have said so for years in enough detail to name the sizes of the effects. The more useful finding sits one layer down, in a mathematical result rather than an empirical one: there is no single thing called being unbiased. Several reasonable definitions of fairness cannot all hold at once, so every system that judges people has already chosen between them, usually without anybody noticing that a choice was made.
The answer, in one line
It reproduces human bias at scale, and this has been measured with large effects. It also cannot be unbiased in the way the question implies: Kleinberg, Mullainathan and Raghavan proved that, outside highly constrained special cases, several reasonable fairness criteria cannot all be satisfied at once.
The short answer#
It reproduces human bias, at scale, and this has been measured repeatedly with large effects. It cannot be made unbiased in the way the question implies, because the criteria people mean by that word are provably incompatible outside narrow special cases. What it can be is audited, which a person cannot, and that is the real difference between machine judgement and human judgement.
So the question that decides outcomes is not whether the system is biased. It is whether anyone is checking, and what happens when a check finds something. On the evidence from the one jurisdiction that made checking compulsory, the answers are mostly nobody and mostly nothing.
The bias is measured, not alleged#
Three studies give the shape of it, and none rests on anecdote.
A resume audit run through a document retrieval pipeline, published at the AAAI/ACM conference on AI, ethics and society, used over 500 real resumes and 500 job descriptions across nine occupations, with 120 first names associated with male, female, Black and white candidates. The embedding models favoured White-associated names in 85.1 per cent of comparisons and female-associated names in 11.1 per cent. Black male candidates were disadvantaged in up to 100 per cent of cases. Part of the effect traced to how often a name appears in training data, which has nothing to do with any candidate.
The authors are clear about the limit, and it matters: these are open models in a simulated pipeline, not the proprietary products vendors sell. Those cannot be tested independently, a point this page returns to.
A study of linguistic bias found responses to non-standard English varieties carried 19 per cent more stereotyping, 25 per cent more demeaning content and 15 per cent more condescension, all significant. The sharper number is retention: a Standard American English input keeps 77.9 per cent of its dialect features in the model's reply, and five minoritised varieties keep two to three per cent. The default is erasure rather than caricature, which is the opposite of the claim usually attached to this work.
AI-detection tools misclassified more than half of essays by non-native English writers as machine-written, a false positive rate averaging 61.22 per cent, while handling US eighth-grade essays almost perfectly. The mechanism is that detectors read predictability, and second-language writing is more predictable. A tool built to catch cheating turned out to be a tool for catching foreigners.
The comparison the question leaves out#
Every one of those findings invites the response that human decision-makers are biased too, and the response is correct. Human hiring, human grading and human lending carry documented disparities, and the systems being replaced were never neutral.
The reason that observation settles less than it appears to is asymmetry of visibility. A human screener's bias is distributed across thousands of separate people, unrecorded, and inaccessible after the fact. A model's bias is one artefact, applied identically to every applicant, and available for measurement by anyone with the access. The machine concentrates the harm and simultaneously makes it findable. Which of those two properties dominates depends entirely on whether anybody looks.
Why unbiased is not a coherent target#
Underneath the empirical work is a formal result that most public argument has never absorbed. Kleinberg, Mullainathan and Raghavan set out three fairness conditions that recur in these debates and proved that, except in highly constrained special cases, no method satisfies all three at once. Satisfying them even approximately requires the data to sit close to one of those special cases.
The consequence for anybody buying or governing such a system is direct. A vendor claiming their tool is fair has satisfied some criterion, and by the theorem it is failing another one that a reasonable person would also call fairness. The question to ask is which criterion was chosen and who chose it, because that decision is about values rather than engineering and it is currently being made by product teams.
The theorem does not say bias cannot be reduced, and it does not make auditing pointless. It says the word unbiased conceals a choice.
What happened when a city made checking compulsory#
New York City passed Local Law 144, the first algorithmic bias audit law in the world, requiring an annual independent bias audit of any automated employment decision tool before it is used, publication of the summary, and notice to candidates. The rules define the trigger with unusual precision, including the use of a simplified output to overrule conclusions reached by other means.
Then two independent measurements arrived. A field audit in which 155 student investigators acted as job seekers across 391 employers found that 18 employers, about 5 per cent, had posted a bias audit, and 13 had posted a transparency notice. The authors name the resulting condition null compliance: non-compliance cannot even be established, because the statute's design makes it impossible to tell from outside whether an employer uses a covered tool at all.
Then the New York State Comptroller audited the enforcing department. Reviewing 32 companies, the department had identified one issue of non-compliance. The state auditors reviewed the same 32 companies and identified at least seventeen instances of potential non-compliance. Two complaints had been received in two years.
Read those together and the lesson is not that the law failed. It is that the measurable question moved. Whether a hiring tool is biased was never going to be settled by legislation; whether anyone would find out was, and did not.
Questions to put to a vendor, and the clause to negotiate#
- Stop asking vendors whether their system is fair. Ask which fairness criterion they optimised, what it trades against, and who inside their company decided. A vendor who cannot answer has not thought about it; one who says all of them has not read the literature.
- Treat auditability as the procurement requirement. The bias is a given. The ability to measure it on your own population, and a contractual right to do so, is the thing that is actually negotiable at purchase and impossible afterwards.
- Check the tool on the people you actually decide about. Every study here measures a population that is not yours. Base rates differ, and by the theorem the fairness properties move with them.
- Assume the compliance paperwork is not the check. One jurisdiction mandated audits and got five per cent publication, and its own enforcer found one problem where independent reviewers found seventeen in the same sample.
Nobody outside a vendor has tested a deployed system#
None of the studies here tested a deployed commercial product, because deployed commercial products cannot be tested by outsiders. That gap runs through the whole field: the systems that make real decisions about real people are the ones nobody outside the vendor has measured, so the published evidence describes the layer underneath them rather than them.
Nor does any of this establish that a human alternative would be better. The comparison almost nobody runs is the one that matters, a like-for-like measurement of the same decision made both ways on the same population, and the reason it is rare is that measuring the human side well is harder than measuring the machine.
Evidence review · SS-2026-210 · Graded against the published rubric
Hirji, R. (2026). Can AI be unbiased, or does it reproduce human bias?. The SuperSkills evidence base, SS-2026-210. https://thesuperskills.com/research/can-ai-be-unbiased. Last reviewed 10 September 2026.
An evidence review by Rahim Hirji, not peer-reviewed research. For a material claim, cite the underlying study as well; every study here carries its own permanent link.
How citations and IDs work