← Research
Research

Does AI detection work?

Not well enough to accuse anyone, and the errors fall hardest on the students least able to contest them.

Last reviewed: 26 August 2026

The Stanford finding, why the base rate makes it worse than it sounds, what serious institutions are doing instead, and the rules if you must use one anyway.

Question this page answersAll 616 questions this research covers

Not well enough to accuse anyone. The most important study on this found that detectors misclassified more than half of essays written by non-native English speakers as AI-generated, an average false positive rate of 61.22 per cent, while achieving near-perfect accuracy on essays by US eighth-graders. The same tool, the same threshold, and a wildly different error rate depending on who wrote the text.

That is not a tool with a calibration problem. That is a tool that systematically penalises one group of students, and any institution using it is running a disciplinary process with a known and severe demographic bias built into the first step.

What the study found#

Liang, Yuksekgonul, Mao, Wu and Zou at Stanford ran seven widely used GPT detectors against TOEFL essays written by non-native English speakers and against essays by US eighth-grade students. The non-native essays were flagged as AI-generated more than half the time. The eighth-grade essays were classified almost perfectly.

Their proposed explanation is the part that makes this structural rather than fixable by tuning. Detectors largely work on perplexity, a measure of how predictable the text is. Writing with less linguistic variability and a narrower vocabulary is more predictable, and that is what writing in a second language often looks like. The detector identifies limited linguistic range and then reports it as machine authorship.

The researchers also showed that simple prompting strategies both mitigate the bias and let genuinely AI-generated text bypass the detectors. So the tool is worst at the thing it is sold for, and its errors fall hardest on students least able to contest them.

Why the base rate makes it worse than it sounds#

Even a much better detector runs into arithmetic. Suppose a detector is 98 per cent accurate and 5 per cent of a 1,000-student cohort actually used AI. That is 50 true cases, of which it catches 49, and 950 honest students, of whom it wrongly flags 19. Almost a third of your accusations are against innocent people, with a tool most institutions would consider excellent.

Now apply the real-world error rates, and the arrangement stops being a detection system and becomes a mechanism for generating false accusations at scale.

What institutions are actually doing#

The serious ones have stopped pretending. The University of Sydney states in its own guidance that an unsecured no-AI condition is only a temporary measure because it cannot be enforced, and has moved to a two-lane model separating secured in-person assessment from AI-permitted work. Denmark went further, requiring from August 2026 that home-written examinations be defended orally.

Both responses accept the same premise: if you cannot verify the artefact, you have to examine the person. Detection is an attempt to avoid that conclusion, and it does not work.

What has moved since the Stanford tests#

Detector technology moves, and the Stanford work tested tools available at the time rather than every product now on the market. Some vendors claim substantially better performance, and commissioned evaluations should be read with the obvious interest in mind. It is entirely possible that detection accuracy improves.

What will not improve is the base-rate arithmetic, and what is unlikely to improve is the fairness problem, because the signal detectors rely on is correlated with linguistic range. A better detector still flags second-language writers more often unless something fundamental changes about how it works.

Detection is a way of not deciding#

Detection is an attempt to keep an assessment model that has already stopped working. The essay was never the point; it was a proxy for a capability, and the proxy has broken. Spending institutional effort on catching people is spending it on the wrong problem.

There is also a cost nobody counts. A detection regime teaches students that the institution's primary relationship with them is suspicion, and it does that most forcefully to international students, who are frequently the ones paying most for the privilege. That is a reputational and ethical exposure, not just a methodological one.

The better question is what you are trying to certify. If it is that a person can think in a domain, examine the person: orally, in secured conditions, or through work you watched them do. If it is that a piece of work meets a standard, judge the work and stop caring who typed it. Most assessment is currently doing neither and hoping a detector will resolve the ambiguity.

If you must use one#

The full argument on assessment is in assessing students when AI can do the assignment. On learning, how humans learn with AI and desirable difficulty. On the workplace version of the same problem, who owns verification.

Key sources

About this research#

Rahim Hirji is the author of SuperSkills (Kogan Page, 2026), keynote speaker on AI and human capability, and founder of The SuperSkills Intelligence Company. Findings are attributed to the studies that produced them and kept separate from the interpretation. This page describes evidence and offers a position; institutions should take their own advice on academic-integrity procedure.

How this research works  ·  Reviewed quarterly  ·  Found an error? Tell me and it is corrected on the page.

Cite this

Hirji, R. (2026). Does AI detection work? The SuperSkills Intelligence Company. Last reviewed 26 August 2026. thesuperskills.com/research/does-ai-detection-work

Questions answered on this page

Does AI detection work?

Not reliably enough to base an accusation on. Liang and colleagues at Stanford found seven widely used GPT detectors misclassified more than half of TOEFL essays written by non-native English speakers as AI-generated, an average false positive rate of 61.22 per cent, while classifying US eighth-grade essays almost perfectly. The same tool and threshold produced wildly different error rates depending on who wrote the text.

Why are AI detectors biased against non-native English speakers?

Because detectors largely work on perplexity, a measure of how predictable text is, and writing with less linguistic variability and a narrower vocabulary is more predictable. That is what writing in a second language often looks like. The detector is not identifying machine authorship; it is identifying limited linguistic range and reporting it as machine authorship. The same researchers showed simple prompting both mitigates the bias and lets genuinely AI-generated text bypass detection.

Why does the base rate matter?

Because even an excellent detector generates mostly false accusations when cheating is uncommon. A 98 per cent accurate detector applied to 1,000 students where 5 per cent used AI catches 49 of 50 real cases and wrongly flags 19 of 950 honest students, so almost a third of accusations are against innocent people. With real-world error rates it stops being a detection system and becomes a mechanism for generating false accusations at scale.

What should institutions do instead?

Accept that if you cannot verify the artefact, you have to examine the person. The University of Sydney states in its own guidance that unsecured no-AI conditions cannot be enforced and has moved to a two-lane model separating secured in-person assessment from AI-permitted work. Denmark requires home-written examinations to be defended orally from August 2026. If a detector is used at all, a flag should be a prompt for a conversation and never evidence, the false positive rate should be published to students in advance, and outcomes should be tracked by first language.

In this hub

Thinking, learning and capability

What sustained AI use does to thinking, and how capability is built and kept.

The work

Where the writing comes from.

These essays draw on research across more than 200 organisations in 30 countries. See the wider body of work, or bring it into your organisation.

All research →

Schools and universities are where this question is least theoretical. There is the schools and education version, and the full range of topics and audiences.

Box of Amazing

Rahim’s free weekly letter on AI and human capability

If this was useful, the weekly letter is where the thinking happens first. Most of what ends up on this site starts there. Weekly essays on AI, capability and the future of work. Read by 25,000 people, every week since 2017. Free, and one click to stop.

Opens Substack to confirm. No pitch in it, unsubscribe in one click, and nobody follows up because you read something.