Not well enough to accuse anyone. The most important study on this found that detectors misclassified more than half of essays written by non-native English speakers as AI-generated, an average false positive rate of 61.22 per cent, while achieving near-perfect accuracy on essays by US eighth-graders. The same tool, the same threshold, and a wildly different error rate depending on who wrote the text.
That is not a tool with a calibration problem. That is a tool that systematically penalises one group of students, and any institution using it is running a disciplinary process with a known and severe demographic bias built into the first step.
What the study found
Liang, Yuksekgonul, Mao, Wu and Zou at Stanford ran seven widely used GPT detectors against TOEFL essays written by non-native English speakers and against essays by US eighth-grade students. The non-native essays were flagged as AI-generated more than half the time. The eighth-grade essays were classified almost perfectly.
Their proposed explanation is the part that makes this structural rather than fixable by tuning. Detectors largely work on perplexity, a measure of how predictable the text is. Writing with less linguistic variability and a narrower vocabulary is more predictable, and that is exactly what writing in a second language often looks like. The detector is not identifying machine authorship. It is identifying limited linguistic range, and then reporting it as machine authorship.
The researchers also showed that simple prompting strategies both mitigate the bias and let genuinely AI-generated text bypass the detectors. So the tool is worst at the thing it is sold for, and its errors fall hardest on students least able to contest them.
Why the base rate makes it worse than it sounds
Even a much better detector runs into arithmetic. Suppose a detector is 98 per cent accurate and 5 per cent of a 1,000-student cohort actually used AI. That is 50 true cases, of which it catches 49, and 950 honest students, of whom it wrongly flags 19. Almost a third of your accusations are against innocent people, with a tool most institutions would consider excellent.
Now apply the real-world error rates, and the arrangement stops being a detection system and becomes a mechanism for generating false accusations at scale.
What institutions are actually doing
The serious ones have stopped pretending. The University of Sydney states in its own guidance that an unsecured no-AI condition is only a temporary measure because it cannot be enforced, and has moved to a two-lane model separating secured in-person assessment from AI-permitted work. Denmark went further, requiring from August 2026 that home-written examinations be defended orally.
Both responses accept the same premise: if you cannot verify the artefact, you have to examine the person. Detection is an attempt to avoid that conclusion, and it does not work.
Where this is uncertain
Detector technology moves, and the Stanford work tested tools available at the time rather than every product now on the market. Some vendors claim substantially better performance, and commissioned evaluations should be read with the obvious interest in mind. It is entirely possible that detection accuracy improves.
What will not improve is the base-rate arithmetic, and what is unlikely to improve is the fairness problem, because the signal detectors rely on is genuinely correlated with linguistic range. A better detector still flags second-language writers more often unless something fundamental changes about how it works.
The SuperSkills view
Detection is an attempt to keep an assessment model that has already stopped working. The essay was never the point; it was a proxy for a capability, and the proxy has broken. Spending institutional effort on catching people is spending it on the wrong problem.
There is also a cost nobody counts. A detection regime teaches students that the institution's primary relationship with them is suspicion, and it does that most forcefully to international students, who are frequently the ones paying most for the privilege. That is a reputational and ethical exposure, not just a methodological one.
The better question is what you are trying to certify. If it is that a person can think in a domain, examine the person: orally, in secured conditions, or through work you watched them do. If it is that a piece of work meets a standard, judge the work and stop caring who typed it. Most assessment is currently doing neither and hoping a detector will resolve the ambiguity.
If you must use one
- Never as evidence. A flag is a prompt to have a conversation, never a finding. Treating a probability score as proof is indefensible given the error rates.
- Publish the false positive rate you are working with, to students, before the assessment. If you are unwilling to publish it, you should not be using the tool.
- Track outcomes by first language. If your flags concentrate among second-language writers, the Stanford finding is reproducing itself in your institution and you now know.
- Give the student the work back and ask them to explain it. This is more accurate than any detector and it is what you were going to have to do anyway.
Related SuperSkills research
The full argument on assessment is in assessing students when AI can do the assignment. On learning, how humans learn with AI and desirable difficulty. On the workplace version of the same problem, who owns verification.
Key sources
- Liang, W., Yuksekgonul, M., Mao, Y., Wu, E. and Zou, J. (2023). GPT detectors are biased against non-native English writers. Patterns, 4(7).
- Bastani, H. et al. (2025). Generative AI can harm learning. PNAS.
- Bjork, R. A. and Bjork, E. L. Desirable difficulties in theory and practice.
About this research
Rahim Hirji is the author of SuperSkills (Kogan Page, 2026) and founder of The SuperSkills Intelligence Company. Findings are attributed to the studies that produced them and kept separate from the interpretation. This page describes evidence and offers a position; institutions should take their own advice on academic-integrity procedure. Reviewed quarterly.
Cite this
Hirji, R. (2026). Does AI detection work? The SuperSkills Intelligence Company. Last reviewed 26 August 2026. thesuperskills.com/research/does-ai-detection-work