You cannot. Not reliably, and not by detection. Every institution that has tried to police its way out of this has arrived at the same place, and the University of Sydney has been unusually honest about it, conceding in its own guidance that an unsecured no-AI condition is "only a temporary measure, noting that it is actually not possible to enforce this." Once you accept that, the question changes from how to catch students to what you are actually assessing. If you cannot verify the artefact, you have to examine the person. That is expensive, it does not scale comfortably, and it is what medicine and aviation concluded decades ago. The two serious responses in the world right now are Sydney's two-lane model and Denmark's national reinstatement of the oral defence, and both start from the same admission.
What the evidence says about the stakes
This is not a policing problem dressed up as a pedagogy problem. It is a pedagogy problem, and there is now a field experiment that shows why. Bastani and colleagues, publishing in PNAS in 2025, gave nearly a thousand high-school mathematics students access to a GPT-4 tutor in three arms: a plain chat interface, a version designed with guardrails to give hints rather than answers, and a control group with textbook and notes.
While the tool was available, both AI groups did far better: grades up 48 percent for the plain interface and 127 percent for the guardrailed tutor. Then the researchers took the tool away and tested the students alone. The plain-interface group scored 17 percent lower than students who had never had access at all. The guardrailed tutor largely eliminated that harm.
Read carefully, that is not a finding about AI. It is a finding about interface design. The same underlying model produced the best learning outcome in the experiment and the worst one, and the only variable was whether the tool made the student do the work. Any assessment policy written without that distinction is solving the wrong problem.
The learning science explains the mechanism. Bjork and Bjork's work on desirable difficulties shows that conditions making study feel harder, such as spacing, interleaving and retrieval practice, improve long-term retention, while conditions making it feel fluent improve immediate performance and worsen retention. Crucially, learners systematically mistake fluency for learning. AI is a fluency machine, so students experiencing an easy, productive session have every reason to believe they are learning and no reliable way to notice they are not.
The two responses worth studying
Sydney's two-lane model, announced in November 2024 and fully in force from the second semester of 2025, splits assessment explicitly. Lane one is secured and in person. Lane two permits AI. What makes it instructive is not the structure but the candour: the university states that prohibition in unsecured conditions cannot be enforced, and some disciplines are considering grading only lane one. It stops pretending, which is the necessary first move.
Denmark went further and faster. In August 2026 the Ministry of Children and Education announced an immediate package requiring that every examination written at home must be defended orally, affecting roughly nine thousand students a year in the upper-secondary system. The oldest assessment technology there is, restored at national scale. The format was mandated before the method had been fully designed, which tells you how urgent it felt to the people responsible.
Elsewhere the moves are smaller and in both directions. Victoria University of Wellington returned law examinations to handwriting in 2025, while Auckland moved further towards digital. Cambridge's Human, Social and Political Sciences faculty scrapped online examinations in 2024, wanting typed in-person papers with handwriting as the budget fallback. South Korea revoked the legal status of AI digital textbooks in August 2025 after roughly 533 billion won had been spent, with adoption stalling near thirty percent.
Where this is genuinely uncertain
Nobody knows what a good oral examination at scale looks like. Denmark has mandated one without publishing the method, and viva assessment carries well-known problems: it is expensive in staff time, it advantages the confident and the fluent over the thoughtful, and it is harder to moderate for consistency and for bias than a marked script. Reinstating it solves the authorship problem and imports a fairness problem.
The Bastani result is also school mathematics over a bounded period, with clear right answers. It establishes that the design variable exists and matters. It does not tell a history department what a guardrailed essay tool should look like, and the honest position is that nobody has built one yet.
And detection remains unreliable in both directions. False accusations fall hardest on students writing in a second language, whose prose is more likely to be flagged. An institution that leans on detection is not only failing to catch the problem, it is manufacturing a different injustice.
The SuperSkills view
Separate two measurements permanently, and stop conflating them. Performance with the tool is what the work requires and what employers will want. Capability without the tool is what the person has become. Both matter. They are not the same number, and an education system that only records the first has stopped measuring the thing it exists to produce.
That reframing dissolves most of the policy argument. You do not need to ban AI, which is unenforceable, or permit it everywhere, which stops assessing the person. You need some assessment of each kind, declared openly, with students told which is which and why. Sydney's two lanes are exactly this, and its honesty about enforceability is what makes the design work.
The deeper point is about what an assignment was ever for. It was rarely the artefact. A history essay is not valuable because the world needed another history essay; it is valuable because writing it forces the student to assemble an argument from evidence, which is a capability they carry afterwards. That capability is what I call the missed reps when it goes missing, and it is the same mechanism a GP described this year on losing the note-writing through which clinicians build pattern recognition. Assessment reform is not about integrity. It is about protecting the thing the assignment was a proxy for.
What to do
Declare the lane. Every assessment should state whether it measures performance with AI or capability without it. Ambiguity is what produces both cheating and unfair accusations.
Design the tool, do not just permit or ban it. The Bastani result says the interface decides the outcome. A tool that withholds until the student commits, asks rather than answers, and makes difficulty visible is a different educational object from a chat box, even running the same model.
Assess the process, not only the product. Ask for the prompt, what came back, what was changed and why. That is both harder to fake and more useful to learn from than a finished essay.
Bring back some live examination, and design it properly. Oral defence, live problems, structured questioning. Take the fairness problems seriously rather than discovering them later.
Stop relying on detection. It does not work well enough to carry a disciplinary process, and its errors are not randomly distributed.
Related SuperSkills research
The evidence underneath this is in how humans learn with AI. On what is lost when the practice goes, see the missed reps and capability debt. On the same problem arriving in the workplace, will AI replace entry-level jobs and the missing rungs.
Key research and primary sources
- Bastani, H. et al. (2025). Generative AI Without Guardrails Can Harm Learning. Proceedings of the National Academy of Sciences, 122(26).
- Bjork, E. L. and Bjork, R. A. Making Things Hard on Yourself, But in a Good Way. In Psychology and the Real World.
- University of Sydney (2025). The two-lane approach to assessment in the age of AI.
- Danish Ministry of Children and Education (2026). Ny strakspakke mod AI-snyd på gymnasierne.
- UNESCO (2024). AI competency framework for teachers.
- Kosmyna, N. et al. (2025). Your Brain on ChatGPT. MIT Media Lab preprint. Widely quoted, 54 participants, treat with caution.
About this research
Rahim Hirji is the author of SuperSkills: The Seven Human Skills for the Age of AI (Kogan Page, 2026) and the founder of The SuperSkills Intelligence Company. Findings are attributed to the studies and institutions that produced them and kept separate from the interpretation, which is the author's. Desirable difficulties is an established concept from learning science and is not his. This is a living reference on a 90-day review cycle, given how fast institutional policy is moving.
Cite this
Hirji, R. (2026). Assessing students when AI can do the assignment. The SuperSkills Intelligence Company. Last reviewed 26 August 2026. thesuperskills.com/research/how-to-assess-students-when-ai-can-do-the-assignment