← Research
Research

Assessing students when AI can do the assignment

If you cannot verify the artefact, you have to examine the person. That is expensive. It is what medicine and aviation concluded decades ago.

Last reviewed: 26 August 2026

How do you assess students when AI can do the assignment? Not by detection. This page sets out the field evidence, the two serious institutional responses, and the distinction Rahim Hirji argues every assessment should declare.

Questions this page answersAll 616 questions this research covers

You cannot. Not reliably, and not by detection. Every institution that has tried to police its way out of this has arrived at the same place, and the University of Sydney has been unusually honest about it, conceding in its own guidance that an unsecured no-AI condition is "only a temporary measure, noting that it is actually not possible to enforce this." Once you accept that, the question changes from how to catch students to what you are actually assessing. If you cannot verify the artefact, you have to examine the person. That is expensive and it does not scale comfortably. Medicine and aviation reached the same conclusion decades ago. The two serious responses in the world right now are Sydney's two-lane model and Denmark's national reinstatement of the oral defence, and both start from the same admission.

What the evidence says about the stakes#

This is a pedagogy problem that keeps getting handled as a policing one, and there is now a field experiment that shows why. Bastani and colleagues, publishing in PNAS in 2025, gave nearly a thousand high-school mathematics students access to a GPT-4 tutor in three arms: a plain chat interface, a version designed with guardrails to give hints rather than answers, and a control group with textbook and notes.

While the tool was available, both AI groups did far better: grades up 48 percent for the plain interface and 127 percent for the guardrailed tutor. Then the researchers took the tool away and tested the students alone. The plain-interface group scored 17 percent lower than students who had never had access at all. The guardrailed tutor largely eliminated that harm.

Read carefully, that is a finding about interface design rather than about AI. The same underlying model produced the best learning outcome in the experiment and the worst one, and the only variable was whether the tool made the student do the work. Any assessment policy written without that distinction is solving the wrong problem.

The learning science explains the mechanism. Bjork and Bjork's work on desirable difficulties shows that conditions making study feel harder, such as spacing, interleaving and retrieval practice, improve long-term retention, while conditions making it feel fluent improve immediate performance and worsen retention. Crucially, learners systematically mistake fluency for learning. AI is a fluency machine, so students experiencing an easy, productive session have every reason to believe they are learning and no reliable way to notice they are not.

The two responses worth studying#

Sydney's two-lane model, announced in November 2024 and fully in force from the second semester of 2025, splits assessment explicitly. Lane one is secured and in person. Lane two permits AI. What makes it instructive is not the structure but the candour: the university states that prohibition in unsecured conditions cannot be enforced, and some disciplines are considering grading only lane one. It stops pretending, which is the necessary first move.

Denmark went further and faster. In August 2026 the Ministry of Children and Education announced an immediate package requiring that every examination written at home must be defended orally, affecting roughly nine thousand students a year in the upper-secondary system. The oldest assessment technology there is, restored at national scale. The format was mandated before the method had been fully designed, which tells you how urgent it felt to the people responsible.

Elsewhere the moves are smaller and in both directions. Victoria University of Wellington returned law examinations to handwriting in 2025, while Auckland moved further towards digital. Cambridge's Human, Social and Political Sciences faculty scrapped online examinations in 2024, wanting typed in-person papers with handwriting as the budget fallback. South Korea revoked the legal status of AI digital textbooks in August 2025 after roughly 533 billion won had been spent, with adoption stalling near thirty percent.

Where this is genuinely uncertain#

Nobody knows what a good oral examination at scale looks like. Denmark has mandated one without publishing the method, and viva assessment carries well-known problems: it is expensive in staff time, it advantages the confident and the fluent over the thoughtful. It is harder to moderate for consistency and for bias than a marked script. Reinstating it solves the authorship problem and imports a fairness problem.

The Bastani result is also school mathematics over a bounded period, with clear right answers. It establishes that the design variable exists and matters. It does not tell a history department what a guardrailed essay tool should look like, and nobody has built one yet.

And detection remains unreliable in both directions. False accusations fall hardest on students writing in a second language, whose prose is more likely to be flagged. An institution that leans on detection fails to catch the problem and manufactures a different injustice at the same time.

Two measurements, permanently separated#

Separate two measurements permanently, and stop conflating them. Performance with the tool is what the work requires and what employers will want. Capability without the tool is what the person has become. Both matter. They are not the same number, and an education system that only records the first has stopped measuring the thing it exists to produce.

That reframing dissolves most of the policy argument. You do not need to ban AI, which is unenforceable, or permit it everywhere, which stops assessing the person. You need some assessment of each kind, declared openly, with students told which is which and why. Sydney's two lanes are exactly this, and its honesty about enforceability is what makes the design work.

The deeper point is about what an assignment was ever for. It was rarely the artefact. A history essay is not valuable because the world needed another history essay; it is valuable because writing it forces the student to assemble an argument from evidence, which is a capability they carry afterwards. That capability is what I call the missed reps when it goes missing. It is the same mechanism a GP described this year on losing the note-writing through which clinicians build pattern recognition. Assessment reform protects the thing the assignment was a proxy for. Integrity is a side effect.

How to run it#

Declare the lane. Every assessment should state whether it measures performance with AI or capability without it. Ambiguity is what produces both cheating and unfair accusations.

Design the tool, do not just permit or ban it. The Bastani result says the interface decides the outcome. A tool that withholds until the student commits, asks rather than answers, and makes difficulty visible is a different educational object from a chat box, even running the same model.

Assess the process, not only the product. Ask for the prompt, what came back, what was changed and why. That is both harder to fake and more useful to learn from than a finished essay.

Bring back some live examination, and design it properly. Oral defence, live problems, structured questioning. Take the fairness problems seriously rather than discovering them later.

Stop relying on detection. It does not work well enough to carry a disciplinary process, and its errors are not randomly distributed.

The evidence underneath this is in how humans learn with AI. On what is lost when the practice goes, see the missed reps and capability debt. On the same problem arriving in the workplace, will AI replace entry-level jobs and the missing rungs. On the definition, desirable difficulty. See does AI detection work. See should children use AI.

Key research and primary sources

About this research#

Rahim Hirji is the author of SuperSkills (Kogan Page, 2026), keynote speaker on AI and human capability, and founder of The SuperSkills Intelligence Company. Findings are attributed to the studies and institutions that produced them and kept separate from the interpretation, which is the author's. Desirable difficulties is an established concept from learning science and is not his. This is a living reference on a 90-day review cycle, given how fast institutional policy is moving.

How this research works  ·  Reviewed quarterly  ·  Found an error? Tell me and it is corrected on the page.

Cite this

Hirji, R. (2026). Assessing students when AI can do the assignment. The SuperSkills Intelligence Company. Last reviewed 26 August 2026. thesuperskills.com/research/how-to-assess-students-when-ai-can-do-the-assignment

Questions answered on this page

How do you assess students when AI can do the assignment?

Not by detection, which is unreliable and whose errors fall hardest on students writing in a second language. The University of Sydney states in its own guidance that an unsecured no-AI condition is only a temporary measure because it is not possible to enforce. Once you accept that, the answer is to examine the person rather than verify the artefact: Sydney's two-lane model separates secured in-person assessment from AI-permitted work, and Denmark from August 2026 requires every home-written examination to be defended orally.

Does using AI harm student learning?

Unrestricted access can. In a 2025 PNAS field experiment, Bastani and colleagues gave nearly a thousand high-school mathematics students a GPT-4 tutor. Grades rose 48 percent with a plain chat interface and 127 percent with a guardrailed tutor while the tool was available. When it was removed, the plain-interface group scored 17 percent lower than students who never had access. The guardrailed version largely removed that harm. The interface design, not the presence of AI, decided the outcome.

What is the two-lane approach to assessment?

The University of Sydney's model, fully in force from the second semester of 2025. Lane one is secured, in-person assessment. Lane two permits AI use. What makes it useful is the university's candour: its guidance concedes that unsecured no-AI conditions cannot be enforced, and some disciplines are considering grading only lane one.

Should universities go back to handwritten and oral exams?

Some already have. Victoria University of Wellington returned law examinations to handwriting in 2025, Cambridge's HSPS faculty scrapped online examinations in 2024, and Denmark mandated oral defence nationally in August 2026. But reinstating viva assessment solves the authorship problem and imports a fairness problem: it is expensive in staff time, advantages the confident and fluent over the thoughtful, and is harder to moderate for consistency and bias than a marked script. Nobody has yet published what a good oral examination at national scale looks like.

In this hub

Thinking, learning and capability

What sustained AI use does to thinking, and how capability is built and kept.

The work

Where the writing comes from.

These essays draw on research across more than 200 organisations in 30 countries. See the wider body of work, or bring it into your organisation.

All research →

Schools and universities are where this question is least theoretical. There is the schools and education version, and the full range of topics and audiences.

Box of Amazing

Rahim’s free weekly letter on AI and human capability

If this was useful, the weekly letter is where the thinking happens first. Most of what ends up on this site starts there. Weekly essays on AI, capability and the future of work. Read by 25,000 people, every week since 2017. Free, and one click to stop.

Opens Substack to confirm. No pitch in it, unsubscribe in one click, and nobody follows up because you read something.