- Should universities let AI mark students' work?
- Can AI mark exams as accurately as a lecturer?
- Do students have a right to know if AI marked their work?
- Should we let AI mark, grade or score work that our organisation certifies?
Not for the grade, and not without telling the student. The Australian universities that have written a policy have settled there, and the best measurement points the same way. On one computer science exam with 570 double-marked scripts, the best of 171 model configurations was closer to the human mark than the two human markers were to each other; a change to the instructions, telling the model to be strict, made 14 of 17 open models fail, and three stopped marking altogether. A marker that works under one prompt and collapses under another can draft feedback and check consistency; it cannot answer for a mark. The marking question is the assessment question from the other side: what the mark certifies, who answers for it, and whether the person who used to mark still learns what the class knows.
What the universities have decided, as reported#
Joe Hinchliffe at Guardian Australia surveyed the country’s universities on 28 September 2026 and found three positions rather than one. Western Sydney, RMIT and Adelaide allow academic staff to use AI to support parts of assessment and feedback; Western Sydney’s spokesperson told the paper that “responsibility for academic judgment must always remain with university teaching staff”. Deakin permits AI to improve efficiency and support assessment and administrative activities but not to assign grades. Newcastle gives qualified approval and lets students opt out. The University of New South Wales, Melbourne and Sydney do not use AI for marking. Monash said AI is not used in its marking process; Wollongong and James Cook are considering it. Armin Alimardani, a senior lecturer in law and technology at Western Sydney, told the paper “we’ve got to be really careful where we are stepping here”, and put the compliance problem in numbers: “Let’s say 80 do exactly as advised, there are going to be 20 who do not.” Alex Cameron, a creative writing student at Queensland University of Technology, said of the prospect that a machine would mark her work, “I would be absolutely bloody incensed.” No university in the survey said a machine would decide a grade.
Closer than two humans, until the instructions change#
The measurement that bears on the accuracy claim is Habibullah, Alshoibi, Alshiekh, Khan and Khan, a preprint of 24 September 2026. They ran 171 model configurations over a computer vision exam that 570 students sat and two human markers had graded independently. The best configuration’s mean absolute error was 1.64 marks out of 35. The two human markers differed from each other by 2.61 marks out of 35 on average. On the accuracy question alone, then, the machine was inside the range of ordinary human disagreement, and a department that double-marks would find nothing alarming in its numbers. The paper’s second finding is the one a policy has to be written around. Adding a preamble that told the model to be a strict grader made 14 of 17 open-weight models fail badly, and three of them stopped producing marks at all; the phrase “never give partial credit” on its own was enough to make models stop grading. The same failures reproduced on a second exam, in machine learning, sat by 1,038 students, with the direction of the error changing by exam. A light fine-tuning adapter trained on about 3,900 marked answers brought five small open models to parity with a human marker and reduced the sensitivity to wording. What that adds up to is a marker whose accuracy is a property of the prompt, the exam and the tuning, none of which the student can see and few of which the marking lecturer will have tested.
What a marker learns by marking#
The case against automating the grade does not rest on accuracy, and the estate has made it before for schools. On how AI will change teaching, marking is where all four conditions of the deskilling risk converge: the tool produces the judgement rather than the material; marking a set of scripts is one of the ways a teacher builds a picture of what a class knows; nobody measures a marker’s unassisted accuracy; and the feedback loop on a marking error is weak. The same four hold in a lecture theatre. A tutor who reads eighty essays learns which argument the cohort has misunderstood, which source everyone is copying and which student has begun to think; a tutor who reads eighty machine summaries of eighty essays learns what the machine noticed. European law arrived at the same place from a different direction: Annex III of Regulation (EU) 2024/1689 classes as high risk any AI system intended to evaluate learning outcomes, with documented human oversight under Article 14. The Guardian’s reporting suggests the profession’s instinct matches the regulator’s. Every university that permits the tool has confined it to support, and every one that has said where the line is has drawn it at the grade.
The students already use a machine the rules cannot describe#
Two more preprints from the same week set the other half of the picture. Poudyal tested fifteen standard scenarios of student AI use against the published policies of twenty Australian universities, 300 classifications in all: 40 per cent came out clearly prohibited, 32.3 per cent a potential breach, 9 per cent permitted with conditions and 18.7 per cent could not be determined from the documents at all. Not one scenario was clearly permitted anywhere, and two trained coders reading the same policies agreed only 57.3 per cent of the time. Written policies, the paper finds, regulate keeping AI-generated text far more clearly than they regulate help with the process. Eastwood, Narne, Hilby, Denny, Aggarwal and Kapoor randomised 132 introductory programming students across four AI teaching assistants and found that the most guarded one, Socratic and fully aware of the student’s problem, was rated least favourably and went with more use of outside chatbots, though the differences did not reach significance. Put the three papers together and the position is this: universities whose own rules cannot say what a student may do with a machine, whose approved tools students route around, are now deciding whether a machine may read the result. A marking policy written before the student-use policy is legible will be enforced against students by a marker nobody can question.
The decisions a university controls#
The pattern the estate applies to every handover, Rules Before Tools, fits marking without adjustment. Which tasks may the machine do: draft feedback for a human to edit, flag inconsistency across markers, check a rubric has been applied, and never assign the mark; Deakin’s wording does that job in one clause. Who can stop it: a named academic per module, with the right to withdraw the tool mid-semester without a committee. What people must remain able to do: mark unaided, and be seen to, which means a sampled double-marking exercise each year where the human mark is recorded before the machine’s is seen, since the Habibullah result shows how far a marker’s accuracy can move with a change of wording nobody noticed. How anyone would know it went wrong: a student’s right to be told, as Newcastle allows, and to have a human re-mark on request, because the error a prompt introduces is systematic and a student is the only person placed to notice their own case. None of that is expensive. All of it is a decision, and a university that has taken it can use the tool in the open rather than in the twenty rooms out of a hundred where Alimardani expects the rules to be ignored.
What this does not show#
The Guardian’s survey is one country’s universities answering a journalist, and a policy stated to a newspaper is not a policy observed in a department. Habibullah and colleagues measured two computer science exams with structured answers; nothing here shows how the machine marks an essay, a proof, a portfolio or a clinical reflection, and the paper is a preprint that has not been reviewed. Poudyal read documents, not practice, and the 18.7 per cent that could not be determined may be resolved in tutorials the study could not see. Eastwood and colleagues’ differences did not reach significance, so the bypass finding is a pattern, not a result. No study has yet measured what a lecturer who stops marking loses in their ability to teach, which is the claim this page rests on; it is inferred from the schools evidence and from the deskilling record elsewhere, and a university that wants to test it could run the double-marking exercise above and publish the gap.
Essay · SS-2026-368
Hirji, R. (2026). Should universities let AI mark students' work?. The SuperSkills evidence base, SS-2026-368. https://thesuperskills.com/research/should-universities-let-ai-mark-students-work. Last reviewed 28 September 2026.
An evidence review by Rahim Hirji, not peer-reviewed research. For a material claim, cite the underlying study as well; every study here carries its own permanent link.
How citations and IDs work