Mostly by taking the preparation and leaving the room. The strongest evidence in this profession measures teacher workload rather than pupil learning: a school-randomised trial in England found lesson planning time falling by 31 per cent with no detectable change in the quality of what was produced. The evidence on pupils splits, and it splits on whether the tool was engineered to teach or simply made available. A tutor built by a physics lecturer to follow known pedagogy beat his own active-learning class. A chatbot handed to teenagers raised their marks while they had it and left them behind the control group once it was taken away.
The trial asked the question teachers were actually asking#
The Education Endowment Foundation funded a Teacher Choices trial, evaluated independently by the National Foundation for Educational Research, on a deliberately narrow question: does using ChatGPT for Key Stage 3 science lesson preparation reduce the time it takes. 259 teachers across 68 schools in England took part over ten weeks in the summer term of 2024. One arm used ChatGPT with a written guide; the other was asked to use no generative AI at all.
Weekly planning time for the relevant classes came out at 56.2 minutes in the ChatGPT arm against 81.5 minutes in the comparison arm, a saving of 25.3 minutes and a reduction of 31 per cent. The EEF gives the result a high security rating. An expert panel reviewed the resulting lesson resources without being told which had been produced with the tool and found no noticeable difference in quality.
Three details in the report matter more than the headline and travel less well.
- The teachers were given five weeks before anything was measured. Weeks one to five were a familiarisation period; planning time was recorded in weeks six to ten. A school reading the 31 per cent as an immediate return would be reading it wrong.
- Use declined over the course of the trial, as did consultation of the guide. The saving was produced by teachers who were using the tool for one or two parts of a lesson, most often generating questions or quizzes and finding activity ideas, rather than across the whole plan.
- The perception moved further than the clock. The proportion of ChatGPT-arm teachers who felt they spent too much time on preparation fell from 49 per cent to 26 per cent. The comparison arm showed no similar fall. Twenty-five minutes a week does not obviously explain a twenty-three point swing, which suggests part of what is being relieved is the dread of the blank page.
What the trial does not establish is anything about teaching. It measured minutes spent preparing one subject at one key stage. Whether the lessons worked better for pupils was outside its scope, and the resource quality check was a blinded panel reading materials rather than anyone watching a class.
Thirty-five per cent plan with it, five per cent mark with it#
The Department for Education's Technology in Schools survey for 2024 to 2025, conducted by IFF Research across 1,634 schools and covering 795 leaders, 1,211 teachers and 489 IT leads, put numbers on where the profession has actually put the tool. 44 per cent of teachers reported using generative AI for school activities. Within that: lesson planning 35 per cent, delivering live lessons 7 per cent, marking 5 per cent.
The distribution is doing something interesting. Teachers have concentrated the tool on the part of the job that happens before anyone is in front of them, and have almost entirely kept it out of the two parts where a judgement about a particular child is being made. Nobody instructed them to. Only around a fifth of schools had a policy on safe and appropriate use at all, and the survey's own interviews found informal guidance more common than formal policy.
The generational split is the part worth watching. Teachers under 35 used it for planning at 43 per cent against 32 per cent for older colleagues, and for written feedback at 21 per cent against 12 per cent. Teachers with under three years in the classroom used it for written feedback at 27 per cent against 14 per cent for everyone else. The people leaning hardest on it for feedback are the people who have written the fewest reports by hand. That is the missing rungs pattern arriving in a staffroom.
School leaders were also more than twice as likely to be planning investment in AI tools for teachers than in tools for pupils, 58 per cent against 20 per cent. On the pupil side the picture is defensive: 73 per cent of secondary teachers thought pupils had used it for homework, 77 per cent of secondary leaders whose pupils could access it reported issues, and the most common issue reported was plagiarism at 67 per cent.
The best result for an AI tutor came from a lecturer who built one#
Kestin and colleagues ran a randomised crossover experiment in Harvard's largest introductory physics course. Of 233 enrolled students, 194 were eligible, and each experienced both conditions: one topic taught in an active-learning class, another taught at home by a purpose-built tutor, with the conditions reversed the following week.
Median post-test score was 4.5 in the AI condition against 3.5 in the classroom condition, measured against a combined pre-test median of 2.75. Median learning gain in the AI condition was over double. A rank-sum test gave z of -5.6 at p below 10 to the minus 8. The regression estimate of effect size, 0.63, is described by the authors as an underestimate because of a ceiling effect; a quantile regression puts the range at 0.73 to 1.3 standard deviations. Median time on task was 49 minutes against the 60 minutes assumed for the class, and time on task did not correlate with score. Students reported higher engagement, 4.1 against 3.6, and higher motivation, 3.4 against 3.1. Enjoyment and growth mindset showed no significant difference.
Read the methods and the result becomes narrower and more useful. The tutor's accuracy relied on pre-written answers. The prompts were question-specific and written by instructors who knew the content deeply. The lessons carried high-quality instructional video and a structured scaffold that governed how the tutor was allowed to respond. The authors state plainly that they do not presume structured AI tutoring will outperform in-class active learning in all contexts, naming complex synthesis and higher-order critical thinking as the likely exceptions, and they position the finding at the understanding, applying and analysing levels of Bloom's taxonomy.
So the study is evidence that a well-designed tutor beats a well-designed class at the delivery stage of new material. It is not evidence about a pupil opening a chatbot. The gap between those two things is the entire argument about AI in education, and almost every citation of this paper closes it silently.
The experiment that took the tutor away again#
Bastani and colleagues ran a field experiment with Turkish high school students. Marks rose 48 per cent with unrestricted access to a GPT-4 assistant and 127 per cent with a guardrailed tutor designed to withhold answers. Then access was removed for the exam. The unrestricted group scored 17 per cent below students who had never had the tool at all. The guardrailed arm was largely spared that penalty.
These two studies are usually presented as a contradiction and they are not one. Kestin measured performance while a carefully engineered tutor was present. Bastani measured what remained after an ordinary one was withdrawn. Both found the design of the tool mattering more than its presence, and both found the guardrailed configuration, the one that makes the student do the work, producing the durable result. The distinction this research has been drawing at desirable difficulty and productive struggle is the same distinction, arrived at from the other direction.
The uncomfortable corollary for a school: both of Bastani's arms would have looked identical on any usage dashboard. Adoption metrics cannot see the difference between the configuration that taught and the configuration that did the work instead. This research calls that gap usage theatre, and education is where it is cheapest to fall into.
Marking is the exposed part, and teachers have found that out first#
Applying the four conditions from the deskilling risk test, teaching scores medium and asymmetrically. Preparation is a complement: the material is assembled and the teacher still decides what to do with it. The diagnostic judgement of what a particular child has misunderstood is not something current tools substitute for well, because the evidence for that judgement is in the room rather than in the text.
Marking is where the conditions converge. The tool can produce the judgement rather than the material, which is condition one. Marking a set of books is one of the ways a new teacher builds a picture of what a class actually knows, which is condition two. Nobody measures a teacher's unassisted marking accuracy, which is condition three. And the feedback loop on a marking error is weak: a teacher who systematically misreads a cohort's misconception may find out at the end of a key stage, or never, which is condition four.
European law has already reached the same place from a different direction. Annex III of Regulation (EU) 2024/1689 classifies as high risk any AI system intended to evaluate learning outcomes, including where those outcomes steer a pupil's learning, together with systems determining admission and systems monitoring pupils during tests. Marking, in other words, is the part a regulator singled out, and the requirements that follow include documented human oversight under Article 14.
Which makes the DfE's 5 per cent figure the most encouraging number in this page. The profession's own instinct has so far put the tool where the risk is lowest. That instinct is not written down anywhere, it is not protected by policy in four schools out of five, and it is weakest among the teachers with the least experience.
The trials nobody has run#
- No study has measured teacher capability after prolonged reliance. Every finding above is about pupils or about minutes. Whether a teacher who has planned with a model for three years still plans as well without it has not been tested, and would require the unassisted baseline that almost no profession keeps.
- Kestin measured immediately. The post-tests followed the lessons. Retention was not measured, and the authors name spacing and retention studies as future work. A learning gain that survives a fortnight is a different claim from one measured on the day.
- The EEF trial is one subject, one key stage, one country. Time was self-recorded. The trial also over-represented schools in London and the South East and schools rated Outstanding, which the report states.
- Nobody has tested AI marking against blind human marking at scale. The 5 per cent adoption figure means the profession is not generating the data either.
- England has no national position. There is departmental guidance and a survey, and roughly a fifth of schools have a policy. The absence is itself a decision, and it delegates the decision to whichever tool a fourteen-year-old opens at home.
Twenty years of asking a version of this#
Rahim Hirji spent two decades in education and education technology before writing SuperSkills, and has served as a school governor and spoken in hundreds of schools. The questions on this page have a dated public trail in his newsletter well before the current debate: "Reinventing Education" in November 2017, "AI Robot Teachers" in August 2018, "AI in Education" in October 2018 and "Experiments in Education" in August 2020, which were flagging automated tutoring and automated grading years before either was available to a classroom.
In What I Tell Parents About AI in May 2026, he set out the questions he thinks a parent should put to a school, and the position underneath them: the struggle is the lesson. Removing the struggle removes the lesson.
He also makes the point that the school is rarely the villain here, because the sector has been handed no curriculum, no budget and no plan, and is improvising in public.
Six decisions a school can take this term#
- Write down where the tool is allowed, by task rather than by tool. Planning, resource generation, feedback, marking and delivery are five different risk positions. A single policy on "AI" cannot distinguish them.
- Protect marking deliberately, and say why. Not on principle. Because marking is where the diagnostic judgement is built and where nobody would notice its decay.
- Give new teachers the reps. The staff most likely to use AI for written feedback are the ones with under three years' experience. That is precisely inverted from where it should sit.
- Budget five weeks before you expect any saving. The trial that produced 31 per cent gave teachers a familiarisation period first and measured nothing during it.
- If you deploy anything to pupils, ask whether it withholds answers. That single design choice is what separated Bastani's two arms, and it is the only variable in the education evidence with a durable effect.
- Measure something other than usage. Both arms of the strongest negative study looked identical on adoption. See measuring adoption properly.
Key sources
- Education Endowment Foundation and National Foundation for Educational Research (2024). ChatGPT in lesson preparation: a Teacher Choices trial. Evaluation report, 12 December 2024.
- Department for Education and IFF Research (2025). Technology in Schools survey: 2024 to 2025. Research report, November 2025.
- Kestin, G., Miller, K., Klales, A., Milbourne, T. and Ponti, G. (2025). AI tutoring outperforms in-class active learning: an RCT introducing a novel research-based design in an authentic educational setting. Scientific Reports, 15.
- Bastani, H., Bastani, O., Sungu, A., Ge, H., Kabakci, O. and Mariman, R. (2025). Generative AI Without Guardrails Can Harm Learning. Proceedings of the National Academy of Sciences.
- Kosmyna, N. et al. (2025). Your Brain on ChatGPT: Accumulation of Cognitive Debt when Using an AI Assistant for Essay Writing. MIT Media Lab.
- Bjork, R. and Bjork, E. (2011). Making Things Hard on Yourself, But in a Good Way.
- European Parliament and Council (2024). Regulation (EU) 2024/1689, Annex III: High-risk AI systems, point 3 on education and vocational training.
Related SuperSkills research#
On the pupil side, how much teenagers should use AI, should children use AI and assessing students when AI can do the assignment. On the mechanism, how humans learn with AI and productive struggle. On the method, deskilling risk by profession, with the neighbouring cases at medicine and consulting. On policy, official guidance on AI in education.
About this research#
Rahim Hirji is the author of SuperSkills (Kogan Page, 2026), keynote speaker on AI and human capability, and founder of The SuperSkills Intelligence Company. The EEF project record, the DfE survey report and the Scientific Reports paper were each read at source and every figure on this page checked against them, including the sample sizes, the familiarisation period and the ceiling-effect caveat on the effect size. Figures circulating about teacher time savings from commercial tools are not used here, because none could be traced to a controlled trial.
Cite this
Hirji, R. (2026). How will AI change teaching? The SuperSkills Intelligence Company. Last reviewed 1 September 2026. thesuperskills.com/research/how-will-ai-change-teaching
