← Research
Research

How will AI change teaching?

The workload evidence and the learning evidence are answering different questions, and the profession has quietly put the tool in the one place where the risk is lowest.

Last reviewed: 1 September 2026

The EEF planning trial, what English teachers actually use it for, the Harvard tutor result read with its methods attached, the experiment that took the tutor away, and why marking is the exposed part.

Questions this page answersAll 435 questions this research covers

Mostly by taking the preparation and leaving the room. The strongest evidence in this profession measures teacher workload rather than pupil learning: a school-randomised trial in England found lesson planning time falling by 31 per cent with no detectable change in the quality of what was produced. The evidence on pupils splits, and it splits on whether the tool was engineered to teach or simply made available. A tutor built by a physics lecturer to follow known pedagogy beat his own active-learning class. A chatbot handed to teenagers raised their marks while they had it and left them behind the control group once it was taken away.

The trial asked the question teachers were actually asking#

The Education Endowment Foundation funded a Teacher Choices trial, evaluated independently by the National Foundation for Educational Research, on a deliberately narrow question: does using ChatGPT for Key Stage 3 science lesson preparation reduce the time it takes. 259 teachers across 68 schools in England took part over ten weeks in the summer term of 2024. One arm used ChatGPT with a written guide; the other was asked to use no generative AI at all.

Weekly planning time for the relevant classes came out at 56.2 minutes in the ChatGPT arm against 81.5 minutes in the comparison arm, a saving of 25.3 minutes and a reduction of 31 per cent. The EEF gives the result a high security rating. An expert panel reviewed the resulting lesson resources without being told which had been produced with the tool and found no noticeable difference in quality.

Three details in the report matter more than the headline and travel less well.

What the trial does not establish is anything about teaching. It measured minutes spent preparing one subject at one key stage. Whether the lessons worked better for pupils was outside its scope, and the resource quality check was a blinded panel reading materials rather than anyone watching a class.

Thirty-five per cent plan with it, five per cent mark with it#

The Department for Education's Technology in Schools survey for 2024 to 2025, conducted by IFF Research across 1,634 schools and covering 795 leaders, 1,211 teachers and 489 IT leads, put numbers on where the profession has actually put the tool. 44 per cent of teachers reported using generative AI for school activities. Within that: lesson planning 35 per cent, delivering live lessons 7 per cent, marking 5 per cent.

The distribution is doing something interesting. Teachers have concentrated the tool on the part of the job that happens before anyone is in front of them, and have almost entirely kept it out of the two parts where a judgement about a particular child is being made. Nobody instructed them to. Only around a fifth of schools had a policy on safe and appropriate use at all, and the survey's own interviews found informal guidance more common than formal policy.

The generational split is the part worth watching. Teachers under 35 used it for planning at 43 per cent against 32 per cent for older colleagues, and for written feedback at 21 per cent against 12 per cent. Teachers with under three years in the classroom used it for written feedback at 27 per cent against 14 per cent for everyone else. The people leaning hardest on it for feedback are the people who have written the fewest reports by hand. That is the missing rungs pattern arriving in a staffroom.

School leaders were also more than twice as likely to be planning investment in AI tools for teachers than in tools for pupils, 58 per cent against 20 per cent. On the pupil side the picture is defensive: 73 per cent of secondary teachers thought pupils had used it for homework, 77 per cent of secondary leaders whose pupils could access it reported issues, and the most common issue reported was plagiarism at 67 per cent.

The best result for an AI tutor came from a lecturer who built one#

Kestin and colleagues ran a randomised crossover experiment in Harvard's largest introductory physics course. Of 233 enrolled students, 194 were eligible, and each experienced both conditions: one topic taught in an active-learning class, another taught at home by a purpose-built tutor, with the conditions reversed the following week.

Median post-test score was 4.5 in the AI condition against 3.5 in the classroom condition, measured against a combined pre-test median of 2.75. Median learning gain in the AI condition was over double. A rank-sum test gave z of -5.6 at p below 10 to the minus 8. The regression estimate of effect size, 0.63, is described by the authors as an underestimate because of a ceiling effect; a quantile regression puts the range at 0.73 to 1.3 standard deviations. Median time on task was 49 minutes against the 60 minutes assumed for the class, and time on task did not correlate with score. Students reported higher engagement, 4.1 against 3.6, and higher motivation, 3.4 against 3.1. Enjoyment and growth mindset showed no significant difference.

Read the methods and the result becomes narrower and more useful. The tutor's accuracy relied on pre-written answers. The prompts were question-specific and written by instructors who knew the content deeply. The lessons carried high-quality instructional video and a structured scaffold that governed how the tutor was allowed to respond. The authors state plainly that they do not presume structured AI tutoring will outperform in-class active learning in all contexts, naming complex synthesis and higher-order critical thinking as the likely exceptions, and they position the finding at the understanding, applying and analysing levels of Bloom's taxonomy.

So the study is evidence that a well-designed tutor beats a well-designed class at the delivery stage of new material. It is not evidence about a pupil opening a chatbot. The gap between those two things is the entire argument about AI in education, and almost every citation of this paper closes it silently.

The experiment that took the tutor away again#

Bastani and colleagues ran a field experiment with Turkish high school students. Marks rose 48 per cent with unrestricted access to a GPT-4 assistant and 127 per cent with a guardrailed tutor designed to withhold answers. Then access was removed for the exam. The unrestricted group scored 17 per cent below students who had never had the tool at all. The guardrailed arm was largely spared that penalty.

These two studies are usually presented as a contradiction and they are not one. Kestin measured performance while a carefully engineered tutor was present. Bastani measured what remained after an ordinary one was withdrawn. Both found the design of the tool mattering more than its presence, and both found the guardrailed configuration, the one that makes the student do the work, producing the durable result. The distinction this research has been drawing at desirable difficulty and productive struggle is the same distinction, arrived at from the other direction.

The uncomfortable corollary for a school: both of Bastani's arms would have looked identical on any usage dashboard. Adoption metrics cannot see the difference between the configuration that taught and the configuration that did the work instead. This research calls that gap usage theatre, and education is where it is cheapest to fall into.

Marking is the exposed part, and teachers have found that out first#

Applying the four conditions from the deskilling risk test, teaching scores medium and asymmetrically. Preparation is a complement: the material is assembled and the teacher still decides what to do with it. The diagnostic judgement of what a particular child has misunderstood is not something current tools substitute for well, because the evidence for that judgement is in the room rather than in the text.

Marking is where the conditions converge. The tool can produce the judgement rather than the material, which is condition one. Marking a set of books is one of the ways a new teacher builds a picture of what a class actually knows, which is condition two. Nobody measures a teacher's unassisted marking accuracy, which is condition three. And the feedback loop on a marking error is weak: a teacher who systematically misreads a cohort's misconception may find out at the end of a key stage, or never, which is condition four.

European law has already reached the same place from a different direction. Annex III of Regulation (EU) 2024/1689 classifies as high risk any AI system intended to evaluate learning outcomes, including where those outcomes steer a pupil's learning, together with systems determining admission and systems monitoring pupils during tests. Marking, in other words, is the part a regulator singled out, and the requirements that follow include documented human oversight under Article 14.

Which makes the DfE's 5 per cent figure the most encouraging number in this page. The profession's own instinct has so far put the tool where the risk is lowest. That instinct is not written down anywhere, it is not protected by policy in four schools out of five, and it is weakest among the teachers with the least experience.

The trials nobody has run#

Twenty years of asking a version of this#

Rahim Hirji spent two decades in education and education technology before writing SuperSkills, and has served as a school governor and spoken in hundreds of schools. The questions on this page have a dated public trail in his newsletter well before the current debate: "Reinventing Education" in November 2017, "AI Robot Teachers" in August 2018, "AI in Education" in October 2018 and "Experiments in Education" in August 2020, which were flagging automated tutoring and automated grading years before either was available to a classroom.

In What I Tell Parents About AI in May 2026, he set out the questions he thinks a parent should put to a school, and the position underneath them: the struggle is the lesson. Removing the struggle removes the lesson. He also makes the point that the school is rarely the villain here, because the sector has been handed no curriculum, no budget and no plan, and is improvising in public.

Six decisions a school can take this term#

Key sources

On the pupil side, how much teenagers should use AI, should children use AI and assessing students when AI can do the assignment. On the mechanism, how humans learn with AI and productive struggle. On the method, deskilling risk by profession, with the neighbouring cases at medicine and consulting. On policy, official guidance on AI in education.

About this research#

Rahim Hirji is the author of SuperSkills (Kogan Page, 2026), keynote speaker on AI and human capability, and founder of The SuperSkills Intelligence Company. The EEF project record, the DfE survey report and the Scientific Reports paper were each read at source and every figure on this page checked against them, including the sample sizes, the familiarisation period and the ceiling-effect caveat on the effect size. Figures circulating about teacher time savings from commercial tools are not used here, because none could be traced to a controlled trial.

How this research works  ·  Reviewed quarterly  ·  Found an error? Tell me and it is corrected on the page.

Cite this

Hirji, R. (2026). How will AI change teaching? The SuperSkills Intelligence Company. Last reviewed 1 September 2026. thesuperskills.com/research/how-will-ai-change-teaching

In this hub

Professions and sectors

Where the pressure lands first, profession by profession.

The work

Where the writing comes from.

These essays draw on research across more than 200 organisations in 30 countries. See the wider body of work, or bring it into your organisation.

All research →
Box of Amazing

Rahim’s free weekly letter on AI and human capability

If this was useful, the weekly letter is where the thinking happens first. Most of what ends up on this site starts there. Weekly essays on AI, capability and the future of work. Read by 25,000 people, every week since 2017. Free, and one click to stop.

Opens Substack to confirm. No pitch in it, unsubscribe in one click, and nobody follows up because you read something.

Running an event, or responsible for how AI arrives in your organisation? Keynotes  ·  Advisory and coaching  ·  Enquire