- Is an AI tutor as good as a human tutor?
- Should I use an AI tutor instead of a human one to revise for an exam?
In one large randomised trial, on an immediate test, yes. StudentBench, a preprint posted on 23 September 2026 by Handshake AI Research, randomised 2,383 adults to an hour of AI tutoring, an hour with an expert human tutor on video, or an hour of videos, and found the AI arm’s GRE gain statistically equivalent to the human arm’s at a fraction of the cost. It joins a Harvard physics trial in which an engineered AI tutor beat the class. But every such result is measured minutes after the session. The one field experiment that waited found students who had used an unrestricted assistant 17 per cent worse off when it was withdrawn, and the guardrailed tutor arm neither better nor worse. Whether an AI tutor is as good as a human one depends on what you want left behind.
The answer, in one line
On a test sat straight after the session, one large randomised trial says yes: StudentBench, a September 2026 preprint from Handshake AI Research with 2,383 adults, found an hour of AI tutoring statistically equivalent to an hour with an expert human tutor on GRE gains, within a quarter of a standard deviation, and a Harvard physics trial found an expert-built AI tutor beat the class.
The trial, and what it compared#
The StudentBench paper is by Curtis Northcutt, Inaara Hasmani, Kevin Feng, Trevor Khangi, Andreas Plesner and Jonas Mueller; Handshake AI, a company that supplies human experts to AI developers, funded the work and recruited about 70,000 students from its own platform to reach the 2,383 who took part. Each sat a 27-question GRE pre-test, spent one hour in the condition they were assigned, and sat a post-test of different questions with no tutor present. Thirteen AI tutors were tested, from Google, OpenAI, Anthropic and others, doing lesson planning, conversational tutoring and practice problems in text. The human tutors were expert GRE tutors on live video. The control watched videos. Against control, AI tutoring added 6.86 percentage points on the quantitative section and 5.47 on the verbal, with confidence intervals well clear of zero. Against the human tutors, the AI tutors’ gains fell within a quarter of a standard deviation, about four percentage points, either way, and the authors report equivalence at p = .015. In five of seven domains the best AI tutor exceeded the human average. The smallest model to reach equivalence did so, on the authors’ costing, at 918 times lower cost per point gained.
What the equivalence test does and does not say#
An equivalence test reverses the usual burden. Rather than asking whether two things differ, it asks whether the difference between them is small enough to ignore, and the researcher sets the size of small in advance. Here small was a quarter of a standard deviation. So the finding is that an hour of AI tutoring and an hour with a paid expert produced gains within four points of each other, on average, on a test sat straight afterwards. It does not say the two hours were the same. The human tutors were the ones the company recruited at a reference rate of 75 dollars an hour, so the bar is whatever those tutors were. And there was no arm in which students worked through practice questions on their own, which the authors name as a limitation: the gain that either kind of tutor adds over self-study is not in the data. The paper is not peer reviewed and its authors’ employer sells the human expertise the AI tutors were compared against, which cuts in no obvious direction but should be known.
The earlier trials#
Two published trials set the frame. Kestin and colleagues, in Scientific Reports in June 2025, ran a randomised crossover in Harvard’s largest introductory physics course: 194 students, each topic taught either by in-class active learning or by an AI tutor at home built with expert-written scaffolds and pre-written answers. The median learning gain in the AI condition was more than double the class’s, in less time. The tutor was heavily engineered by subject experts; it was not a chatbot with a syllabus. Bastani and colleagues, in PNAS in June 2025, gave nearly a thousand Turkish high-school students either unrestricted GPT-4, a hints-only tutor, or nothing. While the tools were present, grades rose 48 per cent with unrestricted access and 127 per cent with the tutor. When they were withdrawn, the unrestricted group scored 17 per cent below students who had never had the tool. The guardrailed tutor group showed no such loss, and no lasting gain. The whole result is on how humans learn with AI.
The benchmark the word tutor carries#
Since Benjamin Bloom’s 1984 paper in Educational Researcher, one-to-one tutoring has been the standard everything else is measured against: the tutored students in the two dissertations he drew on scored about two standard deviations above a conventional class, and he asked how group instruction could ever match it. The figure did not survive replication at that size. Kurt VanLehn’s 2011 review in Educational Psychologist put human tutoring at 0.79 standard deviations over no tutoring and the best computer tutors of that era at 0.76, and concluded that the two-sigma effect was not observed. So when a 2026 trial finds AI tutoring equivalent to human tutoring, the human bar it clears is closer to 0.8 than to 2. That is still a large effect in education. It is not the folk figure.
What “as good as” leaves out#
Three things are missing from every equivalence result, and the StudentBench authors name the first themselves: they “measured immediate learning gains without evaluating whether they persist over months”. The second is transfer, whether a gain on GRE items moves to any other task; nothing was tested. The third is what the estate is about: what the learner can do when the tutor is not there. Bastani’s withdrawal test is the only one of the three studies to ask it, and its answer split by design. An assistant that gives answers produced gains that reversed into a loss; a tutor that gave hints produced gains that vanished without harm. A human tutor who does a student’s homework would be judged a bad tutor whatever the post-test said. The same standard applies to a machine, and the immediate post-test cannot apply it. The mechanisms are on productive struggle and the illusion of competence.
For a learner, a parent, or whoever pays for the tutoring#
The evidence supports using an AI tutor for practice with feedback, which is what all three trials tested, and supports the guardrailed design over the open one. It does not support judging the tutor by how the session felt or by the score straight afterwards. The test that matters is the one Bastani ran: after a few weeks, sit a set of questions with everything closed and compare it with where you started. If the score holds, the tutoring built something. If it falls back, the tool was doing the work. The same test belongs in any school’s or employer’s contract with a tutoring provider, human or machine. That is the estate’s position in a classroom: the risk that can be measured sits in the handover of the work to the machine and in what remains of the learner’s capability afterwards, and the decision a learner or a school controls is which parts of the work the machine may do, who can switch it off, what the student must still be able to do unaided, and how anyone would know if it had stopped working. Rules Before Tools, applied to a revision session.
What this does not show#
It does not show that AI tutors are as good as human tutors in general. StudentBench measured adults preparing for one standardised test, in one hour, in English, on a text interface, with an immediate post-test and no retention or transfer measure, in a preprint funded by a company with a commercial interest in the question, and its authors say the retention study is still to come. Kestin’s result is one course with a tutor built by experts for that course. Bastani’s is school mathematics over a bounded period. None of the three compares an AI tutor with a human tutor over a term, and no study yet does. The 918-fold cost figure is the authors’ and depends on a 75-dollar hourly rate and on their choice of model. And it does not show that human tutoring is the gold standard; VanLehn’s review is fifteen years old and the field has not re-run it at scale.
Essay · SS-2026-304
Hirji, R. (2026). Is an AI tutor as good as a human tutor?. The SuperSkills evidence base, SS-2026-304. https://thesuperskills.com/research/is-an-ai-tutor-as-good-as-a-human-tutor. Last reviewed 24 September 2026.
An evidence review by Rahim Hirji, not peer-reviewed research. For a material claim, cite the underlying study as well; every study here carries its own permanent link.
How citations and IDs work