← Research
Research · Question

Is an AI tutor as good as a human tutor?

The StudentBench equivalence result, the Harvard and Turkish trials that frame it, the two-sigma folk figure and what replaced it, and the withdrawal test that none of the equivalence studies has run.

Last reviewed: 24 September 2026 · Next review due: 24 September 2027

On an immediate test, in one 2,383-person randomised trial, an hour of AI tutoring matched an hour with an expert human tutor on GRE gains, at a fraction of the cost. No equivalence trial has measured what remains weeks later, and the one experiment that withdrew the tool found a 17 per cent loss for unrestricted use. An evidence review by Rahim Hirji; every figure resolves to a graded entry in the evidence base that says what it does not show.

Questions this page answersAll 895 questions this research covers

In one large randomised trial, on an immediate test, yes. StudentBench, a preprint posted on 23 September 2026 by Handshake AI Research, randomised 2,383 adults to an hour of AI tutoring, an hour with an expert human tutor on video, or an hour of videos, and found the AI arm’s GRE gain statistically equivalent to the human arm’s at a fraction of the cost. It joins a Harvard physics trial in which an engineered AI tutor beat the class. But every such result is measured minutes after the session. The one field experiment that waited found students who had used an unrestricted assistant 17 per cent worse off when it was withdrawn, and the guardrailed tutor arm neither better nor worse. Whether an AI tutor is as good as a human one depends on what you want left behind.

The answer, in one line

On a test sat straight after the session, one large randomised trial says yes: StudentBench, a September 2026 preprint from Handshake AI Research with 2,383 adults, found an hour of AI tutoring statistically equivalent to an hour with an expert human tutor on GRE gains, within a quarter of a standard deviation, and a Harvard physics trial found an expert-built AI tutor beat the class.

Share as a card

The trial, and what it compared#

The StudentBench paper is by Curtis Northcutt, Inaara Hasmani, Kevin Feng, Trevor Khangi, Andreas Plesner and Jonas Mueller; Handshake AI, a company that supplies human experts to AI developers, funded the work and recruited about 70,000 students from its own platform to reach the 2,383 who took part. Each sat a 27-question GRE pre-test, spent one hour in the condition they were assigned, and sat a post-test of different questions with no tutor present. Thirteen AI tutors were tested, from Google, OpenAI, Anthropic and others, doing lesson planning, conversational tutoring and practice problems in text. The human tutors were expert GRE tutors on live video. The control watched videos. Against control, AI tutoring added 6.86 percentage points on the quantitative section and 5.47 on the verbal, with confidence intervals well clear of zero. Against the human tutors, the AI tutors’ gains fell within a quarter of a standard deviation, about four percentage points, either way, and the authors report equivalence at p = .015. In five of seven domains the best AI tutor exceeded the human average. The smallest model to reach equivalence did so, on the authors’ costing, at 918 times lower cost per point gained.

What the equivalence test does and does not say#

An equivalence test reverses the usual burden. Rather than asking whether two things differ, it asks whether the difference between them is small enough to ignore, and the researcher sets the size of small in advance. Here small was a quarter of a standard deviation. So the finding is that an hour of AI tutoring and an hour with a paid expert produced gains within four points of each other, on average, on a test sat straight afterwards. It does not say the two hours were the same. The human tutors were the ones the company recruited at a reference rate of 75 dollars an hour, so the bar is whatever those tutors were. And there was no arm in which students worked through practice questions on their own, which the authors name as a limitation: the gain that either kind of tutor adds over self-study is not in the data. The paper is not peer reviewed and its authors’ employer sells the human expertise the AI tutors were compared against, which cuts in no obvious direction but should be known.

The earlier trials#

Two published trials set the frame. Kestin and colleagues, in Scientific Reports in June 2025, ran a randomised crossover in Harvard’s largest introductory physics course: 194 students, each topic taught either by in-class active learning or by an AI tutor at home built with expert-written scaffolds and pre-written answers. The median learning gain in the AI condition was more than double the class’s, in less time. The tutor was heavily engineered by subject experts; it was not a chatbot with a syllabus. Bastani and colleagues, in PNAS in June 2025, gave nearly a thousand Turkish high-school students either unrestricted GPT-4, a hints-only tutor, or nothing. While the tools were present, grades rose 48 per cent with unrestricted access and 127 per cent with the tutor. When they were withdrawn, the unrestricted group scored 17 per cent below students who had never had the tool. The guardrailed tutor group showed no such loss, and no lasting gain. The whole result is on how humans learn with AI.

The benchmark the word tutor carries#

Since Benjamin Bloom’s 1984 paper in Educational Researcher, one-to-one tutoring has been the standard everything else is measured against: the tutored students in the two dissertations he drew on scored about two standard deviations above a conventional class, and he asked how group instruction could ever match it. The figure did not survive replication at that size. Kurt VanLehn’s 2011 review in Educational Psychologist put human tutoring at 0.79 standard deviations over no tutoring and the best computer tutors of that era at 0.76, and concluded that the two-sigma effect was not observed. So when a 2026 trial finds AI tutoring equivalent to human tutoring, the human bar it clears is closer to 0.8 than to 2. That is still a large effect in education. It is not the folk figure.

What “as good as” leaves out#

Three things are missing from every equivalence result, and the StudentBench authors name the first themselves: they “measured immediate learning gains without evaluating whether they persist over months”. The second is transfer, whether a gain on GRE items moves to any other task; nothing was tested. The third is what the estate is about: what the learner can do when the tutor is not there. Bastani’s withdrawal test is the only one of the three studies to ask it, and its answer split by design. An assistant that gives answers produced gains that reversed into a loss; a tutor that gave hints produced gains that vanished without harm. A human tutor who does a student’s homework would be judged a bad tutor whatever the post-test said. The same standard applies to a machine, and the immediate post-test cannot apply it. The mechanisms are on productive struggle and the illusion of competence.

For a learner, a parent, or whoever pays for the tutoring#

The evidence supports using an AI tutor for practice with feedback, which is what all three trials tested, and supports the guardrailed design over the open one. It does not support judging the tutor by how the session felt or by the score straight afterwards. The test that matters is the one Bastani ran: after a few weeks, sit a set of questions with everything closed and compare it with where you started. If the score holds, the tutoring built something. If it falls back, the tool was doing the work. The same test belongs in any school’s or employer’s contract with a tutoring provider, human or machine. That is the estate’s position in a classroom: the risk that can be measured sits in the handover of the work to the machine and in what remains of the learner’s capability afterwards, and the decision a learner or a school controls is which parts of the work the machine may do, who can switch it off, what the student must still be able to do unaided, and how anyone would know if it had stopped working. Rules Before Tools, applied to a revision session.

What this does not show#

It does not show that AI tutors are as good as human tutors in general. StudentBench measured adults preparing for one standardised test, in one hour, in English, on a text interface, with an immediate post-test and no retention or transfer measure, in a preprint funded by a company with a commercial interest in the question, and its authors say the retention study is still to come. Kestin’s result is one course with a tutor built by experts for that course. Bastani’s is school mathematics over a bounded period. None of the three compares an AI tutor with a human tutor over a term, and no study yet does. The 918-fold cost figure is the authors’ and depends on a 75-dollar hourly rate and on their choice of model. And it does not show that human tutoring is the gold standard; VanLehn’s review is fifteen years old and the field has not re-run it at scale.

Essay · SS-2026-304

Cite this page

Hirji, R. (2026). Is an AI tutor as good as a human tutor?. The SuperSkills evidence base, SS-2026-304. https://thesuperskills.com/research/is-an-ai-tutor-as-good-as-a-human-tutor. Last reviewed 24 September 2026.

An evidence review by Rahim Hirji, not peer-reviewed research. For a material claim, cite the underlying study as well; every study here carries its own permanent link.

How citations and IDs work
Questions answered on this page

Is an AI tutor as good as a human tutor?

On a test sat straight after the session, one large randomised trial says yes: StudentBench, a September 2026 preprint from Handshake AI Research with 2,383 adults, found an hour of AI tutoring statistically equivalent to an hour with an expert human tutor on GRE gains, within a quarter of a standard deviation, and a Harvard physics trial found an expert-built AI tutor beat the class. No trial has yet compared the two on what students retain weeks later, and the one field experiment that withdrew the AI found unrestricted users 17 per cent below students who never had it.

What did the StudentBench study find?

Against a control group that watched videos, one hour of AI tutoring added 6.86 percentage points on GRE quantitative questions and 5.47 on verbal, and the gains were equivalent to those from expert human tutors on video within plus or minus 0.25 standard deviations, about four points, at p = .015. In five of seven domains the best AI tutor exceeded the human average, and the cheapest model to reach equivalence did so at 918 times lower cost per point on the authors' costing. The study is not peer reviewed, was funded by Handshake AI, measured only immediate gains, and had no self-study arm.

Is one-to-one human tutoring really worth two standard deviations?

Not on the later evidence. Benjamin Bloom's 1984 paper reported tutored students about two standard deviations above a conventional class, from two dissertations, and set the benchmark the field still quotes. Kurt VanLehn's 2011 review in Educational Psychologist put human tutoring at 0.79 standard deviations over no tutoring and the best computer tutors of the time at 0.76, and found the two-sigma effect was not observed. An AI tutor equivalent to a human one is clearing a bar closer to 0.8 than to 2.

How should I judge whether an AI tutor is working?

Not by the score straight after the session or by how it felt. Wait a few weeks, sit a set of questions with every tool closed, and compare with where you started; if the score holds the tutoring built something, and if it falls back the tool was doing the work. Prefer a tutor that gives hints and asks questions over one that gives answers, which is the design difference that separated harm from no harm in the Turkish high-school experiment, and put the same withdrawal test into any contract with a tutoring provider, human or machine.

In this hub

Thinking, learning and capability

What sustained AI use does to thinking, and how capability is built and kept.

Ask the evidence
What does the evidence actually show?What should our board be asking about this?Where does Rahim disagree with the consensus?
Bring this into your organisation

If this describes something happening in your teams, say so.

Keynotes, board sessions and advisory work, drawing on research across more than 200 organisations in 30 countries. Tell me the room, the date and the shift you need. A reply within 24 hours.

Start a conversation

Topics and audiences  ·  All research

The withdrawal test, run on a school's or an employer's own tutoring or training tool, so the decision to buy rests on what remains rather than on the post-session score. Designing and running that test is the engagement. AI advisory for CEOs and boards.

Schools and universities are where this question is least theoretical. There is AI keynote for schools and education, and the full range of topics and audiences.

Box of Amazing

Rahim’s free weekly letter on AI and human capability

If this was useful, the weekly letter is where the thinking happens first. Most of what ends up on this site starts there. Weekly essays on AI, capability and the future of work. Read by 25,000 people, every week since 2017. Free, and one click to stop.

Opens Substack to confirm. No pitch in it, unsubscribe in one click, and nobody follows up because you read something.