← Research
Research

Does using AI stop you learning?

Every study that takes the tool away before measuring finds the same split, and it does not run between users and non-users.

Last reviewed: 14 September 2026

What skill formation means and why assisted output cannot measure it, the four designs that withdraw assistance before testing, the interaction patterns that scored 65 per cent and the ones that scored under 40, two seventeens that are not the same number, and the description of the Shen and Tamkin sample that the paper's own balance table contradicts.

Question this page answersQuestion this page partly answersAll 811 questions this research covers

Using AI does not stop you learning. Handing it the part you were about to struggle with does. Four studies now separate those two things, using four different methods on four different populations, and they agree: people who kept working at their own pace while the tool was available came out roughly where they would have anyway, and people who let the tool finish the work came out measurably worse once it was taken away. The variable that separates them is what happens in the minutes the window is open.

The answer, in one line

No. Delegating the work to it does. Four studies that withdraw the assistance before measuring all split on the same line: participants who let the tool complete the task performed worse unaided afterwards, and participants who kept working at their own pace while it was available were largely unaffected.

Share as a card

Definition#

Skill formation: the process by which practice turns into capability a person still holds when the help is removed. It is measured by taking the help away and testing unaided, so assisted output tells you almost nothing about it. Every study on this page uses an unaided measure, and that is the single feature that makes them comparable.

Share this definition as a card

Four designs pointing one way#

The reason this question has felt unresolved is that most of what gets quoted measures assisted output, which rises. The four studies below all withdraw the assistance before measuring, and they run from a one-hour randomised trial on professionals to thirty months of panel data on 26,811 schoolchildren. Assembling them is the argument here. No single one of them would settle the question. Reading the set together makes access a difficult variable to keep hold of.

Fifty-two developers and an unfamiliar library#

Judy Hanwen Shen and Alex Tamkin randomised 52 professional developers, 26 in each arm, to learn Trio, a Python library for asynchronous programming that none of them had used, through a self-guided tutorial with a maximum of 35 minutes on two coding tasks. One arm had an AI assistant in the sidebar with access to their code, able to produce the correct solution on request. Then the assistant was taken away and everybody sat the same quiz on concepts they had used minutes earlier.

The assisted arm averaged 50 per cent against 67 per cent for those who had coded by hand, a difference of 4.15 points on a 27-point quiz, Cohen's d of 0.738 at p equals 0.01. The widest gap was on the debugging questions, which is the awkward part: debugging is the skill you need in order to catch the assistant when it is wrong. The assisted arm finished about two minutes faster and that difference did not reach significance, partly because several participants spent up to eleven minutes composing queries.

The finding that matters more is inside the treatment arm. The authors annotated screen recordings and grouped participants by how they used the assistant. Three patterns scored under 40 per cent: wholesale delegation, a slide into delegation after an opening question or two, and leaning on the assistant to debug. Three scored 65 per cent or higher: generating code and then asking follow-up questions about it, asking for code and explanation together, and asking only conceptual questions while writing the code themselves. Shen and Tamkin say plainly that this second analysis was not pre-registered and draws no causal link. It points at the behaviour rather than proving it.

Ten minutes, and persistence is what goes first#

Liu and colleagues ran a series of randomised trials across mathematical reasoning and reading comprehension with 1,222 participants, offering AI assistance during practice and then withdrawing it. Assisted performance improved. Unassisted performance afterwards was significantly worse, and participants were more likely to give up. Their reported threshold is about ten minutes of interaction, which is faster than anyone had assumed the effect could appear.

Their account of the mechanism is the part worth carrying. They attribute the loss to being conditioned to expect an immediate answer, so the experience of working through something on your own never happens, and they note that persistence predicts long-term learning about as well as anything does. That relocates the damage. The casualty is the disposition to keep going when an answer does not arrive, trained out inside a single session, and a forgotten fact is only the visible residue of it.

Same model, opposite outcomes#

Bastani and colleagues gave nearly a thousand high-school students access to GPT-4 for mathematics practice in three arms: unrestricted chat, a guardrailed tutor that gave hints rather than answers, and a control. While the tool was present, grades rose 48 per cent in the unrestricted arm and 127 per cent with the tutor. When access was removed, the unrestricted group scored 17 per cent lower than students who had never had it at all. The guardrailed tutor largely removed that harm.

One model, one subject, one cohort, and the outcome inverted on interface design alone. This is the most useful result in the literature because it takes the question away from whether people should use AI and puts it on how the thing in front of them is built.

Thirty months, and the bill arrives in year two#

Stromberg, Lei and Wu tracked 26,811 Chinese secondary students across thirty months, using staggered AI adoption in a difference-in-differences design against closed-book monthly exams and invigilated entrance exams. Homework scores rose 18 per cent and homework time fell 30 per cent. Monthly exam scores fell 20 per cent within six months. Entrance-exam scores fell 18 and 24 per cent, with the full penalty appearing only after roughly two years.

The split inside that result is the one this page rests on. The losses concentrate among the roughly 80 per cent of users whose pattern looks like outsourcing, indicated by very short homework time coupled with high homework marks. Students who carried on spending about as long on homework as non-users suffered small losses. The authors infer that split from time-on-task rather than observing it, adoption is self-selected, and the paper is not peer reviewed, so it is the weakest of the four causally and the strongest on duration and scale.

Two seventeens that are not the same number#

Shen and Tamkin report a 17 per cent score difference and Bastani reports a 17 per cent shortfall, and those figures will end up in the same sentence somewhere. They are different quantities. The first is 4.15 points of a 27-point quiz, a gap between two arms measured once. The second is a relative fall in grades against students who never had access, measured after withdrawal in a field setting. Running them together would turn a coincidence of arithmetic into a false replication.

What "mostly junior" meant in the balance table#

Anthropic's write-up of the Shen and Tamkin trial describes the participants as "52 (mostly junior) software engineers", and that description has travelled: secondary coverage and commentary repeat it, usually as the reason the result matters for people early in a career. The paper's own balance table says something different. Fifteen of 26 in the treatment arm and 14 of 26 in the control had seven or more years of experience. Two in each arm had one to three years. Participants were crowd workers paid a flat 150 US dollars.

The correction does not weaken the result. It arguably strengthens it, because a comprehension penalty that shows up in experienced developers is a harder thing to explain away than one in beginners. What it does change is the inference people are drawing. Read against the balance table, this describes what happens to anybody meeting unfamiliar material with an assistant standing by to finish the job, and the career stage of the person meeting it drops out. Both descriptions are on the public record and neither is hidden. They simply do not agree, and the paper is the one to quote.

Five hours in a lab and thirty months in a school#

None of this was measured at work, and that gap deserves a paragraph of its own rather than a clause at the end of somebody's summary. Three of the four studies run for under an hour on a task the participant will never do again. The fourth runs for years but on secondary-school students in one country. Nothing here follows a professional through the acquisition of a real skill over months, which is the thing everyone actually wants to know about, and until somebody runs that study the evidence here amounts to four converging proxies for it.

Two further limits belong on the page. Shen and Tamkin used a chat assistant in a sidebar, and note themselves that an agentic tool doing more of the work would be expected to produce a larger effect rather than a smaller one, which makes their result a floor rather than a ceiling. And every one of these designs tests recall or performance shortly after the tool goes away. Whether an early quiz score predicts whether somebody can still do the work in a year is a question none of them resolves.

Keeping the skill and the tool#

Four things follow, and only the first three follow from the evidence. Decide before you prompt whether this is a task you need to be able to do later, because the studies split on that and nothing else. Form a view before you ask, so that the answer arrives into a mind that already has a candidate to compare it against; the participants who did that scored 65 per cent and above. Ask for the explanation as well as the output, which was the single behaviour separating the high scorers from the low ones. And where you are responsible for how other people learn, the design of the tool is the lever: the guardrailed tutor and the plain chat window were the same model, and one of them produced a durable gain and the other a deficit.

The fourth is an argument rather than a measurement, and is marked as such. If persistence is what degrades first, then the useful thing to protect is not any particular skill but the willingness to sit with something unresolved. That follows from Liu and colleagues, it is consistent with the rest, and nothing here tests it directly.

Key sources

The mechanism underneath all four results is desirable difficulty, and the act itself is cognitive offloading. On why the loss is so hard to notice while it is happening, the illusion of competence. For the organisational version of the same split, capability debt and how juniors become senior. On recovery, can you regain a skill you have lost. The practical habits are at using AI without dependency and how humans learn with AI.

Evidence review · SS-2026-236 · Graded against the published rubric

Cite this page

Hirji, R. (2026). Does using AI stop you learning?. The SuperSkills evidence base, SS-2026-236. https://thesuperskills.com/research/does-using-ai-stop-you-learning. Last reviewed 14 September 2026.

An evidence review by Rahim Hirji, not peer-reviewed research. For a material claim, cite the underlying study as well; every study here carries its own permanent link.

How citations and IDs work
Questions answered on this page

Does using AI stop you learning?

No. Delegating the work to it does. Four studies that withdraw the assistance before measuring all split on the same line: participants who let the tool complete the task performed worse unaided afterwards, and participants who kept working at their own pace while it was available were largely unaffected. In a randomised trial of 52 developers, those who asked only conceptual questions or asked for explanations alongside generated code scored 65 per cent or above, while those who delegated scored under 40.

What did the Anthropic coding study actually find?

Judy Hanwen Shen and Alex Tamkin randomised 52 professional developers, 26 per arm, to learn the Python library Trio with or without an AI assistant, then removed the assistant and tested them. The assisted arm averaged 50 per cent against 67 per cent, a gap of 4.15 points on a 27-point quiz, Cohen's d of 0.738 at p equals 0.01, with the widest difference on debugging questions. The assisted arm finished about two minutes faster and that difference was not statistically significant.

How quickly can the effect appear?

Liu and colleagues report it after approximately ten minutes of interaction, across randomised trials with 1,222 participants on mathematical reasoning and reading comprehension. Assisted performance improved, unassisted performance afterwards was significantly worse, and participants were more likely to give up. They attribute the effect to a loss of persistence rather than to forgotten knowledge.

Can the way a tool is built change the outcome?

Yes, and it inverted the result in one experiment. Bastani and colleagues gave nearly a thousand high-school students GPT-4 in three arms. While the tool was present, grades rose 48 per cent with unrestricted chat and 127 per cent with a guardrailed tutor that gave hints rather than answers. Once access was removed, the unrestricted group scored 17 per cent below students who had never had access, and the guardrailed version largely removed that harm. Same model, opposite outcomes.

Were the developers in the Shen and Tamkin trial junior?

Mostly not, despite the description that circulates. Anthropic's write-up calls them 52 mostly junior software engineers. The paper's own balance table reports 15 of 26 in the treatment arm and 14 of 26 in the control with seven or more years of experience, and two per arm with one to three years. They were crowd workers paid a flat 150 US dollars. The correction does not weaken the finding; it widens who it applies to.

Has any of this been measured at work?

No. Three of the four studies run for under an hour on a task the participant will never repeat, and the fourth runs for thirty months but on secondary-school students in one country. No study follows a professional acquiring a real skill over months. Every design also tests shortly after the tool is withdrawn, so whether an early score predicts capability a year later is unresolved.

In this hub

Thinking, learning and capability

What sustained AI use does to thinking, and how capability is built and kept.

Ask the evidence
What does the evidence actually show?What should our board be asking about this?Where does Rahim disagree with the consensus?
Bring this into your organisation

If this describes something happening in your teams, say so.

Keynotes, board sessions and advisory work, drawing on research across more than 200 organisations in 30 countries. Tell me the room, the date and the shift you need. A reply within 24 hours.

Start a conversation

Topics and audiences  ·  All research

Box of Amazing

Rahim’s free weekly letter on AI and human capability

If this was useful, the weekly letter is where the thinking happens first. Most of what ends up on this site starts there. Weekly essays on AI, capability and the future of work. Read by 25,000 people, every week since 2017. Free, and one click to stop.

Opens Substack to confirm. No pitch in it, unsubscribe in one click, and nobody follows up because you read something.

Running an event, or responsible for how AI arrives in your organisation? Keynotes  ·  Advisory for CEOs and boards  ·  Enquire