Using AI does not stop you learning. Handing it the part you were about to struggle with does. Four studies now separate those two things, using four different methods on four different populations, and they agree: people who kept working at their own pace while the tool was available came out roughly where they would have anyway, and people who let the tool finish the work came out measurably worse once it was taken away. The variable that separates them is what happens in the minutes the window is open.
The answer, in one line
No. Delegating the work to it does. Four studies that withdraw the assistance before measuring all split on the same line: participants who let the tool complete the task performed worse unaided afterwards, and participants who kept working at their own pace while it was available were largely unaffected.
Definition#
Skill formation: the process by which practice turns into capability a person still holds when the help is removed. It is measured by taking the help away and testing unaided, so assisted output tells you almost nothing about it. Every study on this page uses an unaided measure, and that is the single feature that makes them comparable.
Four designs pointing one way#
The reason this question has felt unresolved is that most of what gets quoted measures assisted output, which rises. The four studies below all withdraw the assistance before measuring, and they run from a one-hour randomised trial on professionals to thirty months of panel data on 26,811 schoolchildren. Assembling them is the argument here. No single one of them would settle the question. Reading the set together makes access a difficult variable to keep hold of.
Fifty-two developers and an unfamiliar library#
Judy Hanwen Shen and Alex Tamkin randomised 52 professional developers, 26 in each arm, to learn Trio, a Python library for asynchronous programming that none of them had used, through a self-guided tutorial with a maximum of 35 minutes on two coding tasks. One arm had an AI assistant in the sidebar with access to their code, able to produce the correct solution on request. Then the assistant was taken away and everybody sat the same quiz on concepts they had used minutes earlier.
The assisted arm averaged 50 per cent against 67 per cent for those who had coded by hand, a difference of 4.15 points on a 27-point quiz, Cohen's d of 0.738 at p equals 0.01. The widest gap was on the debugging questions, which is the awkward part: debugging is the skill you need in order to catch the assistant when it is wrong. The assisted arm finished about two minutes faster and that difference did not reach significance, partly because several participants spent up to eleven minutes composing queries.
The finding that matters more is inside the treatment arm. The authors annotated screen recordings and grouped participants by how they used the assistant. Three patterns scored under 40 per cent: wholesale delegation, a slide into delegation after an opening question or two, and leaning on the assistant to debug. Three scored 65 per cent or higher: generating code and then asking follow-up questions about it, asking for code and explanation together, and asking only conceptual questions while writing the code themselves. Shen and Tamkin say plainly that this second analysis was not pre-registered and draws no causal link. It points at the behaviour rather than proving it.
Ten minutes, and persistence is what goes first#
Liu and colleagues ran a series of randomised trials across mathematical reasoning and reading comprehension with 1,222 participants, offering AI assistance during practice and then withdrawing it. Assisted performance improved. Unassisted performance afterwards was significantly worse, and participants were more likely to give up. Their reported threshold is about ten minutes of interaction, which is faster than anyone had assumed the effect could appear.
Their account of the mechanism is the part worth carrying. They attribute the loss to being conditioned to expect an immediate answer, so the experience of working through something on your own never happens, and they note that persistence predicts long-term learning about as well as anything does. That relocates the damage. The casualty is the disposition to keep going when an answer does not arrive, trained out inside a single session, and a forgotten fact is only the visible residue of it.
Same model, opposite outcomes#
Bastani and colleagues gave nearly a thousand high-school students access to GPT-4 for mathematics practice in three arms: unrestricted chat, a guardrailed tutor that gave hints rather than answers, and a control. While the tool was present, grades rose 48 per cent in the unrestricted arm and 127 per cent with the tutor. When access was removed, the unrestricted group scored 17 per cent lower than students who had never had it at all. The guardrailed tutor largely removed that harm.
One model, one subject, one cohort, and the outcome inverted on interface design alone. This is the most useful result in the literature because it takes the question away from whether people should use AI and puts it on how the thing in front of them is built.
Thirty months, and the bill arrives in year two#
Stromberg, Lei and Wu tracked 26,811 Chinese secondary students across thirty months, using staggered AI adoption in a difference-in-differences design against closed-book monthly exams and invigilated entrance exams. Homework scores rose 18 per cent and homework time fell 30 per cent. Monthly exam scores fell 20 per cent within six months. Entrance-exam scores fell 18 and 24 per cent, with the full penalty appearing only after roughly two years.
The split inside that result is the one this page rests on. The losses concentrate among the roughly 80 per cent of users whose pattern looks like outsourcing, indicated by very short homework time coupled with high homework marks. Students who carried on spending about as long on homework as non-users suffered small losses. The authors infer that split from time-on-task rather than observing it, adoption is self-selected, and the paper is not peer reviewed, so it is the weakest of the four causally and the strongest on duration and scale.
Two seventeens that are not the same number#
Shen and Tamkin report a 17 per cent score difference and Bastani reports a 17 per cent shortfall, and those figures will end up in the same sentence somewhere. They are different quantities. The first is 4.15 points of a 27-point quiz, a gap between two arms measured once. The second is a relative fall in grades against students who never had access, measured after withdrawal in a field setting. Running them together would turn a coincidence of arithmetic into a false replication.
What "mostly junior" meant in the balance table#
Anthropic's write-up of the Shen and Tamkin trial describes the participants as "52 (mostly junior) software engineers", and that description has travelled: secondary coverage and commentary repeat it, usually as the reason the result matters for people early in a career. The paper's own balance table says something different. Fifteen of 26 in the treatment arm and 14 of 26 in the control had seven or more years of experience. Two in each arm had one to three years. Participants were crowd workers paid a flat 150 US dollars.
The correction does not weaken the result. It arguably strengthens it, because a comprehension penalty that shows up in experienced developers is a harder thing to explain away than one in beginners. What it does change is the inference people are drawing. Read against the balance table, this describes what happens to anybody meeting unfamiliar material with an assistant standing by to finish the job, and the career stage of the person meeting it drops out. Both descriptions are on the public record and neither is hidden. They simply do not agree, and the paper is the one to quote.
Five hours in a lab and thirty months in a school#
None of this was measured at work, and that gap deserves a paragraph of its own rather than a clause at the end of somebody's summary. Three of the four studies run for under an hour on a task the participant will never do again. The fourth runs for years but on secondary-school students in one country. Nothing here follows a professional through the acquisition of a real skill over months, which is the thing everyone actually wants to know about, and until somebody runs that study the evidence here amounts to four converging proxies for it.
Two further limits belong on the page. Shen and Tamkin used a chat assistant in a sidebar, and note themselves that an agentic tool doing more of the work would be expected to produce a larger effect rather than a smaller one, which makes their result a floor rather than a ceiling. And every one of these designs tests recall or performance shortly after the tool goes away. Whether an early quiz score predicts whether somebody can still do the work in a year is a question none of them resolves.
Keeping the skill and the tool#
Four things follow, and only the first three follow from the evidence. Decide before you prompt whether this is a task you need to be able to do later, because the studies split on that and nothing else. Form a view before you ask, so that the answer arrives into a mind that already has a candidate to compare it against; the participants who did that scored 65 per cent and above. Ask for the explanation as well as the output, which was the single behaviour separating the high scorers from the low ones. And where you are responsible for how other people learn, the design of the tool is the lever: the guardrailed tutor and the plain chat window were the same model, and one of them produced a durable gain and the other a deficit.
The fourth is an argument rather than a measurement, and is marked as such. If persistence is what degrades first, then the useful thing to protect is not any particular skill but the willingness to sit with something unresolved. That follows from Liu and colleagues, it is consistent with the rest, and nothing here tests it directly.
Key sources
- Shen, J. H. and Tamkin, A. (2026). How AI Impacts Skill Formation. arXiv:2601.20245, submitted 28 January 2026. Pre-registered at osf.io/pk6a5. The arm means of 50 and 67 per cent appear in Anthropic's write-up of 29 January 2026; the paper's own text gives the 4.15-point gap. Graded entry.
- Liu, G., Christian, B., Dumbalska, T., Bakker, M. A. and Dubey, R. (2026). AI Assistance Reduces Persistence and Hurts Independent Performance. arXiv:2604.04721. Preprint. Graded entry.
- Bastani, H., Bastani, O., Sungu, A., Ge, H., Kabakı, Ö. and Mariman, R. (2025). Generative AI Without Guardrails Can Harm Learning. PNAS, 122(26), e2422633122. Graded entry.
- Stromberg, D., Lei, V. and Wu, Y. (2026). The Generative AI Learning Penalty: Evidence from Chinese Secondary Education. CEPR Discussion Paper 21577. Working paper. Graded entry.
Related SuperSkills research#
The mechanism underneath all four results is desirable difficulty, and the act itself is cognitive offloading. On why the loss is so hard to notice while it is happening, the illusion of competence. For the organisational version of the same split, capability debt and how juniors become senior. On recovery, can you regain a skill you have lost. The practical habits are at using AI without dependency and how humans learn with AI.
Evidence review · SS-2026-236 · Graded against the published rubric
Hirji, R. (2026). Does using AI stop you learning?. The SuperSkills evidence base, SS-2026-236. https://thesuperskills.com/research/does-using-ai-stop-you-learning. Last reviewed 14 September 2026.
An evidence review by Rahim Hirji, not peer-reviewed research. For a material claim, cite the underlying study as well; every study here carries its own permanent link.
How citations and IDs work