← Research
Research

How humans learn with AI

Performance while the tool is present, and capability once it is gone, are two different measurements. Almost nobody takes the second one.

Last reviewed: 26 August 2026

How do humans learn when AI does the practice? This page separates what the learning science and the field evidence show from how Rahim Hirji reads it, and sets out the rule that keeps capability building while AI is in the room.

Humans learn by doing a thing badly, repeatedly, until they stop doing it badly. That is the whole mechanism, and it is the one AI is most efficient at removing. The best field evidence available says the effect is real and that it cuts both ways. With a generative model in front of them, people perform substantially better. With the model taken away, some of them perform worse than people who never had it at all. Performance while the tool is present and capability once it is gone are two different measurements, and AI separates them further than any tool we have had. So the question is not whether people should use AI while they learn. It is whether the work still contains the difficulty through which capability is built. Remove the difficulty and you get better output this quarter and a thinner person behind it in three years.

What the evidence shows

The cleanest result comes from a field experiment rather than a survey. Bastani and colleagues at Wharton and Penn gave nearly a thousand high-school mathematics students access to a GPT-4 tutor during practice sessions, in three arms: a plain chat interface much like ChatGPT, which they called GPT Base; a version with prompts designed to protect learning, which gave hints rather than answers, called GPT Tutor; and a control group with only textbook and notes. While the tool was available, both AI groups did far better, with grades up 48 percent for GPT Base and 127 percent for GPT Tutor. Then the researchers took the tool away and tested the students alone. The GPT Base group scored 17 percent lower than students who had never had access. Unfettered access had not merely failed to teach them. It had left them worse off than if they had struggled unaided. The guardrails in GPT Tutor largely removed that harm. The study was published in the Proceedings of the National Academy of Sciences in 2025.

Read carefully, that is not a finding about artificial intelligence. It is a finding about interface design, and it is the most useful result in the whole literature precisely because it is not a verdict. The same underlying model produced both the best learning outcome and the worst one. What differed was whether the tool made the student do the work.

The mechanism underneath was described long before any of this. Ericsson, Krampe and Tesch-Römer, writing in Psychological Review in 1993, studied violinists and pianists in Berlin and set out the concept of deliberate practice: effortful, targeted activity at the edge of current ability, with feedback, sustained over years. Not repetition, and not performance. Practice that is uncomfortable on purpose. Bjork and Bjork sharpened the point for anyone designing learning, with what they called desirable difficulties. Conditions that make study feel harder and slower, such as spacing sessions apart, interleaving different problem types, and testing yourself rather than re-reading, produce better long-term retention. Conditions that make study feel fluent and easy produce better immediate performance and worse retention later. The critical part of their work is the second half: learners consistently misjudge which is which. Fluency feels like learning. It is often the opposite.

That is the exact trap generative AI sets. It makes the work feel fluent. The answer arrives complete, well organised and plausible, and the sensation of understanding it is very close to the sensation of having produced it. A 2025 study by Microsoft Research and Carnegie Mellon, surveying 319 knowledge workers about 936 real uses of AI at work, described the shift in what the thinking consists of: from gathering information to verifying the machine's output, from solving the problem to integrating the answer, from doing the task to supervising it. The thinking does not vanish. It changes shape and it thins. A 2025 study by Michael Gerlich, across 666 participants, found a negative correlation between frequent AI use and critical-thinking scores, with cognitive offloading as the mediating mechanism and the effect strongest among the youngest users.

A team at the MIT Media Lab took a more direct measurement, comparing people writing essays with a language model, with a search engine, or unaided, while recording brain activity. The language-model group showed the weakest connectivity and the lowest sense of ownership over what they had written, and the authors named the effect cognitive debt. It is a striking result and it should be held lightly, for reasons set out in the next section.

None of this makes AI bad for learners. The opposite result is equally well established. Brynjolfsson, Li and Raymond, studying 5,179 customer-support agents, found that access to an AI assistant raised productivity by fourteen percent on average and by thirty-four percent among the newest and least experienced staff, while barely moving the most skilled. AI transfers the patterns of expert workers to inexperienced ones, immediately and at scale. That is a genuine and valuable gain. It also poses the question this page exists to ask. If the tool carries the novice to expert-looking output on day one, what is happening to the process by which the novice was supposed to become an expert? The support study measured performance. It did not measure who those agents had become two years later, because nobody has yet had two years to measure.

Where the evidence is uncertain

Take the strongest study first, because it has the clearest limits. The Bastani experiment was mathematics practice, in schools, over a bounded period, with a single subject and a single age group. Professional judgement is not high-school algebra, and the fact that a tutoring interface protected learning in one setting does not tell you what a guardrailed interface would look like for a trainee solicitor or a graduate analyst. It tells you the design variable exists and that it matters, which is a great deal more than we had before, but not what to build.

The theory the interpretation leans on is itself contested. Deliberate practice has been re-examined, notably by Macnamara and Maitra in 2019, who revisited the original Berlin data and found that accumulated practice explained considerably less of the difference between performers than the 1993 paper is usually taken to claim. That matters here in a specific way. It means the answer is not simply more repetitions. It means the design and quality of practice carries more weight than the count, which strengthens the case for protecting the right reps rather than all of them, and weakens any confident arithmetic about how many are enough.

The MIT Media Lab result rests on 54 participants and remains a preprint, and its own authors and later commentators have flagged sample size and reproducibility. Treat anyone citing it as settled with the caution the authors themselves ask for. The Gerlich study shows correlation, not causation, and it carries a published correction, issued in September 2025, which anyone citing it should read alongside the original. And the productivity findings are not in dispute at all on their own terms. Output and speed rise. The capability question runs on a longer clock, over the years across which professional judgement actually forms, and that is precisely the horizon no workplace study has yet had time to reach.

The honest position is not that the evidence is conclusive. It is that the direction is consistent across learning science, laboratory work and now one good field experiment, that the mechanism is well understood, and that the damage is slow, invisible and expensive to reverse. Those are exactly the conditions under which you design before you have proof, not after.

The SuperSkills interpretation

Almost every organisation is measuring the wrong one of the two numbers. Performance-with-the-tool is visible, immediate and flattering: it appears in output, in cycle time, in the quality of the deck. Capability-without-the-tool is invisible, deferred and unflattering, and virtually nobody measures it, because measuring it means asking people to work without the thing you just bought them. The gap between those two numbers is where the risk lives, and the Bastani experiment is the first study to put a figure on it in a real setting. Both groups looked excellent while the model was open. One of them had been taught and one had been carried.

I call the repetitions that go missing in this process the missed reps. They are not the reps you notice skipping. They are the small, dull, unglamorous ones that nobody defends in a redesign, because individually none of them look load-bearing: the first draft you rewrite four times, the analysis you get wrong before you get it right, the client note you labour over for an hour that a model now produces in nine seconds. Judgement is assembled from thousands of those encounters, and it is assembled invisibly, which is why it can be removed invisibly. The accumulated organisational version is capability debt: the loss of human knowledge, skill and judgement that builds up when an organisation automates work faster than it redesigns how people learn through doing.

There is a second effect that the learning literature captures better than the AI literature does. Constraint is not an obstacle to formation. It is frequently the mechanism of it. Bjork and Bjork's desirable difficulties are, in a laboratory, the same thing that scarcity used to do in ordinary life: force attention, force depth, force the slow route. Remove the constraint and you do not get the same learning faster. You get a different, thinner thing that feels the same at the time. That is the argument I made in the Box of Amazing essay The Flake 99 Theory of Being Human, in a domestic register rather than an academic one: the constraint was never the enemy, it was the curriculum.

Which leads to the practical conclusion, and it is not the one either camp wants. The answer to AI in learning is not restriction, which is unenforceable and would forfeit a real gain. Nor is it unfettered access, which the evidence now says actively harms the learner. It is design, and the design variable is the same one GPT Tutor changed: does the tool make the person do the work, or does it do the work for them? That is the difference between drift and design, tested in a randomised trial rather than asserted from a stage.

Reps before delegation: a working rule

The rule I give individuals and teams is deliberately crude, because a rule people can remember at the moment of temptation beats a framework they read once. Do it before you delegate it. Before handing a task to a model, ask three questions.

The seniority asymmetry is the part worth holding on to. The novice needs to do the task many times before automating it; the expert who has done it many thousands of times can automate it safely, because they retain the judgement to check the output. Any specific numbers I use for that when teaching, a hundred and ten thousand, are illustrative heuristics rather than measured thresholds, and I would not want them quoted as findings. The shape of the rule is what the evidence supports: delegation is safe in proportion to the capability you already hold, and dangerous in proportion to the capability you were still building.

What this looks like in practice

A first-year analyst builds the model, gets the assumptions wrong, is corrected in a review, and rebuilds it. Repeat forty times and the analyst can smell a wrong number in someone else's model at a glance. That smell is the entire product being sold at the senior level, and it is not taught anywhere. It is a residue of the forty rebuilds. When a model builds the model, the deck ships faster and the residue never forms.

A trainee lawyer reads a hundred contracts before they can see, in ten seconds, that clause eleven is the problem. A junior doctor takes several hundred histories before pattern recognition begins to do the work that conscious reasoning did at the start. A copywriter writes a great many bad headlines to be able to tell, instantly, that this one is nearly right. In every case, the visible output of the expert is fast, confident and apparently effortless, and it is the compressed product of a large volume of slow, effortful and largely invisible work. AI reproduces the visible half perfectly and the invisible half not at all.

The organisational version, done well, looks unremarkable. A consultancy that keeps a rule that the first draft of a client hypothesis is written unaided, then improved with AI, and that the two versions are both kept, so a reviewer can see what the person thought before the machine spoke. A trading floor that runs the same trade unaided once a quarter. A hospital that examines judgement directly, through simulation and oral questioning, because it worked out a century ago that good outputs do not prove a capable practitioner.

What to do

If you are learning. Write your own answer before you ask. It can be a bad answer and it can take four minutes; the point is that you have committed to a position the model can then challenge, which is a different cognitive act from receiving one. Use AI to critique, extend and stress-test your work rather than to originate it. Notice fluency and distrust it, because the feeling of an easy session is the single most reliable signal that little was retained.

If you manage people who are learning. Decide, explicitly, which tasks are protected reps and say so, because in the absence of a decision the answer defaults to whatever is fastest that afternoon. Ask for the thinking, not the artefact: what did you prompt, what came back, what did you change and why. Review the second draft alongside the first. And accept the cost honestly. Protected practice is slower this quarter and it is the only thing that produces someone who can do the senior job in 2031.

If you design learning or education. The Bastani result is your brief. Guardrails are not a compromise between learning and access; in that experiment the guardrailed tool beat both unfettered access and no access at all. Build tools that ask rather than answer, that withhold until the learner commits, and that make the difficulty visible rather than removing it. And separate your two measurements, permanently: assess performance with the tool for the work, and assess capability without it for the person.

If you are accountable for capability at the top of an organisation. Put the second measurement on a dashboard. Nothing else in this page will survive contact with a quarter-end unless someone senior is answerable for a number that goes down when people are being carried.

Development of the idea

The argument on this page developed publicly over several years. In The Great Unbundling of Work (25 May 2025), I argued that AI was not removing jobs but unbundling them into tasks, and named the consequence for people whose expertise was being commoditised out from under them. In The Case for Being Bad at Things (18 January 2026), I set out the learning form of the argument directly, using Federer's record of winning roughly eighty percent of matches while winning barely half of all points, and wrote that the risk is not that we become bad at our jobs but that we become bad at becoming good. The Flake 99 Theory of Being Human (8 March 2026) took the same mechanism into ordinary life and the disappearance of constraint. The framework is developed in SuperSkills (Kogan Page, 2026).

Key research and primary sources

Where a claim matters, go to the study rather than the article reporting it. These are the primary sources behind this page.

Related SuperSkills research

The mechanisms in this page are developed in their own right elsewhere: the missed reps, the missing rungs, synthetic seniority, capability debt and drift versus design. On the wider question of what AI does to thinking, see AI and human judgement and AI and critical thinking. On using the tool without the dependency, see using AI without dependency, and on the pipeline consequences, will AI replace entry-level jobs.

About this research

Rahim Hirji is the author of SuperSkills: The Seven Human Skills for the Age of AI (Kogan Page, 2026) and the founder of The SuperSkills Intelligence Company. This work draws on research across more than 200 organisations in 30 countries over seven years. The page is written to a deliberate rule: findings are attributed to the studies that produced them and kept separate from the interpretation, which is the author's. Deliberate practice, desirable difficulties and cognitive offloading are established concepts from the research literature and are not his. The missed reps, the missing rungs, synthetic seniority, capability debt and drift versus design are part of the SuperSkills lexicon. This is a living reference, reviewed and updated as significant new evidence appears, rather than a dated article left to stand.

Cite this

Hirji, R. (2026). How humans learn with AI. The SuperSkills Intelligence Company. Last reviewed 26 August 2026. thesuperskills.com/research/how-humans-learn-with-ai

The work

Where the writing comes from.

These essays draw on research across more than 200 organisations in 30 countries. See the wider body of work, or bring it into your organisation.

All research →