Hand a beginner a fully worked example and they will learn more from it than from being made to solve the problem themselves. Hand the same example to someone who has already done twenty of them and they will learn less than if you had left them alone. The support has not changed. The person has. Educational psychologists named this pattern in 2003, after a decade of producing it by accident. It remains the closest thing any literature holds to an answer for the question everybody now asks about AI: who should be using the help, and when.
The answer, in one line
The expertise reversal effect is the finding that instructional support which helps a beginner learn can hinder someone who already knows the material.
Definition#
The expertise reversal effect: the finding that instructional support which helps a beginner learn can hinder someone who already knows the material. Named by Slava Kalyuga, Paul Ayres, Paul Chandler and John Sweller in 2003, it holds that guidance becomes redundant once a learner's own knowledge supplies the same thing, and that redundant guidance still has to be read and reconciled, which costs the working memory that learning needs. The term is theirs and is not claimed by this research.
Named in 2003, after a decade of results nobody could explain#
Cognitive load theory had spent the 1990s producing instructional designs that worked. Put the text inside the diagram instead of beside it and people learn faster, because they stop searching back and forth between the two. Speak the explanation rather than printing it and people learn faster still, because the ear and the eye do not compete for the same channel. Show a solved problem instead of setting an unsolved one and people learn faster again, because they are not burning attention on a search strategy that teaches them nothing.
Then the same laboratories started finding that these designs stopped working, and eventually turned harmful, on learners who had done the topic before. Kalyuga and three colleagues at the University of New South Wales gathered eleven of those studies into one paper for Educational Psychologist and gave the pattern a name. Their summary of it is unusually flat for a journal abstract: techniques that are highly effective with inexperienced learners can lose their effectiveness and even have negative consequences when used with more experienced learners
.
The mechanism they propose is redundancy, and the detail that carries it is a claim about attention rather than about preference. A beginner has no stored pattern for the task, so external guidance stands in for the pattern they do not yet have. An expert has the pattern. Offer them the guidance anyway and they cannot simply disregard it: in the authors' words, redundant information is frequently difficult to ignore
. They read it, they hold it alongside what they already know, and they do the work of checking that the two agree. That reconciliation consumes the same limited working memory that building new knowledge would have used. The guidance has not become useless. It has become a second thing to process.
The electricians who did better once the labels came off#
Two of the studies in the review do most of the argumentative work, and both were run on people learning a trade rather than on undergraduates.
In the first, apprentice electricians were given circuit diagrams either with the explanatory text integrated into the diagram or with the diagram alone. Inexperienced trainees needed the text and could not make sense of the bare diagram at all. More experienced trainees performed significantly better with the diagram on its own, and rated the bare version as taking less mental effort. The same page of material produced opposite results in the same trade at two points in the same training programme.
In the second, mechanical trade apprentices were given either worked examples to study or the equivalent problems to solve. The authors' own abstract states the outcome without hedging: inexperienced trainees benefited most from the worked examples, and with more experience in the domain the worked examples became redundant and problem solving proved superior. That paper is titled When Problem Solving Is Superior to Studying Worked Examples, which tells you how surprising it was to the field at the time.
Two things are worth holding on to about those results. The reversal was measured on what people could do afterwards, on a later test, without the instruction in front of them. And the learners reported the harmful condition as easier. Experienced trainees given redundant text rated their mental effort higher than experienced trainees given the bare diagram, so in that direction the felt experience and the outcome agreed. The estate carries the case where they come apart at desirable difficulty.
Fade the guidance, and do not simply stop it#
If support has to come off, the review points at when and how. Alexander Renkl and colleagues, whose studies Kalyuga's paper draws on, found that detailed worked examples suited novices and should be gradually faded as knowledge grows, replaced step by step with problems. Renkl, Robert Atkinson, Uwe Maier and Richard Staley then tested the removal itself and found that a fading procedure beat an abrupt switch from examples to problems. The review calls this the guidance-fading effect and treats it as the direct instructional application of the reversal.
That is a more useful finding for an organisation than the reversal on its own, because it names a shape. Support is not a switch with an on position for juniors and an off position for seniors. It is a taper, and the taper has to be run deliberately by somebody who knows where the learner currently is. Nothing about a general-purpose chat interface tapers. It offers the same completeness to a graduate in week two and to a partner of twenty years, and the completeness is the thing the 2003 literature says should have been coming off.
The middle is the part a chat window removes#
There is a distinction inside this research that the popular reading of it tends to lose, and it decides whether the effect transfers to AI at all.
A worked example demonstrates the route to an answer: the problem statement followed by every solution step, laid out so the learner can see where each move came from. Cognitive load theory's explanation for why it works on beginners depends entirely on that structure. The example removes the aimless search for a strategy and directs attention to the steps, which is where the pattern gets built. What it hands over is the reasoning.
A generative answer removes the search and the steps together. Rahim Hirji put the point in Show Your Working on 12 July 2026, through the red pen in the margin of a British school exercise book. The answer was never the thing being marked. The answer was only proof that the working had happened.
His reading of the machine follows from that: it hands you the answer with the middle removed. No steps. No crossings-out. No line where it nearly went wrong.
So a worked example and a chat response sit on opposite sides of the variable Kalyuga and colleagues were manipulating. Both reduce effort. One of them reduces effort by showing the learner the structure they are meant to acquire; the other reduces effort by supplying the finished object and keeping the structure. On the 2003 account, only the first should build anything in a beginner. That is an inference from this literature and not a result in it. The inference belongs to this research; the authors make no such claim.
The trap in reading this as permission for beginners#
Read quickly, the expertise reversal effect says that support helps novices and hurts experts, so the sensible policy is to give AI to the junior people and withhold it from the senior ones. The measured evidence on AI runs both with and against that, depending on what is being measured, and the split is the most useful thing on this page.
On immediate output the reversal pattern appears cleanly. In Fabrizio Dell'Acqua and colleagues' field experiment at Boston Consulting Group, consultants below the average gained most from the model and the gap between them and the strongest performers narrowed. Erik Brynjolfsson, Danielle Li and Lindsey Raymond found the same shape in a customer support centre, where novice and low-skilled agents gained most. Shakked Noy and Whitney Zhang found it in writing tasks. That is four separate settings in which the help was worth more to the person who knew less.
On what people can then do unaided, the direction inverts. Hamsa Bastani and colleagues gave nearly a thousand high-school students unrestricted access to GPT-4 and grades rose by 48 per cent while the tool was there. With the tool removed, that group scored 17 per cent lower than students who had never had it. A second arm of the same experiment, given a tutor that offered hints instead of answers, gained more while using it and largely escaped the loss afterwards. Budzyn and colleagues found the equivalent in a hospital: endoscopists' unassisted detection rate fell after a period of working with the AI system.
Those two bodies of work are not in conflict, because they measure different things. The productivity studies measure output with the tool in hand. The expertise reversal literature measures capability with the tool taken away, on a later test, which is also what Bastani and Budzyn measure. Once the dependent variables are separated, the 2003 finding and the 2025 findings say the same thing from two directions: the effect of support on performance and its effect on learning can point in opposite directions at once, and a beginner is where they are furthest apart.
What a 2003 result cannot settle about a 2026 tool#
The review is a review. It generated no data, it selected its own eleven studies, roughly half of them the authors' own, and it pooled no effect sizes, so nobody can say from it what was left out or how large the reversal is on average. The underlying experiments ran on trade apprentices and school pupils learning bounded tasks with a correct answer the experimenter already held, which is a long way from a lawyer deciding whether a clause is acceptable.
Nothing in this literature tested a generative system, because none existed in this form. The two treatments it compares are a diagram with labels and a diagram without them, and neither is a fluent collaborator that will answer whatever it is asked. No study known to this research has yet run the direct test: give matched groups the same task, vary the level of AI scaffolding against measured prior expertise in that domain, and test unaided performance afterwards. Bastani's two arms are the closest thing to it and they varied the interface rather than the learner.
One more limit belongs on the record, because it cuts against the argument. Kalyuga and colleagues report one case, Pollock, Chandler and Sweller on isolated elements, where the advantage disappeared with expertise without ever reversing. An effect that sometimes flattens rather than flipping is a weaker claim than the name suggests, and the paper says so.
Three questions this puts to whoever decides who gets the tool#
The practical content of the 2003 conclusion is a warning about defaults. Without tailoring to the learner's experience, the authors write, the effectiveness of instructional designs is likely to be random
. A single AI policy applied across a firm is exactly an untailored instructional design, and the reversal predicts that it will help some people, do nothing for others and cost a third group, with no way to tell which from the outside.
Three questions follow, and all three are answerable inside an organisation without new research. First, for this task, does the person already hold the pattern, and how would anybody know? Expertise in the reversal literature is domain-specific and task-specific, so a fifteen-year veteran can be a novice on the thing in front of them this morning. Second, does the support show the working or supply the result? That is the difference between a worked example and an answer. Bastani's two arms turned that one variable into a 17-point swing. Third, is the support tapering? If the same interface serves the first week and the tenth year, nobody is running the fade that the evidence says is the part that works.
The question of what a junior in particular should do with this is taken up at should juniors use AI at all, and the organisational version at how juniors become senior.
Key sources
- Kalyuga, S., Ayres, P., Chandler, P. and Sweller, J. (2003). The Expertise Reversal Effect. Educational Psychologist, 38(1), 23-31. DOI 10.1207/S15326985EP3801_4. The DOI resolves to Taylor and Francis, which serves an abstract and returns an empty body to both of this site's fetchers; the full text was read at the University of Wollongong Research Online copy, whose own cover sheet gives the authors in a different order from the article's byline. Graded entry.
- Kalyuga, S., Chandler, P., Tuovinen, J. and Sweller, J. (2001). When Problem Solving Is Superior to Studying Worked Examples. Journal of Educational Psychology, 93(3), 579-588. Author abstract read at source; full text gated, so no figure from it appears here. Graded entry.
- Bastani, H., Bastani, O., Sungu, A., Ge, H., Kabakci, O. and Mariman, R. (2025). Generative AI Without Guardrails Can Harm Learning: Evidence from High School Mathematics. PNAS, 122(26), e2422633122. Graded entry.
- Dell'Acqua, F. et al. (2023). Navigating the Jagged Technological Frontier. Harvard Business School and BCG working paper. Graded entry.
- Brynjolfsson, E., Li, D. and Raymond, L. (2025). Generative AI at Work. Quarterly Journal of Economics, 140(2), 889-942. Graded entry.
- Noy, S. and Zhang, W. (2023). Experimental evidence on the productivity effects of generative artificial intelligence. Science, 381(6654), 187-192. Graded entry.
- Budzyn, K., Roman'czyk, M., Kitala, D. et al. (2025). Endoscopist deskilling risk after exposure to artificial intelligence in colonoscopy. The Lancet Gastroenterology and Hepatology, 10(10), 896-903. Graded entry.
- Hirji, R. (2026). Show Your Working. Box of Amazing, 12 July 2026. Source of the red-pen argument and of the line about the answer with the middle removed.
Related SuperSkills research#
On what happens when the work is handed over rather than shown, cognitive offloading and does using AI stop you learning. On the gap between feeling fluent and being able, desirable difficulty, productive struggle and the illusion of competence. On the version of this question a person actually asks, should juniors use AI at all and how do juniors become senior. On the task-by-task unevenness of where the help lands, the jagged frontier. On the practice a person skips while the tool is doing it, the missed reps. On the design question for a whole organisation, how humans learn with AI.
Explainer · SS-2026-267 · Graded against the published rubric
Hirji, R. (2026). What is the expertise reversal effect?. The SuperSkills evidence base, SS-2026-267. https://thesuperskills.com/research/what-is-the-expertise-reversal-effect. Last reviewed 18 September 2026.
An evidence review by Rahim Hirji, not peer-reviewed research. For a material claim, cite the underlying study as well; every study here carries its own permanent link.
How citations and IDs work