Yes. The useful question is not whether a junior should use AI but what they do in the first ten minutes of a task, and on that the evidence is unusually clear. Three experiments took the tool away afterwards and measured what was left. In every one, the harm depended on whether the person had done any of the thinking before the machine did.
That makes both of the common answers wrong. Telling a graduate not to use AI hands the work to someone who will. Telling them to use it freely produces the result below, which is the most uncomfortable number in this research.
Seventeen per cent below the people who never had it
Bastani and colleagues gave nearly a thousand high-school students one of three conditions: unrestricted GPT-4, a tutor built to give hints rather than answers, and no tool. While the tool was present, grades rose 48 per cent with unrestricted access and 127 per cent with the tutor. Then access was withdrawn and everyone sat the same assessment. The unrestricted group scored 17 per cent lower than students who had never had the tool at all. The guardrailed tutor largely removed that harm. Graded entry.
Read the two halves of that together, because either on its own misleads. Unrestricted access produced a real 48 per cent gain in performance and left students worse than if they had never used it. Both things are true, and only one of them appears in a report on how the pilot went.
Sankaranarayanan produced the same shape in adults doing professional work. Seventy-eight participants built something in one of three conditions: manual, unrestricted AI, and a scaffolded version designed to make the user do part of the thinking. Both AI groups beat the manual control on the work itself and were statistically indistinguishable from each other. Then the AI was removed and they had to maintain what they had built. The unrestricted group failed at 77 per cent, against 39 per cent for the scaffolded group. Graded entry.
Two groups that looked identical on delivery, separated by a factor of two the moment the tool was gone. It is one session on one task in novice programming with no follow-up, so it establishes that the gap can be produced rather than how long it lasts.
The gain is real, and it is largest for exactly the people at risk
Brynjolfsson, Li and Raymond studied 5,179 customer-support agents through a staged rollout. Productivity rose 14 per cent on average, 34 per cent for the newest and least experienced staff, and barely at all for the most skilled, because the tool transfers expert patterns to novices. Graded entry.
This is why the advice to abstain fails. A junior who refuses AI is choosing to be a third less effective than a colleague who does not, on measures their employer can see. The gain is genuine and it is largest for them.
It is also the precise moment the problem starts. The novice ships expert-looking work without the experience that expert-looking work used to require, and the study measures output over months rather than development over years, so what happens to those agents afterwards is the one thing it cannot tell us. That is the gap this page sits in: a large measured short-term gain, and an unmeasured long-term cost that the short-term gain actively conceals.
Why the first attempt is the part that matters
The learning science explains why the guardrailed versions kept the gain and dropped the harm, and it predates AI by decades.
Retrieval practice: pulling something out of your own memory strengthens it, and reading it again mostly strengthens the feeling of knowing. Roediger and Karpicke found restudying beat testing at five minutes, 81 per cent against 75, and testing beat restudying at one week, 56 against 42. Graded entry. Asking a model is not retrieval. The answer arrives from outside and the strengthening does not happen.
Productive struggle: attempting a problem before being taught produces better transfer than being taught first, even though the attempt usually fails. Sinha and Kapur's meta-analysis of 53 studies puts it at g = 0.36, rising to between 0.37 and 0.58 where the design follows the principles closely. Graded entry. A person who asks before trying never has the failed attempt, which is where the value sat.
Both mechanisms point at the same thirty seconds. The question is not how much AI a junior uses across a week. It is whether anything happened in their own head before the first prompt.
You will not be able to feel this happening
The reason this needs a rule rather than judgement in the moment is that the effect is invisible from the inside. Fluent material feels learned and a good output feels like evidence of a good performer, which is the illusion of competence. Fisher and colleagues showed that searching the internet inflates people's estimates of their own unaided knowledge, including on questions the search never touched. Graded entry.
In Sankaranarayanan's experiment the two AI groups could not be told apart while the tool was present. Nobody in the unrestricted group knew they were the fragile ones. Anyone asking a junior whether AI is harming their development is asking them to report on the one thing the effect prevents them from seeing.
Experience does not protect against this either. Nineteen endoscopists averaging 27.6 years of practice lost six percentage points of unassisted detection within months of routine exposure to an AI tool. Graded entry. If that happens to people with three decades behind them, a graduate has no margin at all.
Four rules that follow from the evidence
Attempt before you ask. Even badly, even for two minutes. This is the single instruction that separates the two arms of both experiments, and it costs almost nothing.
Use it to check and challenge, not to produce. Write the thing, then ask what is wrong with it. Ask it to argue the other side. The tool is at its most useful where you already have a position for it to attack, and at its most expensive where you have none.
Keep some work unaided, on purpose. Not out of principle, as a measurement. You cannot tell what you can still do from work you did with help.
Test yourself by removing it. Every study on this page found the gap only when the tool was taken away. That is the diagnostic, and it is available to anyone willing to be uncomfortable for an afternoon.
And four for whoever is managing them
The rules above put the burden on the least powerful person in the arrangement, which is the wrong place for it. A graduate told to work more slowly than their tools allow, in an organisation that measures throughput, will lose that argument every week.
So: say which tasks are for learning rather than for delivery, and protect their timelines accordingly. Ask juniors to explain and defend work rather than only to produce it, because that is the check the illusion of competence cannot survive. Prefer tools that make the user do part of the thinking where a choice exists, since that is the difference the two experiments actually measured. And measure capability directly rather than reading it off output, which is the argument in assessing capability rather than output.
The organisational version of this problem, where the junior tasks that built senior judgement are removed before anyone notices they were load-bearing, is the missing rungs. What accumulates when it goes unaddressed is capability debt.
What this page cannot tell you
Every experiment here is short. Bastani ran over a bounded period of school mathematics, Sankaranarayanan over a single session, Brynjolfsson over months. The claim juniors actually care about is about years, and nobody has measured that. Whether the gap closes with later practice, persists, or compounds is unknown, and can you regain a skill you have lost sets out how little is established.
The transfer is also an inference. School mathematics and a programming maintenance task are not law, medicine, consulting or design, and the guardrail that worked in a tutoring interface may not have an equivalent in professional software.
There is no study here on apprenticeships specifically. The estate holds no graded evidence on whether the apprenticeship model survives contact with this, which is a real gap on a question people keep asking, and it is named rather than filled.
Key sources
- Bastani, H. et al. (2025). Generative AI Without Guardrails Can Harm Learning. Proceedings of the National Academy of Sciences, 122(26). Graded entry.
- Sankaranarayanan, S. (2026). Mitigating "Epistemic Debt" in Generative AI-Scaffolded Novice Programming. Graded entry.
- Brynjolfsson, E., Li, D. and Raymond, L. (2023). Generative AI at Work. NBER Working Paper 31161. Graded entry.
- Roediger, H. L. III and Karpicke, J. D. (2006). Test-Enhanced Learning. Psychological Science, 17(3), 249-255. Graded entry.
- Sinha, T. and Kapur, M. (2021). When Problem Solving Followed by Instruction Works. Review of Educational Research, 91(5), 761-798. Graded entry.
- Fisher, M., Goddu, M. K. and Keil, F. C. (2015). Searching for Explanations. Graded entry.
- Budzyn, K. et al. (2025). Endoscopist deskilling risk after exposure to artificial intelligence in colonoscopy. The Lancet Gastroenterology and Hepatology, 10(10), 896-903. Graded entry.
- Institute for Fiscal Studies (2026). New Estimates of the Impact of Undergraduate Degrees on Lifetime Earnings. Graded entry.
Related SuperSkills research
The organisational side of this question is the missing rungs and will AI replace entry-level jobs. The mechanisms are retrieval practice, productive struggle, desirable difficulty and the illusion of competence. What accumulates is capability debt, and what it looks like from outside is synthetic seniority and the missed reps. On subject choice and returns, what should I tell my children to study. On the individual habit, using AI without dependency and am I becoming dependent on AI. On what schools are being told to do, the official guidance on AI in education.
About this research
Rahim Hirji is the author of SuperSkills (Kogan Page, 2026), keynote speaker on AI and human capability, and founder of The SuperSkills Intelligence Company. He has worked in and around education technology for a decade and wrote on early careers and internships in Minterns taking Minternships? (25 August 2019), three years before the missing-rungs argument took its current form. Findings are attributed to the studies that produced them and kept separate from the interpretation. This is a living reference, reviewed and updated as significant new evidence appears.
Cite this
Hirji, R. (2026). Should juniors use AI at all? The SuperSkills Intelligence Company. Last reviewed 30 August 2026. thesuperskills.com/research/should-juniors-use-ai
