← Research
Research

How do juniors become senior if AI does the junior work?

Junior work was never a training scheme. It was a by-product of senior workload, and the by-product has to be chosen now or it goes quietly.

Last reviewed: 6 September 2026

The ten-minute effect, the interface changes that halve the damage, what thirty months of Chinese exam data show about outsourcing, and the four things that still build a senior.

Questions this page answersQuestions this page partly answersAll 811 questions this research covers

The same way they always did, by doing work that was hard enough to change them. What has changed is that nobody is now forced to give them that work. The junior tasks existed because senior people needed them done and could not do them all; the training was a by-product of somebody else's workload. AI removes the reason and leaves the by-product to be chosen deliberately or lost quietly. The evidence is unusually clear about which version of use builds a senior and which version does not, and the dividing line is not how much AI a junior uses.

The answer, in one line

The same way they always did, by doing work hard enough to change them. What has changed is that nobody is now forced to give them that work.

Share as a card

Definition#

Seniority: the ability to judge whether work is right without being told, built by doing the work and being wrong about it under supervision. AI can produce the work but not the being wrong.

Share this definition as a card

Ten minutes is enough to change what somebody does next#

The speed of the effect is the finding that should reorganise how a team is run. In randomised trials with 1,222 people across mathematical reasoning and reading comprehension, assistance was available during practice and then taken away. Performance improved while the tool was there, and afterwards those participants did worse unassisted and gave up sooner. The authors report the effect emerging after roughly ten minutes of interaction, and locate it in persistence: people conditioned to expect an immediate answer stop sitting with a problem, and sitting with problems is one of the strongest predictors of long-term learning.

These are short online tasks and a preprint, so the effect is a carry-over within a session and nothing about sustained professional practice. It still matters, because a ten-minute mechanism does not need a policy to take hold. It happens in the gap between a junior being handed a task and the first thing they do about it.

Access is not the harm. Outsourcing is#

Four studies converge here from different directions, which is the reason to trust the shape of the answer even where each one is limited.

Nearly a thousand high-school students were split three ways: unrestricted GPT-4, a hints-only tutor, or nothing. Grades rose 48 per cent with unrestricted access and 127 per cent with the tutor. Then the tool was taken away, and the unrestricted group scored 17 per cent lower than students who had never had it. The tutor group kept most of their gain. Same model, same students, different interface.

Developers learning an unfamiliar programming library scored 17 per cent lower on comprehension when they had an assistant, while finishing only marginally faster. The authors identify six patterns of interaction, and three of them preserve learning outcomes with the assistant still switched on.

In a study of 78 novice programmers, both AI groups beat the manual control on getting the code working, and did not differ from each other. Then the AI was cut off for a thirty-minute maintenance task. Unrestricted users failed at 77 per cent; the scaffolded group failed at 39. The author's phrase for the first group is fragile experts, and their fragility was invisible in everything measured up to that point.

And in the study with the longest run, thirty months of data on 26,811 Chinese secondary students, homework scores rose 18 per cent and homework time fell 30 per cent, while closed-book monthly exam scores fell 20 per cent within six months. Entrance-exam scores fell 18 and 24 per cent, with the full penalty appearing only after about two years. The load-bearing detail: the losses concentrated in the roughly 80 per cent of users whose homework time collapsed while scores rose, the pattern of outsourcing. Students who kept working at their normal pace were largely spared. That is self-selected adoption, one school system, secondary students and a working paper, and it says nothing directly about professional work.

The result that runs the other way, and it should#

A randomised experiment with 1,174 adults on a workplace-style problem found the opposite of a penalty. Without the assistant, the more educated group outperformed the less educated by 0.548 standard deviations. With it, the gap fell to 0.139. Once the assistant was removed, treated participants did not do worse than controls, and the less educated kept part of their gain, though a sizeable gap returned.

One session with an immediate unassisted module tests transfer within a sitting and not skill formation over months. What it establishes is enough to settle one argument: giving a junior the tool does not automatically leave them worse off. The studies that found post-removal deficits had taught a body of knowledge and removed the tool afterwards. Anyone running a blanket ban is treating access as the variable, and the variable is what the person does in the first ten minutes.

The cockpit worked this out first#

Sixteen airline pilots flew routine and non-routine scenarios in a 747-400 simulator with automation varied. Instrument scanning and manual control held up well, even where pilots reported little recent hand-flying. What degraded was cognitive: tracking position without a map, deciding the next navigational step, spotting an instrument failure.

The hands survive disuse better than the judgement does. Sixteen pilots in a simulator is a small base for a general rule, and the direction matches the wider decay literature, where a meta-analysis of 189 data points found cognitive and accuracy-based skills decaying faster than physical and speed-based ones. Applied to a junior, it points the wrong way from where most delegation decisions get made. The tasks a manager is most comfortable handing to AI, because they look mechanical, are the ones a junior loses least by losing. The ones that feel wasteful, working out what the problem actually is, are the ones the seniority was made of.

Four things that build a senior now#

Keep the reps that were hard and delegate the ones that were only long. The distinction is not seniority of task, it is whether the junior had to decide anything. Formatting a deck is length. Working out which three of eleven findings belong in it is a rep.

Design the interaction, not the policy. Bastani's tutor and Sankaranarayanan's scaffold both cut the damage roughly in half without removing the tool, and both worked by making the person produce something before the model did. A rule about when juniors may use AI is weaker than a habit about what they do in the first minute of using it.

Measure unaided, on a schedule, and write the number down. The Polish endoscopy result exists only because those departments still ran colonoscopies without the tool and could compare. An organisation with no unassisted baseline has no way of learning what it is losing, which is what a capability audit is for.

Make being wrong survivable. This one is an argument and not a measurement, so it is offered as such. If the definition above holds, seniority comes from being wrong under supervision, and a team where a junior's first draft is checked by a model before a person sees it has removed the supervision and kept the wrongness private. Nothing here measures that. It follows from the rest and it is testable by anyone willing to ask their juniors when they were last corrected by a human being.

Nobody has watched anyone do it yet#

The honest limit is large. Every result above is short, or young, or both. Liu measures minutes. Sankaranarayanan measures one session. Stromberg measures schoolchildren over thirty months, which is the longest run available and still not a career. No study has followed a cohort from a first job to a senior role with the tool present throughout, because there has not been time.

Sixteen authors in Nature Medicine named the risk that this describes, never-skilling, the failure to form competence at all when a tool substitutes for the effort that would have built it, and separated it from deskilling and mis-skilling. The terms are theirs. Their own caution is the part to carry: "Direct causal evidence linking AI exposure during training to competency failure in medical trainees does not exist." It is a risk model. Anyone quoting it as proof has misread it.

So the practical answer sits at the level of the individual team and not the level of the profession. The mechanisms are measured, the interfaces that blunt them are measured, and the long-run outcome is unknown. That is enough to act on and not enough to be confident about, and saying otherwise in either direction would be an overreach.

Key sources

On what has gone from the ladder, the missing rungs and the missed reps. On what the result looks like from outside, synthetic seniority. On the day-to-day version of this question, should juniors use AI at all and whether apprenticeships still work. On the counter-argument, whether this is deskilling or a rescaling of what counts as skill. On finding out before it matters, assessing capability rather than output.

About this research#

The definition of seniority above is the working definition used on this page and is not offered as a coined term. Never-skilling, mis-skilling and desirable difficulty belong to the researchers credited above. The reading offered here, that junior work was a by-product of senior workload and has to become a deliberate choice, and that the tasks safest to delegate are the long ones and not the hard ones, is an interpretation by Rahim Hirji, author of SuperSkills (Kogan Page, 2026), and is marked as an interpretation and not a finding.

Evidence review · SS-2026-188 · Graded against the published rubric

Cite this page

Hirji, R. (2026). How do juniors become senior if AI does the junior work?. The SuperSkills evidence base, SS-2026-188. https://thesuperskills.com/research/how-do-juniors-become-senior. Last reviewed 6 September 2026.

An evidence review by Rahim Hirji, not peer-reviewed research. For a material claim, cite the underlying study as well; every study here carries its own permanent link.

How citations and IDs work
Questions answered on this page

How do juniors become senior if AI does the junior work?

The same way they always did, by doing work hard enough to change them. What has changed is that nobody is now forced to give them that work. Junior tasks existed because senior people needed them done and could not do them all, so the training was a by-product of somebody else's workload. AI removes the reason and leaves the by-product to be chosen deliberately or lost quietly. The evidence is clear that the dividing line is not how much AI a junior uses. In one randomised study of nearly a thousand school students, an unrestricted group scored 17 per cent lower after the tool was removed than students who never had it, while a hints-only tutor group kept most of their gain.

Should juniors be banned from using AI so they learn properly?

The evidence does not support a ban and does support designing the interaction. A randomised experiment with 1,174 adults on a workplace-style task found that treated participants did not perform worse than controls once the assistant was removed, and that the assistant closed about three quarters of an education-based performance gap while present. Where harm appears, it appears through outsourcing. Across thirty months of data on 26,811 Chinese secondary students, exam losses concentrated in the roughly 80 per cent whose homework time collapsed while their homework scores rose, and students who kept working at their normal pace were largely spared.

Which junior tasks are safe to give to AI and which are not?

Delegate the tasks that were only long and keep the ones that were hard. The test is whether the junior had to decide anything: formatting a deck is length, working out which three of eleven findings belong in it is a rep. Aviation research points the same way. In a 747-400 simulator study, pilots' instrument scanning and manual control held up despite little recent practice, while the cognitive tasks degraded, and a meta-analysis of 189 data points found cognitive and accuracy-based skills decaying faster than physical ones. The tasks that feel most delegable because they look mechanical are the ones a junior loses least by losing.

How would you know if your juniors are not actually developing?

Measure unaided performance on a schedule and record the number. The strongest deskilling result available exists only because Polish endoscopy departments continued to run colonoscopies without AI and could compare: unassisted adenoma detection fell from 28.4 to 22.4 per cent in the same doctors. A study of 78 novice programmers makes the point at the other end of a career: both AI groups produced working code and looked identical on every measure taken, until the tool was cut off for a maintenance task and the unrestricted users failed at 77 per cent against 39 for a scaffolded group. Fragility is invisible until something is measured without the tool.

Is there evidence that people who learn with AI from day one end up less capable?

Not yet, and the gap is the main limit on everything else. No study has followed a cohort from a first job into a senior role with the tool present throughout, because there has not been time. The longest run available is thirty months of secondary school data. Sixteen authors in Nature Medicine named the risk as never-skilling, the failure to form competence at all when a tool substitutes for the effort that would have built it, and wrote in the paper itself that direct causal evidence linking AI exposure during training to competency failure in medical trainees does not exist. It is a risk model and should not be quoted as proof of harm.

In this hub

Work, careers and the labour market

What happens to jobs, careers and the first rung.

Ask the evidence
What does the evidence actually show?What should our board be asking about this?Where does Rahim disagree with the consensus?
Bring this into your organisation

If this describes something happening in your teams, say so.

Keynotes, board sessions and advisory work, drawing on research across more than 200 organisations in 30 countries. Tell me the room, the date and the shift you need. A reply within 24 hours.

Start a conversation

Topics and audiences  ·  All research

Box of Amazing

Rahim’s free weekly letter on AI and human capability

If this was useful, the weekly letter is where the thinking happens first. Most of what ends up on this site starts there. Weekly essays on AI, capability and the future of work. Read by 25,000 people, every week since 2017. Free, and one click to stop.

Opens Substack to confirm. No pitch in it, unsubscribe in one click, and nobody follows up because you read something.

Running an event, or responsible for how AI arrives in your organisation? Keynotes  ·  Advisory for CEOs and boards  ·  Enquire