← Research
Research

What happens to medical and legal training?

Two professions that codified apprenticeship, meeting a technology that removes the tasks the apprenticeship runs on.

Last reviewed: 19 September 2026

The three failures Ke and colleagues separate, the one measured clinical result and what it is not, the trial where the machine's gain never reached the doctor, the hallucination rates in the tools trainees are handed, and what the Divisional Court said about pupils.

Question this page answersAll 829 questions this research covers

Nobody has measured it, and the strongest statement available on the question is a denial. Sixteen authors writing in Nature Medicine in May 2026 put it in one sentence: direct causal evidence linking AI exposure during training to competency failure in medical trainees does not exist. What does exist is a set of measurements around the edges of that hole, and they are sharp enough to act on. Medicine and law are the right places to look, because both professions wrote down how many repetitions qualification takes.

The answer, in one line

Share as a card

Definition#

Supervised repetition: the qualifying mechanism in both professions. A trainee performs real work under someone who can see it, often enough that the judgement forms and can be assessed. Medicine and law are unusual in counting the repetitions, which makes them the two places where the effect of removing them would show first.

Share this definition as a card

Ke and colleagues split one word into three, and the third one is new#

The vocabulary for this question was thin until 22 May 2026, when sixteen authors across sixteen institutions published a Perspective in Nature Medicine separating three failures that the word deskilling had been carrying at once. Graded entry.

Deskilling is the degradation of a competence a clinician already had. Mis-skilling is the acquisition of incorrect reasoning patterns through uncritical adoption of erroneous or biased output. Never-skilling is the failure to form a foundational competence at all, when AI substitutes for the effort that would have built it. The authors predict a state they call false proficiency: competence that looks real and depends on the machine remaining available.

The third category is the one that matters for training, and the distinction is structural rather than a matter of severity. A deskilled clinician has a baseline to be recovered. A never-skilled one has nothing to return to, so every remedy built on retraining assumes a thing that was never there.

The paper reports no data of its own, and the estate grades it accordingly. The authors disclaim the strong reading more than once, and their sentence about the absence of causal evidence is quoted above because a great deal of secondary writing repeats their taxonomy as though it were a finding. Use it for the categories. Do not use it as evidence that the harm has occurred.

The one clinical measurement, and the people it is about#

Budzyn and colleagues examined 1,443 colonoscopies performed without AI assistance at four Polish centres, 795 before the centres adopted an AI detection system and 648 after, by 19 endoscopists averaging 27.6 years of experience. Adenoma detection in those unassisted procedures fell from 28.4 per cent to 22.4 per cent, a drop of 6.0 percentage points, with an adjusted odds ratio of 0.69. Graded entry.

This is the strongest direct evidence in the literature that capability degrades when a tool takes over the judgement, and it says nothing about trainees. The endoscopists had thirty years each. What it establishes for a training argument is the timescale: months of routine exposure were enough to move the unassisted performance of people whose skill was thoroughly formed. A profession with a view about how long it takes to build a capability should notice how little time it took to shift one.

The study is observational rather than randomised, covers one procedure in one country, and uses detection rate as a proxy for skill. Ehsan and colleagues give a second clinical setting, twelve months inside a five-site hospital group adopting an AI radiotherapy planning system, where dosimetrists reported by month nine that their unaided proficiency had worsened; that account is qualitative, measures no skill, and is widely misdescribed elsewhere as a study of radiologists using generative AI, which it is not. Graded entry.

Goh's trial: the machine's advantage did not reach the doctor#

In a single-blind randomised trial run over a month from November 2023, 50 US-licensed physicians, 26 attendings and 24 residents, worked through clinical vignettes with GPT-4 or with conventional resources, graded blind. Median diagnostic reasoning was 76 per cent with the model and 74 per cent without, an adjusted difference of 2 percentage points that did not reach significance. GPT-4 working alone scored a median 92 per cent, sixteen points clear of the physicians using conventional resources. Graded entry.

Set that beside the training question and it becomes uncomfortable. A tool substantially better than the doctors at the task produced almost no improvement in their work, which means the doctors were mostly not taking what it had. If the gain does not transfer to an attending physician with a validated rubric and a controlled hour, the assumption that a resident absorbs reasoning by watching a model do it has nothing behind it. Whatever the trainee is learning from the interaction, this trial gives no reason to think it is diagnosis.

In law the tools handed to juniors get the law wrong between a sixth and a third of the time#

Magesh, Surani, Dahl, Suzgun, Manning and Ho ran the first preregistered empirical evaluation of retrieval-augmented legal research products, over 200 handwritten queries registered with the Open Science Foundation in March 2024. Each of the three commercial tools hallucinated between 17 and 33 per cent of the time. Lexis+ AI answered 65 per cent of queries accurately and 18 per cent incompletely; Westlaw AI-Assisted Research was accurate 42 per cent of the time; Ask Practical Law AI was incomplete on 62 per cent of queries. Graded entry.

These are the paid, citation-grounded products, sold on the claim that grounding solves the problem. The free models a trainee reaches for at eleven at night are worse, and the Divisional Court of England and Wales said so in terms on 6 June 2025, hearing the Ayinde and Al-Haroun referrals together. At paragraph 6: tools "trained on a large language model such as ChatGPT are not capable of conducting reliable legal research". In Al-Haroun, a claim for £89.4 million, a schedule of forty-five citations put before the court contained eighteen cases that did not exist. Graded entry.

Paragraph 8 belongs on a page about training. Read it twice:

This duty rests on lawyers who use artificial intelligence to conduct research themselves or rely on the work of others who have done so. This is no different from the responsibility of a lawyer who relies on the work of a trainee solicitor or a pupil barrister for example, or on information obtained from an internet search.

The court reached for the pupil as the obvious analogy for a machine whose output must be checked. That is the profession's own account of what a junior is for, stated in the same breath as the thing now doing the junior's task. A supervisor who reads the model's work with the care once given to a pupil's has kept the discipline and removed the pupil. One further figure is worth holding against the panic: Charlotin's register of court decisions involving hallucinated material records 2,022 as at 6 September 2026, and the largest group by a distance is litigants in person at 1,163, against 805 for lawyers. Graded entry.

Two experiments that removed the tool and looked#

Neither is about doctors or lawyers, and together they are the closest thing to a test of never-skilling that anyone has run. Bastani and colleagues gave nearly a thousand high-school students either unrestricted GPT-4, a hints-only tutor, or nothing. Grades rose 48 per cent with unrestricted access and 127 per cent with the tutor while the tool was there. With it removed, the unrestricted group scored 17 per cent below students who had never had it, and the guardrailed tutor largely removed that harm. Graded entry.

Sankaranarayanan ran the same shape with 78 adults learning to program. Both AI groups beat the manual control on the work itself and were indistinguishable from each other. On a later maintenance task with the AI withdrawn, the unrestricted group failed 77 per cent of the time against 39 per cent for the scaffolded group. His phrase for them is fragile experts. Graded entry.

Two findings survive the move from a classroom to a hospital or a chambers. The damage stays invisible while the tool is present, because in both experiments the assisted groups performed at or above everyone else. And the scaffolding is what separates the outcomes: the same technology, configured to withhold the answer, produced a different learner. A training programme choosing between permitting AI and forbidding it is choosing between the two arms that both went wrong, and the arm that worked was neither.

What the assessment machinery assumes, and why it will not notice#

Postgraduate medicine assesses competence by watching. Moonen-van Loon and colleagues, analysing 12,779 workplace-based assessments from 953 residents, found that a reliability coefficient of 0.80 required eight mini-CEX observations, nine directly observed procedures or nine multi-source feedback rounds, falling to seven, eight and one when combined in a portfolio. Graded entry.

Every one of those instruments records what the trainee did, in a setting where the tool is present. Oral examination is the oldest instrument for getting behind the output, and a weak one: a 2025 comparison across three viva formats at a single institution reached 0.663 at best, which is moderate. Graded entry. So a trainee with false proficiency passes an assessment system that was designed, entirely reasonably, when the only way to produce the work was to be able to do it.

The supply of posts is a separate question and it is moving faster#

Hosseini Maasoum and Lichtinger report junior employment at AI-adopting US firms falling about 9 per cent relative to non-adopters six quarters after diffusion, with no comparable break in senior employment, concentrated in exposed occupations. Graded entry. The paper is unreviewed, and its subject is firms in general rather than either profession here.

Medicine and law are partly insulated, because the number of training posts is set by regulators and royal colleges rather than by whoever is doing the hiring. That insulation is also what makes them worth watching: if the number of qualifying places holds while the work inside those places thins out, the count of trainees will keep saying that nothing has happened.

Four things this page cannot establish#

Whether never-skilling occurs. The authors who named it say directly that nobody has shown it, and no study here follows a cohort of trainees through qualification.

Whether the Budzyn result generalises. One procedure, one country, an observational design, and a proxy outcome. It is the best evidence available and a single result.

Whether the classroom experiments transfer. Bastani and Sankaranarayanan measure school mathematics and novice programming, on tasks of hours. Clinical and legal judgement forms over years, through supervision and consequence that no experiment reproduces.

What the right configuration is. The scaffolding result says the middle arm worked in a maths tutor and a code editor. Nothing has tested a scaffolded AI inside a residency or a training contract, and the paragraph below is an argument rather than a finding.

What a training director can do while the evidence catches up#

Separate the three failures before designing anything, because the responses point in different directions. Deskilling in qualified staff is answered by scheduled unassisted practice. Never-skilling in trainees is answered by sequencing, keeping the repetitions unaided until the judgement is formed and only then adding the tool. Mis-skilling is answered by supervision of the reasoning rather than of the output, which is the one thing already built into both professions.

Measure unaided performance on a schedule. This is the single practice every study here supports, and both professions can absorb it, since they already observe trainees directly. A mini-CEX with the tool switched off is a capability measurement and needs no new instrument. See a capability audit.

Assume the tool is wrong often enough to matter, and give trainees the figure. A junior lawyer who knows the product hallucinates up to a third of the time reads its output differently from one who was told it was grounded in case law.

Do not count posts as a proxy for training. The number of places can hold while the work inside them empties, and the profession will discover that at the point those people become the supervisors.

Key sources

On the general version of this question, how juniors become senior and the missing rungs. On the profession-level view, AI and medicine and AI and law. On who is most exposed, which professions face the greatest deskilling risk. On whether trainees should use the tools at all, should juniors use AI. On measuring what a person can still do, assessing capability rather than output.

About this research#

Written by Rahim Hirji, author of SuperSkills (Kogan Page, 2026), keynote speaker on AI and human capability, and founder of The SuperSkills Intelligence Company. Nothing on this page is a SuperSkills coinage: deskilling, mis-skilling, never-skilling and false proficiency belong to Ke and colleagues, and fragile experts to Sankaranarayanan. Every figure was checked against the publisher record on 19 September 2026, and the two quoted paragraphs of the judgment were read in the approved judgment itself. The assembly is the contribution; the findings are other people's.

How this research works  ·  Reviewed quarterly  ·  Found an error? Tell me and it is corrected on the page.

Evidence review · SS-2026-270 · Graded against the published rubric

Cite this page

Hirji, R. (2026). What happens to medical and legal training?. The SuperSkills evidence base, SS-2026-270. https://thesuperskills.com/research/what-happens-to-medical-and-legal-training. Last reviewed 19 September 2026.

An evidence review by Rahim Hirji, not peer-reviewed research. For a material claim, cite the underlying study as well; every study here carries its own permanent link.

How citations and IDs work
Questions answered on this page

What happens to medical and legal training when AI does the junior work?

Nobody has measured it yet, and the profession that says so most clearly is medicine. Ke and colleagues, writing in Nature Medicine in May 2026, state that direct causal evidence linking AI exposure during training to competency failure in medical trainees does not exist. What is established is narrower and still serious: capability degrades in fully trained clinicians within months of routine AI exposure, learners who had unrestricted access perform worse than those who never had it once the tool is withdrawn, and the commercial legal research tools trainees are handed hallucinate between 17 and 33 per cent of the time.

What is never-skilling?

Never-skilling is the failure to form a foundational competence during training, because AI substituted for the cognitive effort that would have built it. Ke and colleagues separate it from deskilling, the degradation of a competence a clinician already had, and from mis-skilling, the acquisition of incorrect reasoning patterns from uncritical use of erroneous output. Never-skilling has no earlier baseline to recover, which is what makes it a different category and not a matter of degree. It is a risk model and not an observed phenomenon; the authors say so themselves.

Has deskilling been measured in doctors?

Once, convincingly, and in already-qualified clinicians rather than trainees. Budzyn and colleagues examined 1,443 colonoscopies performed without AI assistance by 19 endoscopists averaging 27.6 years of experience, at four Polish centres. Adenoma detection in unassisted procedures fell from 28.4 per cent before the centres adopted AI to 22.4 per cent after, a drop of 6.0 percentage points. It is observational rather than randomised, one procedure in one country, and detection rate is a proxy for skill.

Are AI legal research tools safe for trainees to use?

Not unsupervised. Magesh and colleagues ran the first preregistered evaluation of commercial legal research tools and found each of the three hallucinating between 17 and 33 per cent of the time: Westlaw AI-Assisted Research was accurate 42 per cent of the time, Lexis+ AI 65 per cent, and Ask Practical Law AI was incomplete on 62 per cent of queries. In June 2025 the Divisional Court held that freely available generative tools are not capable of conducting reliable legal research, and that the duty to check falls on the lawyer relying on the work, exactly as it would if a trainee solicitor or pupil barrister had produced it.

Are there still enough training posts?

That is the part moving fastest and the part with the weakest evidence. Hosseini Maasoum and Lichtinger report junior employment at AI-adopting US firms falling about 9 per cent relative to non-adopters six quarters after diffusion, with no comparable break in senior employment. The paper is unreviewed, and its subject is firms in general rather than either profession here. Formal qualifying routes in medicine and law are set by regulators rather than by hiring managers, so they move on a different clock.

In this hub

Professions and sectors

Where the pressure lands first, profession by profession.

Ask the evidence
What does the evidence actually show?What should our board be asking about this?Where does Rahim disagree with the consensus?
Bring this into your organisation

If this describes something happening in your teams, say so.

Keynotes, board sessions and advisory work, drawing on research across more than 200 organisations in 30 countries. Tell me the room, the date and the shift you need. A reply within 24 hours.

Start a conversation

Topics and audiences  ·  All research

Box of Amazing

Rahim’s free weekly letter on AI and human capability

If this was useful, the weekly letter is where the thinking happens first. Most of what ends up on this site starts there. Weekly essays on AI, capability and the future of work. Read by 25,000 people, every week since 2017. Free, and one click to stop.

Opens Substack to confirm. No pitch in it, unsubscribe in one click, and nobody follows up because you read something.

Running an event, or responsible for how AI arrives in your organisation? Keynotes  ·  Advisory for CEOs and boards  ·  Enquire