Nobody has measured it, and the strongest statement available on the question is a denial. Sixteen authors writing in Nature Medicine in May 2026 put it in one sentence: direct causal evidence linking AI exposure during training to competency failure in medical trainees does not exist. What does exist is a set of measurements around the edges of that hole, and they are sharp enough to act on. Medicine and law are the right places to look, because both professions wrote down how many repetitions qualification takes.
The answer, in one line
Nobody has measured it yet, and the profession that says so most clearly is medicine. Ke and colleagues, writing in Nature Medicine in May 2026, state that direct causal evidence linking AI exposure during training to competency failure in medical trainees does not exist.
Definition#
Supervised repetition: the qualifying mechanism in both professions. A trainee performs real work under someone who can see it, often enough that the judgement forms and can be assessed. Medicine and law are unusual in counting the repetitions, which makes them the two places where the effect of removing them would show first.
Ke and colleagues split one word into three, and the third one is new#
The vocabulary for this question was thin until 22 May 2026, when sixteen authors across sixteen institutions published a Perspective in Nature Medicine separating three failures that the word deskilling had been carrying at once. Graded entry.
Deskilling is the degradation of a competence a clinician already had. Mis-skilling is the acquisition of incorrect reasoning patterns through uncritical adoption of erroneous or biased output. Never-skilling is the failure to form a foundational competence at all, when AI substitutes for the effort that would have built it. The authors predict a state they call false proficiency: competence that looks real and depends on the machine remaining available.
The third category is the one that matters for training, and the distinction is structural rather than a matter of severity. A deskilled clinician has a baseline to be recovered. A never-skilled one has nothing to return to, so every remedy built on retraining assumes a thing that was never there.
The paper reports no data of its own, and the estate grades it accordingly. The authors disclaim the strong reading more than once, and their sentence about the absence of causal evidence is quoted above because a great deal of secondary writing repeats their taxonomy as though it were a finding. Use it for the categories. Do not use it as evidence that the harm has occurred.
The one clinical measurement, and the people it is about#
Budzyn and colleagues examined 1,443 colonoscopies performed without AI assistance at four Polish centres, 795 before the centres adopted an AI detection system and 648 after, by 19 endoscopists averaging 27.6 years of experience. Adenoma detection in those unassisted procedures fell from 28.4 per cent to 22.4 per cent, a drop of 6.0 percentage points, with an adjusted odds ratio of 0.69. Graded entry.
This is the strongest direct evidence in the literature that capability degrades when a tool takes over the judgement, and it says nothing about trainees. The endoscopists had thirty years each. What it establishes for a training argument is the timescale: months of routine exposure were enough to move the unassisted performance of people whose skill was thoroughly formed. A profession with a view about how long it takes to build a capability should notice how little time it took to shift one.
The study is observational rather than randomised, covers one procedure in one country, and uses detection rate as a proxy for skill. Ehsan and colleagues give a second clinical setting, twelve months inside a five-site hospital group adopting an AI radiotherapy planning system, where dosimetrists reported by month nine that their unaided proficiency had worsened; that account is qualitative, measures no skill, and is widely misdescribed elsewhere as a study of radiologists using generative AI, which it is not. Graded entry.
Goh's trial: the machine's advantage did not reach the doctor#
In a single-blind randomised trial run over a month from November 2023, 50 US-licensed physicians, 26 attendings and 24 residents, worked through clinical vignettes with GPT-4 or with conventional resources, graded blind. Median diagnostic reasoning was 76 per cent with the model and 74 per cent without, an adjusted difference of 2 percentage points that did not reach significance. GPT-4 working alone scored a median 92 per cent, sixteen points clear of the physicians using conventional resources. Graded entry.
Set that beside the training question and it becomes uncomfortable. A tool substantially better than the doctors at the task produced almost no improvement in their work, which means the doctors were mostly not taking what it had. If the gain does not transfer to an attending physician with a validated rubric and a controlled hour, the assumption that a resident absorbs reasoning by watching a model do it has nothing behind it. Whatever the trainee is learning from the interaction, this trial gives no reason to think it is diagnosis.
In law the tools handed to juniors get the law wrong between a sixth and a third of the time#
Magesh, Surani, Dahl, Suzgun, Manning and Ho ran the first preregistered empirical evaluation of retrieval-augmented legal research products, over 200 handwritten queries registered with the Open Science Foundation in March 2024. Each of the three commercial tools hallucinated between 17 and 33 per cent of the time. Lexis+ AI answered 65 per cent of queries accurately and 18 per cent incompletely; Westlaw AI-Assisted Research was accurate 42 per cent of the time; Ask Practical Law AI was incomplete on 62 per cent of queries. Graded entry.
These are the paid, citation-grounded products, sold on the claim that grounding solves the problem. The free models a trainee reaches for at eleven at night are worse, and the Divisional Court of England and Wales said so in terms on 6 June 2025, hearing the Ayinde and Al-Haroun referrals together. At paragraph 6: tools "trained on a large language model such as ChatGPT are not capable of conducting reliable legal research". In Al-Haroun, a claim for £89.4 million, a schedule of forty-five citations put before the court contained eighteen cases that did not exist. Graded entry.
Paragraph 8 belongs on a page about training. Read it twice:
This duty rests on lawyers who use artificial intelligence to conduct research themselves or rely on the work of others who have done so. This is no different from the responsibility of a lawyer who relies on the work of a trainee solicitor or a pupil barrister for example, or on information obtained from an internet search.
The court reached for the pupil as the obvious analogy for a machine whose output must be checked. That is the profession's own account of what a junior is for, stated in the same breath as the thing now doing the junior's task. A supervisor who reads the model's work with the care once given to a pupil's has kept the discipline and removed the pupil. One further figure is worth holding against the panic: Charlotin's register of court decisions involving hallucinated material records 2,022 as at 6 September 2026, and the largest group by a distance is litigants in person at 1,163, against 805 for lawyers. Graded entry.
Two experiments that removed the tool and looked#
Neither is about doctors or lawyers, and together they are the closest thing to a test of never-skilling that anyone has run. Bastani and colleagues gave nearly a thousand high-school students either unrestricted GPT-4, a hints-only tutor, or nothing. Grades rose 48 per cent with unrestricted access and 127 per cent with the tutor while the tool was there. With it removed, the unrestricted group scored 17 per cent below students who had never had it, and the guardrailed tutor largely removed that harm. Graded entry.
Sankaranarayanan ran the same shape with 78 adults learning to program. Both AI groups beat the manual control on the work itself and were indistinguishable from each other. On a later maintenance task with the AI withdrawn, the unrestricted group failed 77 per cent of the time against 39 per cent for the scaffolded group. His phrase for them is fragile experts. Graded entry.
Two findings survive the move from a classroom to a hospital or a chambers. The damage stays invisible while the tool is present, because in both experiments the assisted groups performed at or above everyone else. And the scaffolding is what separates the outcomes: the same technology, configured to withhold the answer, produced a different learner. A training programme choosing between permitting AI and forbidding it is choosing between the two arms that both went wrong, and the arm that worked was neither.
What the assessment machinery assumes, and why it will not notice#
Postgraduate medicine assesses competence by watching. Moonen-van Loon and colleagues, analysing 12,779 workplace-based assessments from 953 residents, found that a reliability coefficient of 0.80 required eight mini-CEX observations, nine directly observed procedures or nine multi-source feedback rounds, falling to seven, eight and one when combined in a portfolio. Graded entry.
Every one of those instruments records what the trainee did, in a setting where the tool is present. Oral examination is the oldest instrument for getting behind the output, and a weak one: a 2025 comparison across three viva formats at a single institution reached 0.663 at best, which is moderate. Graded entry. So a trainee with false proficiency passes an assessment system that was designed, entirely reasonably, when the only way to produce the work was to be able to do it.
The supply of posts is a separate question and it is moving faster#
Hosseini Maasoum and Lichtinger report junior employment at AI-adopting US firms falling about 9 per cent relative to non-adopters six quarters after diffusion, with no comparable break in senior employment, concentrated in exposed occupations. Graded entry. The paper is unreviewed, and its subject is firms in general rather than either profession here.
Medicine and law are partly insulated, because the number of training posts is set by regulators and royal colleges rather than by whoever is doing the hiring. That insulation is also what makes them worth watching: if the number of qualifying places holds while the work inside those places thins out, the count of trainees will keep saying that nothing has happened.
Four things this page cannot establish#
Whether never-skilling occurs. The authors who named it say directly that nobody has shown it, and no study here follows a cohort of trainees through qualification.
Whether the Budzyn result generalises. One procedure, one country, an observational design, and a proxy outcome. It is the best evidence available and a single result.
Whether the classroom experiments transfer. Bastani and Sankaranarayanan measure school mathematics and novice programming, on tasks of hours. Clinical and legal judgement forms over years, through supervision and consequence that no experiment reproduces.
What the right configuration is. The scaffolding result says the middle arm worked in a maths tutor and a code editor. Nothing has tested a scaffolded AI inside a residency or a training contract, and the paragraph below is an argument rather than a finding.
What a training director can do while the evidence catches up#
Separate the three failures before designing anything, because the responses point in different directions. Deskilling in qualified staff is answered by scheduled unassisted practice. Never-skilling in trainees is answered by sequencing, keeping the repetitions unaided until the judgement is formed and only then adding the tool. Mis-skilling is answered by supervision of the reasoning rather than of the output, which is the one thing already built into both professions.
Measure unaided performance on a schedule. This is the single practice every study here supports, and both professions can absorb it, since they already observe trainees directly. A mini-CEX with the tool switched off is a capability measurement and needs no new instrument. See a capability audit.
Assume the tool is wrong often enough to matter, and give trainees the figure. A junior lawyer who knows the product hallucinates up to a third of the time reads its output differently from one who was told it was grounded in case law.
Do not count posts as a proxy for training. The number of places can hold while the work inside them empties, and the profession will discover that at the point those people become the supervisors.
Key sources
- Ke, Y., Jin, L., Ong, J. C. L. et al. (2026). AI-induced never-skilling in medical education. Nature Medicine, 32(6), 1997-2006. Graded entry.
- Budzyn, K., Romanczyk, M., Kitala, D. et al. (2025). Endoscopist deskilling risk after exposure to artificial intelligence in colonoscopy. The Lancet Gastroenterology & Hepatology, 10(10), 896-903. Graded entry.
- Goh, E., Gallo, R., Hom, J. et al. (2024). Large Language Model Influence on Diagnostic Reasoning: A Randomized Clinical Trial. JAMA Network Open, 7(10), e2440969. Graded entry.
- Magesh, V., Surani, F., Dahl, M., Suzgun, M., Manning, C. D. and Ho, D. E. (2025). Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools. Journal of Empirical Legal Studies, 22(2), 216-242. Graded entry.
- Divisional Court of England and Wales (2025). Ayinde v London Borough of Haringey and Al-Haroun v Qatar National Bank [2025] EWHC 1383 (Admin), 6 June 2025. Graded entry.
- Bastani, H., Bastani, O., Sungu, A., Ge, H., Kabakci, O. and Mariman, R. (2025). Generative AI Without Guardrails Can Harm Learning. PNAS, 122(26), e2422633122. Graded entry.
- Moonen-van Loon, J. M. W., Overeem, K., Donkers, H. H. L. M., van der Vleuten, C. P. M. and Driessen, E. W. (2013). Composite reliability of a workplace-based assessment toolbox for postgraduate medical education. Advances in Health Sciences Education, 18(5), 1087-1102. Graded entry.
Related SuperSkills research#
On the general version of this question, how juniors become senior and the missing rungs. On the profession-level view, AI and medicine and AI and law. On who is most exposed, which professions face the greatest deskilling risk. On whether trainees should use the tools at all, should juniors use AI. On measuring what a person can still do, assessing capability rather than output.
About this research#
Written by Rahim Hirji, author of SuperSkills (Kogan Page, 2026), keynote speaker on AI and human capability, and founder of The SuperSkills Intelligence Company. Nothing on this page is a SuperSkills coinage: deskilling, mis-skilling, never-skilling and false proficiency belong to Ke and colleagues, and fragile experts to Sankaranarayanan. Every figure was checked against the publisher record on 19 September 2026, and the two quoted paragraphs of the judgment were read in the approved judgment itself. The assembly is the contribution; the findings are other people's.
Evidence review · SS-2026-270 · Graded against the published rubric
Hirji, R. (2026). What happens to medical and legal training?. The SuperSkills evidence base, SS-2026-270. https://thesuperskills.com/research/what-happens-to-medical-and-legal-training. Last reviewed 19 September 2026.
An evidence review by Rahim Hirji, not peer-reviewed research. For a material claim, cite the underlying study as well; every study here carries its own permanent link.
How citations and IDs work