Unevenly, and in a different order from the one the debate assumes. The strongest randomised evidence sits in image-based screening, where AI-supported mammography has now been tested in a trial of more than 100,000 women and reduced interval cancers. The weakest sits in bedside prediction, where the most widely installed proprietary sepsis model performed close to useless when somebody finally validated it outside the vendor. And the finding that matters most to the profession is neither of those. It is a measured fall in experienced endoscopists' own unassisted detection rate after a few months of working alongside a machine.
Medicine is the profession with the most clinical AI evidence and the least agreement about what it means. That is partly because three very different activities get discussed as one thing: reading images, predicting deterioration, and writing notes. They have three different evidence bases and three different answers.
Screening is where the randomised evidence actually is
The Mammography Screening with Artificial Intelligence trial in Sweden is the only randomised controlled trial of AI inside a national cancer screening programme, and it now has full results.
The interim safety analysis, published in The Lancet Oncology in August 2023, covered 80,033 women aged 40 to 80 at four sites in southwest Sweden, randomised one to one between AI-supported reading and standard double reading by two radiologists. Cancer detection was six per 1,000 screened women with AI support against five per 1,000 without, 41 more cancers found. The false-positive rate was 1.5 per cent in both arms. Radiologists performed 46,345 screen readings in the AI arm against 83,231 in the control arm, a 44 per cent reduction in screen-reading workload.
The full results, published in The Lancet in January 2026 with more than 100,000 women and two years of follow-up, tested the question that matters. Interval cancers, the ones diagnosed between screening rounds and generally the more aggressive, fell from 1.76 per 1,000 women in the control arm to 1.55 per 1,000 in the AI arm, a 12 per cent reduction. There were 16 per cent fewer invasive cancers in the interval, 21 per cent fewer large ones and 27 per cent fewer of the aggressive subtypes. Cancers detected at screening rather than between rounds rose from 74 to 81 per cent of all cases.
Read the design rather than the headline. The AI triaged low-risk examinations to a single radiologist and high-risk ones to two, and marked suspicious findings for the human reader. The radiologist kept the recall decision throughout. The first author's own summary is that the study does not support replacing clinicians, because at least one radiologist still reads every case. The trial ran on one mammography device with one AI system in one country, with moderately to highly experienced readers, and the authors say so.
What was tested here was a redesigned workflow with a human decision-maker inside it, evaluated on patient outcomes over years. Almost nothing else in clinical AI has been held to that standard.
The sepsis model that half of American medicine was already running
In 2021, Wong and colleagues at Michigan Medicine published the first serious external validation of the Epic Sepsis Model, a proprietary early-warning tool then deployed across hundreds of US hospitals. They studied 27,697 patients across 38,455 hospitalisations, of which 7 per cent involved sepsis.
The model achieved an area under the curve of 0.63, against the 0.76 to 0.83 its developer had cited. At the alert threshold the hospital was actually using, sensitivity was 33 per cent, specificity 83 per cent and positive predictive value 12 per cent. It failed to identify 1,709 of the 2,552 septic hospitalisations, 67 per cent, and 60 per cent of those patients received timely antibiotics anyway. Meanwhile it crossed the alert threshold in 18 per cent of all hospitalisations, meaning clinicians would evaluate eight patients to find one who eventually became septic.
The authors' own conclusion is the one worth keeping: widespread adoption of a model performing this poorly raises concerns about sepsis management nationally. A tool had been sold, procured, installed and wired into clinical alerting at scale before anyone outside the vendor measured it against outcomes.
Screening and prediction are not the same problem. Screening AI reads a stable, standardised image against a well-defined target. Prediction models forecast a physiological trajectory from data whose meaning shifts between hospitals, coding practices and years. The first has produced randomised evidence. The second has produced a cautionary tale about buying capability on the strength of internal figures.
Forty-one seconds a note, in one arm out of two
Ambient documentation is the fastest-spreading clinical AI in the world and the one with the loosest claims attached to it. A three-arm pragmatic randomised trial at UCLA, run from November 2024 to January 2025, gave 238 outpatient physicians across 14 specialties either Microsoft DAX, Nabla, or usual care.
Time spent writing a note fell by 41 seconds in the Nabla arm and 18 seconds in the control arm, a relative difference of 9.5 per cent that reached significance. The DAX arm fell by 23 seconds and did not differ significantly from doing nothing. Physicians using either scribe reported better scores on the Mini-Z burnout instrument, lower task load and lower work exhaustion, with no significant difference between the two products on any of those measures.
Three details in the paper deserve more attention than the headline. Roughly 15 per cent of physicians given a scribe never used it once. Usage rate correlated with time saved, so the average conceals a large gap between the doctors who adopted the tool and those who did not. And the authors flag a measurement problem affecting every study of this kind: the electronic record's own time metrics do not count editing done inside the scribe application, so reported savings may be overstated.
A minute a note is a real gain across a career. It is also roughly a hundredth of what the category is usually sold as delivering, and the wellbeing effect is larger and better evidenced than the time effect. Those are different value propositions and a service should know which one it is buying.
Adding a model to a doctor added nothing
The most uncomfortable trial in clinical AI is small and clean. Goh and colleagues randomised 50 US-licensed physicians, 26 attendings and 24 residents, to work through clinical vignettes with either GPT-4 or conventional resources such as UpToDate and Google, graded blind against a validated diagnostic reasoning rubric.
Median score was 76 per cent with the model and 74 per cent without: an adjusted difference of 2 percentage points, confidence interval minus 4 to plus 8. Median time per case was 519 seconds against 565, a difference that did not reach significance either. Every physician in the model arm used it.
Then the finding that made the study famous. The model working alone scored 92 per cent, sixteen percentage points above the physicians using conventional resources. The capability was present in the room and did not transfer to the people.
The authors' explanation is about prompting and interaction design rather than about doctors, and they are careful to say the result does not license autonomous diagnosis. Six curated vignettes are not clinical practice, and the design deliberately strips out interviewing, examination and context, which is most of what a clinician does. Take the study for what it establishes: the performance of a human-plus-model pairing is not the sum of the two, and cannot be assumed from either. That is the same shape as the jagged frontier result in consulting and the meta-analytic finding that human-AI combinations frequently underperform the better of their two parts.
The average radiologist does not exist
Yu and colleagues gave 140 radiologists AI assistance across 15 chest X-ray diagnostic tasks, roughly 5,190 observations, with empirical-Bayes shrinkage to separate genuine individual differences from noise. The effect of the assistance ranged from strongly positive to strongly negative between individuals. Experience, subspecialty and prior familiarity with AI all failed to predict who would benefit, and lower performers did not reliably gain the most.
A department that issues the same tool to everyone is therefore helping some clinicians and harming others, with no current means of telling which in advance. That is a governance problem rather than a technology problem, and it points at monitoring individual performance after deployment rather than at better procurement.
Poland, 1,443 colonoscopies, and the first real number on deskilling
Budzyn and colleagues published the finding this research regards as the most important single result in clinical AI. Nested in the ACCEPT trial across four Polish endoscopy centres, they examined 1,443 colonoscopies performed without AI assistance: 795 before the centres introduced AI and 648 afterwards, by 19 endoscopists averaging 28 years of experience.
Adenoma detection in those unassisted procedures fell from 28.4 per cent before AI exposure to 22.4 per cent after, a drop of 6.0 percentage points, p=0.0089, adjusted odds ratio 0.69.
Note what is being measured. Not performance with the tool, which improved. Performance without it, in highly experienced professionals, within months. The study is observational rather than randomised, other changes over the period cannot be fully excluded, and detection rate is a proxy for skill rather than skill itself. It establishes that the effect is real and reachable in ordinary practice, not how common it is.
This is capability debt with a number attached: the cost is not visible while the tool is present and appears only when it is absent. It is also the clearest available instance of what happens when a professional stops taking the repetitions that built the judgement, which this research has called the problem of missed reps.
Medicine already owns the answer, from another safety-critical trade
Aviation faced this in the 1980s and answered it structurally rather than culturally. Line pilots undergo recurrent proficiency checks that can be failed, with hand-flying assessed periodically without the automation. The finding from that literature that transfers most directly is that the manual skills degrade more slowly than the cognitive ones: what erodes first is knowing what the system is doing and why, not the hands.
Applied to a clinical service, that suggests three things follow from Budzyn rather than from principle. Measure unassisted performance periodically, on a schedule, as a condition of continuing to use the tool. Treat a rising gap between assisted and unassisted performance as a signal about the service rather than about the individual. And record how long a clinician actually spends per AI-flagged item, since time per decision is the most diagnostic and least collected number in any oversight arrangement.
Medicine has the apparatus for this already. It runs revalidation, appraisal, audit and mortality review. What it does not yet have is a standing unassisted baseline for tasks where AI has been introduced. The full argument is set out in what professions can learn from aviation.
What has not been shown, and should be said out loud
A page like this is worth more for what it declines to assert.
- No clinical AI deployment has been shown to reduce deskilling. Budzyn measured the loss. Nothing yet measures a countermeasure. Recurrent unassisted assessment is an inference from aviation, not a tested clinical intervention.
- The MASAI result belongs to one device, one AI system, one country and experienced readers. Its authors say so, and generalisation to other programmes is unevidenced rather than merely uncertain.
- Nobody knows which clinicians AI helps. Yu and colleagues looked for predictors and found none that held.
- Large language model performance in medicine is measured almost entirely on vignettes and examinations. Those strip out history-taking, examination, uncertainty over time and the patient in front of you, which is where most diagnostic error is generated.
- Documentation time savings rest on a metric its own investigators call flawed. Any figure quoted from electronic-record telemetry, including the one on this page, carries that caveat.
- The number of AI-enabled devices authorised for marketing is not stated here. The regulator publishes a list and says on the page that it is not comprehensive. Counts in circulation come from third parties reading that list, and this research does not repeat them.
Six questions before a clinical service deploys anything
Drawn from the failures above rather than from a framework.
- Has this been validated outside the vendor, on our kind of patients? The Epic sepsis case is the entire argument for asking, and the answer is often no.
- What is our unassisted baseline today, and when will we measure it again? If nobody records it before deployment, the comparison can never be made afterwards.
- What is the human decision that remains, and how long does it get? MASAI kept the recall decision with a radiologist. A workflow with no protected human decision has not been tested by any of this evidence.
- Who is harmed by the average? Yu and colleagues found the effect varies by individual with no usable predictor, so post-deployment monitoring has to be individual.
- What is the alert burden, and who absorbs it? Eighteen per cent of hospitalisations, in the sepsis case. Alert fatigue is a clinical safety issue in its own right.
- Are we buying time or wellbeing? In the scribe trial the wellbeing effect was larger and more consistent than the time effect. Procured as a productivity tool, it will be judged against the wrong number.
Key research and primary sources
- Budzyn, K., Roman'czyk, M., Kitala, D. et al. (2025). Endoscopist deskilling risk after exposure to artificial intelligence in colonoscopy. The Lancet Gastroenterology and Hepatology, 10(10).
- Lang, K. et al. (2023). AI-supported screen reading versus standard double reading in the MASAI trial: a clinical safety analysis. The Lancet Oncology, 24(8).
- Gommers, J., Lang, K. et al. (2026). Interval cancer, sensitivity and specificity comparing AI-supported mammography screening with standard double reading in the MASAI study. The Lancet, published 29 January 2026.
- Wong, A., Otles, E., Donnelly, J. P. et al. (2021). External Validation of a Widely Implemented Proprietary Sepsis Prediction Model in Hospitalized Patients. JAMA Internal Medicine, 181(8).
- Goh, E., Gallo, R., Hom, J. et al. (2024). Large Language Model Influence on Diagnostic Reasoning: A Randomized Clinical Trial. JAMA Network Open, 7(10).
- Lukac, P. J., Turner, W., Vangala, S. et al. (2025). Ambient AI Scribes in Clinical Practice: A Randomized Trial. NEJM AI, 2(12).
- Yu, F., Moehring, A., Banerjee, O. et al. (2024). Heterogeneity and predictors of the effects of AI assistance on radiologists. Nature Medicine, 30(3).
- United States Federal Aviation Regulations. 14 CFR 121.441, Proficiency checks.
Every figure on this page was checked against the primary source. Where a number in wide circulation could not be traced to a stated methodology, it was removed rather than repeated.
Related SuperSkills research
On the mechanism, deskilling, capability debt and missed reps. On which professions carry the most exposure, the deskilling risk map. On the structural answer, what professions can learn from aviation and whether a lost skill comes back. On oversight, human in the loop is not a safeguard, meaningful human oversight and who supervises work they cannot do. On the adjacent consumer question, AI for therapy or advice.
About this research
Rahim Hirji is the author of SuperSkills (Kogan Page, 2026) and founder of The SuperSkills Intelligence Company. Every trial cited here was read at the primary source and every figure checked against it; the MASAI outcome figures are taken from the trial publications and the journal's own press summaries, which is the weaker basis of the two and is flagged for that reason. The aviation parallel is an inference this research draws, not a finding from any clinical study. Not medical advice. Reviewed quarterly.
Cite this
Hirji, R. (2026). How will AI change medicine? The SuperSkills Intelligence Company. Last reviewed 28 August 2026. thesuperskills.com/research/how-will-ai-change-medicine