← Research
Research

How will AI change medicine?

Screening has the randomised evidence. Bedside prediction mostly does not. And the best-measured effect so far is on the doctors rather than on the patients.

Last reviewed: 28 August 2026

What the MASAI trial established, why the most widely installed sepsis model failed external validation, the first measured deskilling of experienced clinicians, and the questions a clinical service should be able to answer before it deploys anything.

Unevenly, and in a different order from the one the debate assumes. The strongest randomised evidence sits in image-based screening, where AI-supported mammography has now been tested in a trial of more than 100,000 women and reduced interval cancers. The weakest sits in bedside prediction, where the most widely installed proprietary sepsis model performed close to useless when somebody finally validated it outside the vendor. And the finding that matters most to the profession is neither of those. It is a measured fall in experienced endoscopists' own unassisted detection rate after a few months of working alongside a machine.

Medicine is the profession with the most clinical AI evidence and the least agreement about what it means. That is partly because three very different activities get discussed as one thing: reading images, predicting deterioration, and writing notes. They have three different evidence bases and three different answers.

Screening is where the randomised evidence actually is

The Mammography Screening with Artificial Intelligence trial in Sweden is the only randomised controlled trial of AI inside a national cancer screening programme, and it now has full results.

The interim safety analysis, published in The Lancet Oncology in August 2023, covered 80,033 women aged 40 to 80 at four sites in southwest Sweden, randomised one to one between AI-supported reading and standard double reading by two radiologists. Cancer detection was six per 1,000 screened women with AI support against five per 1,000 without, 41 more cancers found. The false-positive rate was 1.5 per cent in both arms. Radiologists performed 46,345 screen readings in the AI arm against 83,231 in the control arm, a 44 per cent reduction in screen-reading workload.

The full results, published in The Lancet in January 2026 with more than 100,000 women and two years of follow-up, tested the question that matters. Interval cancers, the ones diagnosed between screening rounds and generally the more aggressive, fell from 1.76 per 1,000 women in the control arm to 1.55 per 1,000 in the AI arm, a 12 per cent reduction. There were 16 per cent fewer invasive cancers in the interval, 21 per cent fewer large ones and 27 per cent fewer of the aggressive subtypes. Cancers detected at screening rather than between rounds rose from 74 to 81 per cent of all cases.

Read the design rather than the headline. The AI triaged low-risk examinations to a single radiologist and high-risk ones to two, and marked suspicious findings for the human reader. The radiologist kept the recall decision throughout. The first author's own summary is that the study does not support replacing clinicians, because at least one radiologist still reads every case. The trial ran on one mammography device with one AI system in one country, with moderately to highly experienced readers, and the authors say so.

What was tested here was a redesigned workflow with a human decision-maker inside it, evaluated on patient outcomes over years. Almost nothing else in clinical AI has been held to that standard.

The sepsis model that half of American medicine was already running

In 2021, Wong and colleagues at Michigan Medicine published the first serious external validation of the Epic Sepsis Model, a proprietary early-warning tool then deployed across hundreds of US hospitals. They studied 27,697 patients across 38,455 hospitalisations, of which 7 per cent involved sepsis.

The model achieved an area under the curve of 0.63, against the 0.76 to 0.83 its developer had cited. At the alert threshold the hospital was actually using, sensitivity was 33 per cent, specificity 83 per cent and positive predictive value 12 per cent. It failed to identify 1,709 of the 2,552 septic hospitalisations, 67 per cent, and 60 per cent of those patients received timely antibiotics anyway. Meanwhile it crossed the alert threshold in 18 per cent of all hospitalisations, meaning clinicians would evaluate eight patients to find one who eventually became septic.

The authors' own conclusion is the one worth keeping: widespread adoption of a model performing this poorly raises concerns about sepsis management nationally. A tool had been sold, procured, installed and wired into clinical alerting at scale before anyone outside the vendor measured it against outcomes.

Screening and prediction are not the same problem. Screening AI reads a stable, standardised image against a well-defined target. Prediction models forecast a physiological trajectory from data whose meaning shifts between hospitals, coding practices and years. The first has produced randomised evidence. The second has produced a cautionary tale about buying capability on the strength of internal figures.

Forty-one seconds a note, in one arm out of two

Ambient documentation is the fastest-spreading clinical AI in the world and the one with the loosest claims attached to it. A three-arm pragmatic randomised trial at UCLA, run from November 2024 to January 2025, gave 238 outpatient physicians across 14 specialties either Microsoft DAX, Nabla, or usual care.

Time spent writing a note fell by 41 seconds in the Nabla arm and 18 seconds in the control arm, a relative difference of 9.5 per cent that reached significance. The DAX arm fell by 23 seconds and did not differ significantly from doing nothing. Physicians using either scribe reported better scores on the Mini-Z burnout instrument, lower task load and lower work exhaustion, with no significant difference between the two products on any of those measures.

Three details in the paper deserve more attention than the headline. Roughly 15 per cent of physicians given a scribe never used it once. Usage rate correlated with time saved, so the average conceals a large gap between the doctors who adopted the tool and those who did not. And the authors flag a measurement problem affecting every study of this kind: the electronic record's own time metrics do not count editing done inside the scribe application, so reported savings may be overstated.

A minute a note is a real gain across a career. It is also roughly a hundredth of what the category is usually sold as delivering, and the wellbeing effect is larger and better evidenced than the time effect. Those are different value propositions and a service should know which one it is buying.

Adding a model to a doctor added nothing

The most uncomfortable trial in clinical AI is small and clean. Goh and colleagues randomised 50 US-licensed physicians, 26 attendings and 24 residents, to work through clinical vignettes with either GPT-4 or conventional resources such as UpToDate and Google, graded blind against a validated diagnostic reasoning rubric.

Median score was 76 per cent with the model and 74 per cent without: an adjusted difference of 2 percentage points, confidence interval minus 4 to plus 8. Median time per case was 519 seconds against 565, a difference that did not reach significance either. Every physician in the model arm used it.

Then the finding that made the study famous. The model working alone scored 92 per cent, sixteen percentage points above the physicians using conventional resources. The capability was present in the room and did not transfer to the people.

The authors' explanation is about prompting and interaction design rather than about doctors, and they are careful to say the result does not license autonomous diagnosis. Six curated vignettes are not clinical practice, and the design deliberately strips out interviewing, examination and context, which is most of what a clinician does. Take the study for what it establishes: the performance of a human-plus-model pairing is not the sum of the two, and cannot be assumed from either. That is the same shape as the jagged frontier result in consulting and the meta-analytic finding that human-AI combinations frequently underperform the better of their two parts.

The average radiologist does not exist

Yu and colleagues gave 140 radiologists AI assistance across 15 chest X-ray diagnostic tasks, roughly 5,190 observations, with empirical-Bayes shrinkage to separate genuine individual differences from noise. The effect of the assistance ranged from strongly positive to strongly negative between individuals. Experience, subspecialty and prior familiarity with AI all failed to predict who would benefit, and lower performers did not reliably gain the most.

A department that issues the same tool to everyone is therefore helping some clinicians and harming others, with no current means of telling which in advance. That is a governance problem rather than a technology problem, and it points at monitoring individual performance after deployment rather than at better procurement.

Poland, 1,443 colonoscopies, and the first real number on deskilling

Budzyn and colleagues published the finding this research regards as the most important single result in clinical AI. Nested in the ACCEPT trial across four Polish endoscopy centres, they examined 1,443 colonoscopies performed without AI assistance: 795 before the centres introduced AI and 648 afterwards, by 19 endoscopists averaging 28 years of experience.

Adenoma detection in those unassisted procedures fell from 28.4 per cent before AI exposure to 22.4 per cent after, a drop of 6.0 percentage points, p=0.0089, adjusted odds ratio 0.69.

Note what is being measured. Not performance with the tool, which improved. Performance without it, in highly experienced professionals, within months. The study is observational rather than randomised, other changes over the period cannot be fully excluded, and detection rate is a proxy for skill rather than skill itself. It establishes that the effect is real and reachable in ordinary practice, not how common it is.

This is capability debt with a number attached: the cost is not visible while the tool is present and appears only when it is absent. It is also the clearest available instance of what happens when a professional stops taking the repetitions that built the judgement, which this research has called the problem of missed reps.

Medicine already owns the answer, from another safety-critical trade

Aviation faced this in the 1980s and answered it structurally rather than culturally. Line pilots undergo recurrent proficiency checks that can be failed, with hand-flying assessed periodically without the automation. The finding from that literature that transfers most directly is that the manual skills degrade more slowly than the cognitive ones: what erodes first is knowing what the system is doing and why, not the hands.

Applied to a clinical service, that suggests three things follow from Budzyn rather than from principle. Measure unassisted performance periodically, on a schedule, as a condition of continuing to use the tool. Treat a rising gap between assisted and unassisted performance as a signal about the service rather than about the individual. And record how long a clinician actually spends per AI-flagged item, since time per decision is the most diagnostic and least collected number in any oversight arrangement.

Medicine has the apparatus for this already. It runs revalidation, appraisal, audit and mortality review. What it does not yet have is a standing unassisted baseline for tasks where AI has been introduced. The full argument is set out in what professions can learn from aviation.

What has not been shown, and should be said out loud

A page like this is worth more for what it declines to assert.

Six questions before a clinical service deploys anything

Drawn from the failures above rather than from a framework.

Key research and primary sources

Every figure on this page was checked against the primary source. Where a number in wide circulation could not be traced to a stated methodology, it was removed rather than repeated.

Related SuperSkills research

On the mechanism, deskilling, capability debt and missed reps. On which professions carry the most exposure, the deskilling risk map. On the structural answer, what professions can learn from aviation and whether a lost skill comes back. On oversight, human in the loop is not a safeguard, meaningful human oversight and who supervises work they cannot do. On the adjacent consumer question, AI for therapy or advice.

About this research

Rahim Hirji is the author of SuperSkills (Kogan Page, 2026) and founder of The SuperSkills Intelligence Company. Every trial cited here was read at the primary source and every figure checked against it; the MASAI outcome figures are taken from the trial publications and the journal's own press summaries, which is the weaker basis of the two and is flagged for that reason. The aviation parallel is an inference this research draws, not a finding from any clinical study. Not medical advice. Reviewed quarterly.

Cite this

Hirji, R. (2026). How will AI change medicine? The SuperSkills Intelligence Company. Last reviewed 28 August 2026. thesuperskills.com/research/how-will-ai-change-medicine

In this hub

AI and Human Judgement

Does AI weaken judgement? The evidence, and what to do about it.

The work

Where the writing comes from.

These essays draw on research across more than 200 organisations in 30 countries. See the wider body of work, or bring it into your organisation.

All research →