← Research
Research · Question

Should I use AI for medical questions?

Models can pass the test that people using them fail. The gap sits in the conversation, so that is where to be careful.

Last reviewed: 30 September 2026 · Next review due: 30 September 2027

What a preregistered randomised study of 1,298 UK adults found about using chatbots to work out what is wrong, why the shortfall arose, five limits on applying it to your own symptoms, and what it leaves untested.

Question this page answersAll 1060 questions this research covers

Not as your only guide to what is wrong or how urgent it is. The best test so far is a preregistered randomised study in Nature Medicine in February 2026. Given ten doctor-written scenarios, the language models on their own named the right condition in 94.9 per cent of cases. The 1,298 UK adults who used those same models named a relevant condition in fewer than 34.5 per cent of cases, no better than adults using whatever they would normally use at home. It is one study, built on written vignettes and three specific models, so it locates a gap and does not measure your situation.

The answer, in one line

The best randomised evidence says be careful. In a 2026 Nature Medicine study of 1,298 UK adults, people using language models named a relevant condition in fewer than 34.5 per cent of cases, no better than a control group, although the models alone scored 94.9 per cent.

Share as a card

Definition#

Disposition: the recommended next step for a set of symptoms: how urgently to seek care, and where. It is the second of the two outcomes in Bean and colleagues' study, alongside naming the condition. The models scored 56.3 per cent on it on average when tested alone.

Share this definition as a card

Alone the models passed, and with people they did not#

Bean and ten colleagues recruited 1,298 UK adults and gave each one of ten medical scenarios written and agreed by three doctors. Participants used GPT-4o, Llama 3 or Command R+, or a control condition of any method they would normally use at home. Tested alone, the models identified conditions in 94.9 per cent of cases and the right disposition in 56.3 per cent on average. Participants using the same models identified relevant conditions in fewer than 34.5 per cent of cases and the right disposition in fewer than 44.2 per cent, both no better than the control group. The paper appeared in Nature Medicine 32(2) on 9 February 2026, and a later Publisher Correction fixed the axis labels on one figure and nothing else.

The shortfall sat in the conversation, not in the model#

The authors looked at why. In 16 of 30 sampled interactions the first message contained only partial information about the scenario. The models suggested 2.21 possible conditions per interaction, of which 34.0 per cent were correct, and participants ended up listing 1.33 on average. Correct suggestions did appear in conversations, and users did not consistently carry them into their final answers. So two things go wrong at once: the person may not tell the tool what matters, and may not pick out the right item from what comes back.

What ten vignettes cannot say about your symptoms#

Understanding a result is a different job from triage#

This section is interpretation, kept apart from the evidence above.

The study tested the hardest use, deciding what is wrong and what to do. It did not test explaining a diagnosis you already have, turning a clinician's jargon into plain words, or preparing questions for an appointment. Nothing here shows those uses are safe or unsafe. The reasoning is that a job with a checkable answer and a professional still in the loop asks less of both the tool and you than triage does, and that is an inference from the shape of the failure and not a finding.

Three habits follow from where the study located the failure. Give the tool the whole picture, because the partial first message was the commonest problem. Ask it to say what else it would need to know, and answer that. Treat any list of conditions as questions for a clinician and not as a ranking. For symptoms that could be an emergency, contact emergency services or your health service directly. The study is no reason to delay that call, and it did not test delaying it.

The confident tone that makes a chatbot easy to believe is a separate problem, covered in why does AI sound so confident when it is wrong.

Key sources

On the profession, how will AI change medicine. On emotional support, is it safe to use AI for therapy. On checking outputs, how do I know when AI is wrong. On another high-stakes personal domain, can I trust an AI chatbot for financial advice.

About this research#

Rahim Hirji is the author of SuperSkills (Kogan Page, 2026), keynote speaker on AI and human capability, and founder of The SuperSkills Intelligence Company. The study's authors, sample, figures and correction notice were read at source on 30 September 2026. The study is graded in the evidence base. This page is information about a research finding and is not medical advice. Disposition is used here in the paper's sense and is not a SuperSkills term.

How this research works  ·  Reviewed quarterly  ·  Found an error? Tell me and it is corrected on the page.

Essay · SS-2026-373

Cite this page

Hirji, R. (2026). Should I use AI for medical questions?. The SuperSkills evidence base, SS-2026-373. https://thesuperskills.com/research/should-i-use-ai-for-medical-questions. Last reviewed 30 September 2026.

An evidence review by Rahim Hirji, not peer-reviewed research. For a material claim, cite the underlying study as well; every study here carries its own permanent link.

How citations and IDs work
Questions answered on this page

Can I use ChatGPT to diagnose myself?

The best randomised evidence says be careful. In a 2026 Nature Medicine study of 1,298 UK adults, people using language models named a relevant condition in fewer than 34.5 per cent of cases, no better than a control group, although the models alone scored 94.9 per cent. The scenarios were written vignettes, so it does not measure real symptoms.

How accurate is AI for medical advice?

Tested alone on ten doctor-written scenarios, three models named the condition in 94.9 per cent of cases and chose the right level of care in 56.3 per cent on average. Accuracy in the hands of ordinary users was much lower. The gap came from incomplete information going in and correct suggestions being missed coming out.

Should I use AI to decide whether to see a doctor?

The study found users chose the right level of care in fewer than 44.2 per cent of cases, no better than the control group. That is a reason not to rely on a chatbot alone for urgency. If symptoms could be an emergency, contact emergency services or your health service directly.

Why did people do worse than the AI alone?

In 16 of 30 sampled interactions the first message held only partial information. The models suggested 2.21 conditions per interaction, of which 34.0 per cent were correct, and users did not consistently include the correct ones in their final answers.

Are newer AI models better at medical questions?

The authors call their result a lower bound for newer models, so improvement is plausible. The study tested GPT-4o, Llama 3 and Command R+, and this page has no later randomised measurement to report.

In this hub

Everyday life

The same questions, asked about your own week rather than your organisation.

Ask the evidence
What does the evidence actually show?What should our board be asking about this?Where does Rahim disagree with the consensus?
Bring this into your organisation

If this describes something happening in your teams, say so.

Keynotes, board sessions and advisory work, drawing on research across more than 200 organisations in 30 countries. Tell me the room, the date and the shift you need. A reply within 24 hours.

Start a conversation

Topics and audiences  ·  All research

Box of Amazing

Rahim’s free weekly letter on AI and human capability

If this was useful, the weekly letter is where the thinking happens first. Most of what ends up on this site starts there. Weekly essays on AI, capability and the future of work. Read by 25,000 people, every week since 2017. Free, and one click to stop.

Opens Substack to confirm. No pitch in it, unsubscribe in one click, and nobody follows up because you read something.

Running an event, or responsible for how AI arrives in your organisation? Keynotes  ·  Advisory  ·  Boards  ·  Enquire