Not as your only guide to what is wrong or how urgent it is. The best test so far is a preregistered randomised study in Nature Medicine in February 2026. Given ten doctor-written scenarios, the language models on their own named the right condition in 94.9 per cent of cases. The 1,298 UK adults who used those same models named a relevant condition in fewer than 34.5 per cent of cases, no better than adults using whatever they would normally use at home. It is one study, built on written vignettes and three specific models, so it locates a gap and does not measure your situation.
The answer, in one line
The best randomised evidence says be careful. In a 2026 Nature Medicine study of 1,298 UK adults, people using language models named a relevant condition in fewer than 34.5 per cent of cases, no better than a control group, although the models alone scored 94.9 per cent.
Definition#
Disposition: the recommended next step for a set of symptoms: how urgently to seek care, and where. It is the second of the two outcomes in Bean and colleagues' study, alongside naming the condition. The models scored 56.3 per cent on it on average when tested alone.
Alone the models passed, and with people they did not#
Bean and ten colleagues recruited 1,298 UK adults and gave each one of ten medical scenarios written and agreed by three doctors. Participants used GPT-4o, Llama 3 or Command R+, or a control condition of any method they would normally use at home. Tested alone, the models identified conditions in 94.9 per cent of cases and the right disposition in 56.3 per cent on average. Participants using the same models identified relevant conditions in fewer than 34.5 per cent of cases and the right disposition in fewer than 44.2 per cent, both no better than the control group. The paper appeared in Nature Medicine 32(2) on 9 February 2026, and a later Publisher Correction fixed the axis labels on one figure and nothing else.
The shortfall sat in the conversation, not in the model#
The authors looked at why. In 16 of 30 sampled interactions the first message contained only partial information about the scenario. The models suggested 2.21 possible conditions per interaction, of which 34.0 per cent were correct, and participants ended up listing 1.33 on average. Correct suggestions did appear in conversations, and users did not consistently carry them into their final answers. So two things go wrong at once: the person may not tell the tool what matters, and may not pick out the right item from what comes back.
What ten vignettes cannot say about your symptoms#
- Written cases, not felt ones. The authors note that vignettes leave out the urgency and stress of real symptoms.
- Specific models. GPT-4o, Llama 3 and Command R+ were tested. The authors describe the result as a lower bound for newer models, and this page has no later measurement to add.
- No search-engine or pharmacist comparison. The control group used whatever they normally would, so the study shows models did not beat that baseline, not that they are worse than a named alternative.
- Weak even alone on urgency. A 56.3 per cent average on disposition means the models themselves were unreliable at deciding how quickly to act.
- Ten scenarios and paid participants. Each participant was paid £2.25, which may differ from how someone behaves when worried.
Understanding a result is a different job from triage#
This section is interpretation, kept apart from the evidence above.
The study tested the hardest use, deciding what is wrong and what to do. It did not test explaining a diagnosis you already have, turning a clinician's jargon into plain words, or preparing questions for an appointment. Nothing here shows those uses are safe or unsafe. The reasoning is that a job with a checkable answer and a professional still in the loop asks less of both the tool and you than triage does, and that is an inference from the shape of the failure and not a finding.
Three habits follow from where the study located the failure. Give the tool the whole picture, because the partial first message was the commonest problem. Ask it to say what else it would need to know, and answer that. Treat any list of conditions as questions for a clinician and not as a ranking. For symptoms that could be an emergency, contact emergency services or your health service directly. The study is no reason to delay that call, and it did not test delaying it.
The confident tone that makes a chatbot easy to believe is a separate problem, covered in why does AI sound so confident when it is wrong.
Key sources
- Bean, A. M., Payne, R. E., Parsons, G., Kirk, H. R., Ciro, J., Mosquera-Gomez, R., Hincapie M, S., Ekanayaka, A. S., Tarassenko, L., Rocher, L. and Mahdi, A. (2026). Reliability of LLMs as medical assistants for the general public: a randomized preregistered study. Nature Medicine 32(2), 609 to 615. Graded entry.
Related SuperSkills research#
On the profession, how will AI change medicine. On emotional support, is it safe to use AI for therapy. On checking outputs, how do I know when AI is wrong. On another high-stakes personal domain, can I trust an AI chatbot for financial advice.
About this research#
Rahim Hirji is the author of SuperSkills (Kogan Page, 2026), keynote speaker on AI and human capability, and founder of The SuperSkills Intelligence Company. The study's authors, sample, figures and correction notice were read at source on 30 September 2026. The study is graded in the evidence base. This page is information about a research finding and is not medical advice. Disposition is used here in the paper's sense and is not a SuperSkills term.
Essay · SS-2026-373
Hirji, R. (2026). Should I use AI for medical questions?. The SuperSkills evidence base, SS-2026-373. https://thesuperskills.com/research/should-i-use-ai-for-medical-questions. Last reviewed 30 September 2026.
An evidence review by Rahim Hirji, not peer-reviewed research. For a material claim, cite the underlying study as well; every study here carries its own permanent link.
How citations and IDs work