← Research
Research

Is AI dangerous?

Three questions get asked as one, and public attention has settled on the one with the least evidence.

Last reviewed: 17 September 2026

Some harms are measured now, with samples and effect sizes. Some are plausible and unproven. Some are argued rather than measured. The ranking by evidence runs opposite to the ranking by drama, and inside an organisation that trade is expensive.

Questions this page answersAll 811 questions this research covers

The question gets asked as one question and answered as one question. It is three. Each of the three has a different evidence base, and they are not close in strength. One category is being measured now, in published studies, with numbers attached. One is plausible and largely unproven. One is argued rather than measured. Public attention has settled almost entirely on the third, which is the one that can be discussed indefinitely without anybody having to check anything.

The answer, in one line

It is three questions with three different evidence bases.

Share as a card

The short answer#

Yes, in ways that are already measurable and mostly undiscussed. Probably, in ways that are plausible and still unproven. Possibly, in ways nobody can currently measure at all. Anyone who answers with a single yes or no has picked one of the three and hidden the choice.

The practical problem for an organisation is that the three compete for the same finite attention, and the ranking by drama runs opposite to the ranking by evidence.

One: the harms with numbers on them#

These are documented in peer-reviewed literature, with samples and effect sizes, and they are happening at scale now.

The scientific record is taking fabricated citations. An audit published in The Lancet ran a pipeline over 2.5 million papers and 125.6 million references. It found 4,046 references to studies that do not exist, across 2,810 papers, and the rate of affected papers rose from 1 in 2,828 in 2023 to 1 in 277 in early 2026. At the time of the audit, 98.4 per cent of the affected papers had received no publisher action. The authors are careful that this is a floor rather than a ceiling, and that it covers an open-access collection rather than the whole of PubMed.

Experienced professionals lose capability within months. A study across four Polish endoscopy centres looked at 1,443 colonoscopies performed without AI assistance by nineteen endoscopists averaging 27.6 years of experience. Adenoma detection in the unassisted procedure fell from 28.4 per cent before the centres adopted AI to 22.4 per cent after, six percentage points, at p=0.0089. The study is observational rather than randomised, covering one procedure in one country. It is also the clearest direct measurement anybody has of a skill degrading in people who had spent decades acquiring it.

Writing with a model changes what the writer believes. An experiment with 1,506 participants gave them a writing assistant configured to argue one side of a question. It moved the opinions in their finished text, and it moved their own opinions in an attitude survey afterwards, including among people who had plenty of time to write independently. The authors call it latent persuasion. One topic, one configuration, so the generalisation is open.

Notice what these three share. None of them looks like a risk while it is happening. They look like speed, like assistance, like a better draft. That is the reason they go ungoverned: an organisation watching for danger is watching for something that announces itself.

Two: the harms that are plausible and unproven#

The International AI Safety Report 2026, written by over a hundred experts with an advisory panel nominated by more than thirty countries, is the best available account of this middle category, and its value is that it declines to resolve it.

On cyber, the report finds that general-purpose AI can identify software vulnerabilities and write and execute code to exploit them, and that criminal groups and state-associated attackers are using it in their operations. It also finds that the largest role is in scaling the preparatory stages, and that systems are not executing attacks autonomously. On biological and chemical risk, it finds systems can produce laboratory instructions and troubleshoot procedures, lowering barriers, while stating that substantial uncertainty remains about how far this raises real-world risk given the practical difficulty of actually producing such weapons.

On manipulation the report is more deflating than most coverage of it. AI-generated content produces measurable belief change in experimental settings, and there is still little evidence of manipulation at scale in the wild, partly because manipulative content is hard to detect and so hard to count.

That combination, real capability with unresolved consequence, is uncomfortable to hold and easy to collapse in either direction. Collapsing it upward produces the headline. Collapsing it downward produces the reassurance. The report does neither, and names the position it leaves decision-makers in: an evidence dilemma, where capability moves quickly and evidence about new risks arrives slowly, so acting early risks entrenching the wrong intervention and waiting risks leaving people exposed.

Three: the harms that are argued rather than measured#

Loss of control, and its endpoint in extinction arguments, is the category that dominates the public conversation. The Safety Report's finding on it is short: expert views vary widely, and current systems show at most early signs of the relevant behaviours.

The numbers people quote come from a survey of 2,778 AI researchers, and the survey did something useful with them. Different respondents from the same population were given differently worded versions of the question. Asked about future AI advances causing human extinction or similarly permanent and severe disempowerment, the median answer was 5 per cent. Asked about human inability to control advanced AI causing the same outcome, the median was 10 per cent. One changed clause, double the figure. Between 41.2 and 51.4 per cent gave more than a one in ten chance depending on how it was put.

This is treated as a gotcha in both directions and is neither. It does not show the risk is imaginary; a population of informed people giving five per cent to human extinction is a serious statement whichever wording produced it. What it shows is that the figure is partly an artefact of the question, so quoting one number without its wording drops the part that determined it. The survey authors say as much, note that their respondents are experts in AI rather than trained forecasters, and cite a related study where framing moved lay estimates by nearly six orders of magnitude.

September 2026: the numbers from inside the labs#

The third category acquired new figures in the week of 8 September 2026, and they are worth setting beside the survey. After Jacob Coxon resigned from Anthropic with a post saying the companies were “gambling with our lives”, TIME reported Evan Hubinger, Anthropic’s head of alignment stress testing, putting the chance of AI killing all humans within the decade at more than 10 per cent; on 15 September TIME quoted Geoffrey Irving at 50 per cent and Marcus Williams of OpenAI at 70 per cent, while David Bellamy of the Institute of Foundation Models called the virus scenario “total bogus” and Jensen Huang of Nvidia told the BBC the extinction claims were “complete nonsense”. These are individual estimates from people who, like the survey respondents, are researchers rather than forecasters. They widen the range; they do not change its category. The one systematic measurement remains Grace et al. and its 5 or 10 per cent, and the same week produced no new evidence of the kind that fills the first category. What the week did produce, in the agent incidents, sits in the first two categories and is taken up on its own page; what a board should do with all of it is here.

Why the three get collapsed#

The categories are ranked by evidence in one order and by interest in the opposite one. A citation that does not exist is a boring fact. A clinician six percentage points worse at finding polyps is a boring fact. The end of the species is not boring, and it requires no data to discuss, so it expands to fill whatever attention is available.

Inside an organisation that trade is expensive, because attention to risk is a budget like any other. A board that has spent its AI risk conversation on scenarios has usually spent nothing on the first category, which is the one already showing up in its own work: the verification nobody is doing, the capability quietly leaving, the judgement that was formed by a system before anybody formed their own.

Which category the speaker is arguing from#

The third category is not being dismissed here#

Nothing here says the third category is unreal, or that the people working on it are wasting their time. A low-probability outcome of unlimited severity is a legitimate object of study, and the argument that it deserves attention proportionate to its stakes rather than to its evidence is a serious one that this page does not refute.

The measured harms also have a coverage problem running the other way. They are measured where measurement is cheap: published literature, endoscopy, short writing tasks. Law, management and strategy have the same mechanisms and none of the studies, so the first category is almost certainly larger than the part of it anybody can currently count.

Evidence review · SS-2026-211 · Graded against the published rubric

Cite this page

Hirji, R. (2026). Is AI dangerous?. The SuperSkills evidence base, SS-2026-211. https://thesuperskills.com/research/is-ai-dangerous. Last reviewed 17 September 2026.

An evidence review by Rahim Hirji, not peer-reviewed research. For a material claim, cite the underlying study as well; every study here carries its own permanent link.

How citations and IDs work
Questions answered on this page

Is AI dangerous?

It is three questions with three different evidence bases. Some harms are measured now: an audit of 2.5 million papers found references to studies that do not exist in 1 in 277 papers by early 2026, and nineteen endoscopists averaging 27.6 years of experience lost six percentage points of unassisted detection within months of their centres adopting AI. Some are plausible and unproven, including cyber and biological uplift. Some are argued rather than measured, including loss of control. A single yes or no picks one category and hides the choice.

What are the biggest measured harms from AI so far?

Three stand out because they carry samples and effect sizes. Fabricated references entering the peer-reviewed literature at a rising rate, with 98.4 per cent of affected papers receiving no publisher action at the time of the audit. Measured deskilling in highly experienced clinicians, six percentage points of adenoma detection lost in unassisted procedures. And latent persuasion, where 1,506 participants writing alongside an opinionated model shifted both what they wrote and what they later reported believing. None of the three looks like a risk while it happens; they look like speed and assistance.

Could AI cause human extinction?

The estimates in circulation come from a survey of 2,778 AI researchers, and they move with the wording. Asked about future AI advances causing human extinction or similarly permanent and severe disempowerment, the median was 5 per cent. Asked about human inability to control advanced AI causing the same outcome, the median was 10 per cent. Between 41.2 and 51.4 per cent gave more than a one in ten chance depending on how it was put. The figure is partly an artefact of the question, and the survey authors note their respondents are experts in AI rather than trained forecasters.

Can AI help make cyber or biological weapons?

The International AI Safety Report 2026 places this in the plausible and unresolved category. It finds general-purpose AI can identify software vulnerabilities and write and execute exploit code, and that criminal and state-associated attackers are using it, while also finding the largest role is in scaling preparation rather than executing attacks autonomously. On biological and chemical risk it finds systems can produce laboratory instructions and troubleshoot procedures, and states that substantial uncertainty remains about how far this raises real-world risk given the practical barriers to producing such weapons.

What should a board actually worry about with AI?

The measured category, because it is the one already inside the building and the cheapest to address. Verification nobody is doing, capability leaving without anyone testing for it, and judgement formed by a system before a person forms their own. A board that has spent its AI risk conversation on scenarios has usually spent nothing here. Attention to risk is a budget, and the ranking by drama runs opposite to the ranking by evidence.

In this hub

Judgement, oversight and accountability

Who decides, who checks, and who is answerable when the machine was involved.

Ask the evidence
What does the evidence actually show?What should our board be asking about this?Where does Rahim disagree with the consensus?
Bring this into your organisation

If this describes something happening in your teams, say so.

Keynotes, board sessions and advisory work, drawing on research across more than 200 organisations in 30 countries. Tell me the room, the date and the shift you need. A reply within 24 hours.

Start a conversation

Topics and audiences  ·  All research

Box of Amazing

Rahim’s free weekly letter on AI and human capability

If this was useful, the weekly letter is where the thinking happens first. Most of what ends up on this site starts there. Weekly essays on AI, capability and the future of work. Read by 25,000 people, every week since 2017. Free, and one click to stop.

Opens Substack to confirm. No pitch in it, unsubscribe in one click, and nobody follows up because you read something.

Running an event, or responsible for how AI arrives in your organisation? Keynotes  ·  Advisory for CEOs and boards  ·  Enquire