- Is AI dangerous?
- What are the biggest measured harms from AI so far?
- Could AI cause human extinction?
- Can AI help make cyber or biological weapons?
- How likely is human extinction from AI, according to the people building it?
The question gets asked as one question and answered as one question. It is three. Each of the three has a different evidence base, and they are not close in strength. One category is being measured now, in published studies, with numbers attached. One is plausible and largely unproven. One is argued rather than measured. Public attention has settled almost entirely on the third, which is the one that can be discussed indefinitely without anybody having to check anything.
The short answer#
Yes, in ways that are already measurable and mostly undiscussed. Probably, in ways that are plausible and still unproven. Possibly, in ways nobody can currently measure at all. Anyone who answers with a single yes or no has picked one of the three and hidden the choice.
The practical problem for an organisation is that the three compete for the same finite attention, and the ranking by drama runs opposite to the ranking by evidence.
One: the harms with numbers on them#
These are documented in peer-reviewed literature, with samples and effect sizes, and they are happening at scale now.
The scientific record is taking fabricated citations. An audit published in The Lancet ran a pipeline over 2.5 million papers and 125.6 million references. It found 4,046 references to studies that do not exist, across 2,810 papers, and the rate of affected papers rose from 1 in 2,828 in 2023 to 1 in 277 in early 2026. At the time of the audit, 98.4 per cent of the affected papers had received no publisher action. The authors are careful that this is a floor rather than a ceiling, and that it covers an open-access collection rather than the whole of PubMed.
Experienced professionals lose capability within months. A study across four Polish endoscopy centres looked at 1,443 colonoscopies performed without AI assistance by nineteen endoscopists averaging 27.6 years of experience. Adenoma detection in the unassisted procedure fell from 28.4 per cent before the centres adopted AI to 22.4 per cent after, six percentage points, at p=0.0089. The study is observational rather than randomised, covering one procedure in one country. It is also the clearest direct measurement anybody has of a skill degrading in people who had spent decades acquiring it.
Writing with a model changes what the writer believes. An experiment with 1,506 participants gave them a writing assistant configured to argue one side of a question. It moved the opinions in their finished text, and it moved their own opinions in an attitude survey afterwards, including among people who had plenty of time to write independently. The authors call it latent persuasion. One topic, one configuration, so the generalisation is open.
Notice what these three share. None of them looks like a risk while it is happening. They look like speed, like assistance, like a better draft. That is the reason they go ungoverned: an organisation watching for danger is watching for something that announces itself.
Two: the harms that are plausible and unproven#
The International AI Safety Report 2026, written by over a hundred experts with an advisory panel nominated by more than thirty countries, is the best available account of this middle category, and its value is that it declines to resolve it.
On cyber, the report finds that general-purpose AI can identify software vulnerabilities and write and execute code to exploit them, and that criminal groups and state-associated attackers are using it in their operations. It also finds that the largest role is in scaling the preparatory stages, and that systems are not executing attacks autonomously. On biological and chemical risk, it finds systems can produce laboratory instructions and troubleshoot procedures, lowering barriers, while stating that substantial uncertainty remains about how far this raises real-world risk given the practical difficulty of actually producing such weapons.
On manipulation the report is more deflating than most coverage of it. AI-generated content produces measurable belief change in experimental settings, and there is still little evidence of manipulation at scale in the wild, partly because manipulative content is hard to detect and so hard to count.
That combination, real capability with unresolved consequence, is uncomfortable to hold and easy to collapse in either direction. Collapsing it upward produces the headline. Collapsing it downward produces the reassurance. The report does neither, and names the position it leaves decision-makers in: an evidence dilemma, where capability moves quickly and evidence about new risks arrives slowly, so acting early risks entrenching the wrong intervention and waiting risks leaving people exposed.
Three: the harms that are argued rather than measured#
Loss of control, and its endpoint in extinction arguments, is the category that dominates the public conversation. The Safety Report's finding on it is short: expert views vary widely, and current systems show at most early signs of the relevant behaviours.
The numbers people quote come from a survey of 2,778 AI researchers, and the survey did something useful with them. Different respondents from the same population were given differently worded versions of the question. Asked about future AI advances causing human extinction or similarly permanent and severe disempowerment, the median answer was 5 per cent. Asked about human inability to control advanced AI causing the same outcome, the median was 10 per cent. One changed clause, double the figure. Between 41.2 and 51.4 per cent gave more than a one in ten chance depending on how it was put.
This is treated as a gotcha in both directions and is neither. It does not show the risk is imaginary; a population of informed people giving five per cent to human extinction is a serious statement whichever wording produced it. What it shows is that the figure is partly an artefact of the question, so quoting one number without its wording drops the part that determined it. The survey authors say as much, note that their respondents are experts in AI rather than trained forecasters, and cite a related study where framing moved lay estimates by nearly six orders of magnitude.
September 2026: the numbers from inside the labs#
The third category acquired new figures in the week of 8 September 2026, and they are worth setting beside the survey. After Jacob Coxon resigned from Anthropic with a post saying the companies were “gambling with our lives”, TIME reported Evan Hubinger, Anthropic’s head of alignment stress testing, putting the chance of AI killing all humans within the decade at more than 10 per cent; on 15 September TIME quoted Geoffrey Irving at 50 per cent and Marcus Williams of OpenAI at 70 per cent, while David Bellamy of the Institute of Foundation Models called the virus scenario “total bogus” and Jensen Huang of Nvidia told the BBC the extinction claims were “complete nonsense”. These are individual estimates from people who, like the survey respondents, are researchers rather than forecasters. They widen the range; they do not change its category. The one systematic measurement remains Grace et al. and its 5 or 10 per cent, and the same week produced no new evidence of the kind that fills the first category. What the week did produce, in the agent incidents, sits in the first two categories and is taken up on its own page; what a board should do with all of it is here.
Why the three get collapsed#
The categories are ranked by evidence in one order and by interest in the opposite one. A citation that does not exist is a boring fact. A clinician six percentage points worse at finding polyps is a boring fact. The end of the species is not boring, and it requires no data to discuss, so it expands to fill whatever attention is available.
Inside an organisation that trade is expensive, because attention to risk is a budget like any other. A board that has spent its AI risk conversation on scenarios has usually spent nothing on the first category, which is the one already showing up in its own work: the verification nobody is doing, the capability quietly leaving, the judgement that was formed by a system before anybody formed their own.
Which category the speaker is arguing from#
- Ask which category the speaker is in. Before agreeing or disagreeing, establish whether the claim is measured, plausible or argued. Most AI risk arguments are two people in different categories talking past each other.
- Govern the boring one first. The measured harms are the ones inside your own building, and they are cheaper to address than anything in the other two. Nobody has to solve alignment to check whether a citation exists.
- Quote a risk number with its question attached. Five per cent and ten per cent came from the same researchers in the same fortnight. The wording is part of the finding.
- Do not treat uncertainty as an answer. The evidence dilemma is real and it is not permission to wait. It is a reason to prefer interventions that stay useful whether or not the uncertainty resolves, starting with measuring what people can still do unaided.
The third category is not being dismissed here#
Nothing here says the third category is unreal, or that the people working on it are wasting their time. A low-probability outcome of unlimited severity is a legitimate object of study, and the argument that it deserves attention proportionate to its stakes rather than to its evidence is a serious one that this page does not refute.
The measured harms also have a coverage problem running the other way. They are measured where measurement is cheap: published literature, endoscopy, short writing tasks. Law, management and strategy have the same mechanisms and none of the studies, so the first category is almost certainly larger than the part of it anybody can currently count.
Evidence review · SS-2026-211 · Graded against the published rubric
Hirji, R. (2026). Is AI dangerous?. The SuperSkills evidence base, SS-2026-211. https://thesuperskills.com/research/is-ai-dangerous. Last reviewed 17 September 2026.
An evidence review by Rahim Hirji, not peer-reviewed research. For a material claim, cite the underlying study as well; every study here carries its own permanent link.
How citations and IDs work