Alarm fatigue is what happens to a person who is warned too often. The warnings stop registering, then they get silenced, and eventually the one that mattered goes the same way as the thousands that did not. It is the best measured failure of a safety control anywhere in the record. The standard answer to AI oversight, flag the risky cases and have a human check those, inherits that failure whole.
The answer, in one line
It is the desensitisation of a person to warning signals caused by exposure to a high volume of them, most of which turn out not to require action.
Definition#
Alarm fatigue: the desensitisation of a person to warning signals caused by exposure to a high volume of them, most of which turn out not to require action, leading to slower responses, silenced alarms and missed true events. An established term in patient safety and human factors, used here in its original sense and not claimed by this research.
The number that makes the argument#
Barbara Drew and colleagues instrumented five adult intensive care units at the University of California, San Francisco and stored every monitor signal for the thirty-one days of March 2013: seven ECG leads, pressure, oxygen saturation and respiration waveforms, user settings and every alarm, for 461 consecutive patients. The count was 2,558,760 unique alarms in a single month, made up of 1,154,201 arrhythmia alarms, 612,927 parameter alarms and 791,632 technical alarms. Of those, 381,560 were audible, an audible burden of 187 per bed per day.
Nurse scientists then annotated 12,671 arrhythmia alarms against a defined protocol, with 95 per cent agreement between raters and a Cohen kappa of 0.86. 88.8 per cent of those were false positives. Among the true ventricular tachycardia alarms, 93 per cent were not sustained long enough to warrant treatment.
That figure is quoted carelessly and this page will not join in. The 88.8 per cent applies to the 12,671 annotated arrhythmia alarms and to nothing else. It does not describe the 2.56 million, and a page saying that 88.8 per cent of clinical alarms are false has misread the paper. The study is single-centre, one month, five units, and it was funded by GE Healthcare.
Ignoring them is the rational response#
A clinician hearing 187 audible alarms a bed a day, where nearly nine in ten of the annotated ones carried no signal, is receiving an instruction about what those sounds mean. Treating them as noise is an accurate inference from the evidence available, and the language of fatigue slightly misdescribes it: this is a base rate doing what base rates do. The person has learned the true positive rate of the instrument and adjusted, and the adjustment is correct on average and catastrophic on the exception.
The Joint Commission documented the exception. Its Sentinel Event Alert 50, issued on 8 April 2013, recorded 98 alarm-related events between January 2009 and June 2012, of which 80 resulted in death, 13 in permanent loss of function and five in unexpected additional care or extended stay, with 94 occurring in hospitals. The contributing factors it recorded are the shape of the behaviour: alarm signals inappropriately turned off (36), absent or inadequate alarm system (30), alarm signals not audible in all areas (25), improper alarm settings (21). Clinicians turn the volume down, turn the alarm off, or set it outside safe limits because of the volume of signals. The Alert led to National Patient Safety Goal NPSG.06.01.01, phased in from 1 July 2014, which made the authority to change or silence an alarm something an organisation has to assign to named people.
The Commission's own footnote limits what those 98 events can be used for: reporting is voluntary, represents only a small proportion of actual events, and supports no conclusion about relative frequency or trend. The widely repeated claim that 85 to 99 per cent of alarm signals require no clinical intervention is quoted by the Commission from AAMI Horizons in 2011 and is not a Joint Commission measurement; the estate has not read the AAMI source and does not use the range.
A designer making the trade explicitly, and the cost of it#
The clearest record of an engineer weighing false alarms against missed events sits in a road accident investigation. In its report on the fatal collision in Tempe, Arizona on 18 March 2018, the National Transportation Safety Board found that the developer had disengaged the vehicle's factory forward collision warning and automatic emergency braking during automated operation, and that on detecting an emergency the system entered a one-second period of action suppression, withholding braking while it verified the hazard or waited for the operator to take control. No alert was given to the operator when action suppression began. The system had detected the pedestrian 5.6 seconds before impact and tracked her all the way without ever classifying her correctly. The stated reason for suppressing the response was concern about false alarms.
The NTSB determined the probable cause to be the operator's failure to monitor the driving environment while visually distracted, and found she would likely have had time to react had she been attentive, so nobody should say the missing alert caused the crash. What the report does establish is a design in which a human being was named as the primary countermeasure in an emergency and simultaneously denied the signal that would have let them act as one, with the volume of false alarms given as the reason. That is alarm fatigue anticipated by the engineer and paid for in advance.
Why this is the first thing to say about AI oversight#
Lisanne Bainbridge saw in 1983 that a person cannot watch a quiet system indefinitely, and her proposal was to hand the watching to an automatic alarm system connected to sound signals. Medicine then ran that experiment for forty years at enormous scale. The result is the evidence above: the remedy relocates the problem to whoever must now decide which alarms deserve a response, and it relocates it into a form where the correct everyday behaviour and the correct exceptional behaviour are opposites.
Most proposals for AI oversight have the same shape. Route the uncertain outputs to a reviewer. Flag the high-risk decisions. Surface a confidence score. Each creates a queue whose base rate the reviewer will learn, and the fluency of model output makes the learning worse, because a false flag on a well-written answer looks like a false flag rather than like a near miss. An oversight design that generates more signals than a person can act on has not added a control. It has added a number that an auditor can count.
Four questions that tell you whether a flag is a control#
- What is the true positive rate, and does the reviewer know it? A reviewer working a queue at nine per cent precision will behave accordingly, and the organisation should assume so rather than discover it.
- Is the volume inside what one person can actually action? Count flags per reviewer per day against the time each one needs. If the arithmetic fails, the control is nominal, in the sense set out at designing a stop button people will use.
- Who is allowed to turn it down, and is that written anywhere? The Joint Commission had to make this an explicit assignment because informal silencing was killing people. Most AI review queues have no equivalent rule and get silenced by threshold changes nobody records.
- Does anybody see the misses? A queue that reports only what it caught cannot tell you its own recall. Sampling the unflagged population is the only way to know whether the control works, and the first thing cut when the queue gets long.
Key sources
- Drew, B. J., Harris, P., Zegre-Hemsey, J. K., Mammone, T., Schindler, D., Salas-Boni, R. et al. (2014). Insights into the Problem of Alarm Fatigue with Physiologic Monitor Devices. PLoS ONE, 9(10), e110274. DOI 10.1371/journal.pone.0110274. Counts confirmed at source 16 September 2026. Graded entry.
- The Joint Commission (2013). Sentinel Event Alert 50: Medical device alarm safety in hospitals. Issue 50, 8 April 2013. Graded entry.
- National Transportation Safety Board (2019). Collision Between Vehicle Controlled by Developmental Automated Driving System and Pedestrian, Tempe, Arizona, March 18, 2018. NTSB/HAR-19/03. Graded entry.
- Bainbridge, L. (1983). Ironies of automation. Automatica, 19(6), 775 to 779. Graded entry.
Related SuperSkills research#
The argument this completes is the ironies of automation, and the limit that made an alarm seem necessary is the vigilance decrement. On controls that exist and do not work, designing a stop button people will use and human in the loop is not a safeguard. On the reviewer's side of it, automation complacency and who owns verification when AI does the work. On what a genuine oversight standard requires, meaningful human oversight.
Explainer · SS-2026-255 · Graded against the published rubric
Hirji, R. (2026). What is alarm fatigue?. The SuperSkills evidence base, SS-2026-255. https://thesuperskills.com/research/what-is-alarm-fatigue. Last reviewed 16 September 2026.
An evidence review by Rahim Hirji, not peer-reviewed research. For a material claim, cite the underlying study as well; every study here carries its own permanent link.
How citations and IDs work