The out-of-the-loop performance problem is what happens to a person's ability to take over when the machine they were supervising stops working. They are slower to notice that something has gone wrong, and slower again to establish what it is once they have noticed. Mica Endsley and Esin Kiris named it in Human Factors in 1995 and measured one of its causes. The result that matters most to anyone supervising AI is which half of understanding survived: their participants still saw the information in front of them, and had lost the grasp of what it meant.
The answer, in one line
It is the loss of a person's ability to take over manual operation when an automated system fails, caused by their having been placed in the role of monitor instead of operator.
Definition#
The out-of-the-loop performance problem: the loss of a person's ability to take over manual operation when an automated system fails, caused by their having been placed in the role of monitor instead of operator. Named by Mica Endsley and Esin Kiris in 1995.
The navigation task, and the five levels it was run at#
Endsley and Kiris automated an automobile navigation task using an expert system, and ran it at five separate levels of operator control: manual throughout; the system suggests and the person decides and acts; the system decides and acts with the person's consent required; the system decides and acts unless the person vetoes; and full automation with no operator interaction. Those five levels are Endsley's own scale, set out in her work on expert systems in cockpits in the late 1980s.
Situation awareness was lower under the automated conditions than under manual performance, and decision time after the expert system failed was longer where situation awareness was lower. The out-of-the-loop effect was significantly greater under full automation than under the intermediate levels. Keeping the person inside the decision, at a level below the top of the scale, preserved both their awareness and their ability to resume control.
They still saw the data and no longer knew what it meant#
Endsley's own account of the study, written up in a chapter for Parasuraman and Mouloua's Automation and Human Performance in 1996, reports the split that makes this finding useful outside aviation. Situation awareness in her framework has three levels: perceiving the elements in front of you, comprehending what they mean in relation to your goals, and projecting what the system will do next. Only the second was damaged. Participants remained aware of the low-level data, so they were monitoring the system perfectly well, and had "less comprehension of what the data meant" for the task they were there to accomplish.
Endsley attributes that specifically to passivity. Under the conditions of the experiment the information displayed to operators did not change between conditions, and vigilance and monitoring effects were too small to account for the decrement. Turning a performer into an observer damaged comprehension on its own, in people who were watching attentively and seeing everything they were shown.
Three routes out of the loop#
Endsley sets out three mechanisms by which automation carries a person out of the loop, and they remain the useful decomposition.
Monitoring. People are poor sustained monitors of reliable systems, which is the vigilance and complacency literature the estate covers under automation complacency. Attention drifts towards whatever else is competing for it, and a system that has never failed offers no reason to look.
Passive processing. Observing a decision is a different cognitive act from making one, and it lays down a weaker model of the situation. This is the route Endsley's own experiment implicates, and the one that survives any amount of diligence.
Feedback. Automating a function tends to remove the information that used to arrive as a by-product of doing it. Designers assume the operator no longer needs what the machine now handles, so the cues that would have supported a takeover are gone at the moment they are wanted.
What a 1995 navigation study cannot settle#
This is one laboratory task, run on an expert system that produced route recommendations, thirty years before the tools this site is otherwise about. A generative model differs from that system in the way that matters here: it produces fluent output across every task rather than one recommendation on a defined one, so the operator's comprehension problem may be larger or smaller and nobody has measured it.
Two further limits belong on the record. The paper's sample size is not given in its abstract and the full text sits behind a publisher paywall, so no participant count appears on this page. And the Level 1 against Level 2 split, which is the most quoted thing here, is read from Endsley's 1996 chapter summarising her own study rather than from the results section of the paper itself. Both are stated because the alternative is to imply a closer reading than was possible.
Comprehension is what the arrangement is quietly delegating#
Almost every oversight arrangement in commercial use is written as though the risk were that the person fails to look. Sign-offs, review steps and second pairs of eyes all address attention. Endsley and Kiris measured people who looked, who were correctly monitoring, and who could not act well when the system failed anyway. If that generalises, then the standard remedy treats the symptom the arrangement does not have.
This is the specific mechanism underneath the human-in-the-loop claim, and it gives Article 14 of the EU AI Act a harder edge than the drafting suggests. Article 14 requires that an overseer be enabled to understand the system's capacities and limitations and to intervene. The out-of-the-loop result says the capacity to intervene falls as the level of automation rises, in people who understand the system and are paying attention. No amount of training or instruction reaches it, because the cause is where the person sits in the decision.
The design lever Endsley identified is the level of control. That is the same lever as drift versus design, seen from the human-factors side. Nobody in an organisation chooses level five. Products arrive configured at the top of the scale, approval flows are added to make them feel supervised, and the arrangement that results asks a person to consent to decisions they did not participate in making. Consent and veto both leave the operator passive, which is the condition the experiment was measuring.
Designing the level down#
Name the level for every automated decision. Take the five-point scale above and place each one on it. Most organisations have never asked the question, and the answer is usually four or five for work nobody would have chosen to fully automate.
Give the person something to decide, not something to approve. A step that only permits yes or no keeps them out of the loop while producing a governance record that says otherwise. This is the difference between the second and the fourth level on the scale.
Test the takeover, not the output. Accuracy under normal running says nothing about the state this literature is concerned with. The measure is how long it takes someone to notice a failure and correct it, which almost nobody records and which is the number that would have shown the problem in advance.
Keep the feedback that automation makes redundant. Information that only mattered because a person was doing the work is exactly the information they will need in order to resume it.
Related SuperSkills research#
On the monitoring route, automation complacency and automation bias. On the regulatory test, meaningful human oversight and why human in the loop is not a safeguard. On who is left holding the failure, the moral crumple zone. On the capability question underneath it, who supervises work they cannot do and what is deskilling.
Key sources
- Endsley, M. R. and Kiris, E. O. (1995). The Out-of-the-Loop Performance Problem and Level of Control in Automation. Human Factors, 37(2), 381-394. DOI 10.1518/001872095779064555.
- Endsley, M. R. (1996). Automation and Situation Awareness. In R. Parasuraman and M. Mouloua (Eds.), Automation and Human Performance: Theory and Applications, 163-181. Lawrence Erlbaum.
- Bainbridge, L. (1983). Ironies of Automation. Automatica, 19(6), 775-779.
- Parasuraman, R. and Manzey, D. H. (2010). Complacency and Bias in Human Use of Automation: An Attentional Integration. Human Factors, 52(3), 381-410.
About this research#
Rahim Hirji is the author of SuperSkills (Kogan Page, 2026), keynote speaker on AI and human capability, and founder of The SuperSkills Intelligence Company. The out-of-the-loop performance problem is an established human-factors concept and no SuperSkills term. Findings are attributed to the studies that produced them and kept separate from the interpretation, which is the author's. The graded evidence, including what each study does not support, is in the evidence base. Last reviewed: 7 September 2026.
Explainer · SS-2026-203 · Graded against the published rubric
Hirji, R. (2026). What is the out-of-the-loop performance problem?. The SuperSkills evidence base, SS-2026-203. https://thesuperskills.com/research/what-is-the-out-of-the-loop-performance-problem. Last reviewed 7 September 2026.
An evidence review by Rahim Hirji, not peer-reviewed research. For a material claim, cite the underlying study as well; every study here carries its own permanent link.
How citations and IDs work