← Research
Research

What is the out-of-the-loop performance problem?

Endsley and Kiris measured people who were monitoring correctly and still could not take over. Comprehension went, perception stayed.

Last reviewed: 7 September 2026

Where the term was named, the five levels of control it was measured across, which half of situation awareness the automation damaged, and why an approval step is the wrong remedy for it.

Questions this page answersAll 811 questions this research covers

The out-of-the-loop performance problem is what happens to a person's ability to take over when the machine they were supervising stops working. They are slower to notice that something has gone wrong, and slower again to establish what it is once they have noticed. Mica Endsley and Esin Kiris named it in Human Factors in 1995 and measured one of its causes. The result that matters most to anyone supervising AI is which half of understanding survived: their participants still saw the information in front of them, and had lost the grasp of what it meant.

The answer, in one line

It is the loss of a person's ability to take over manual operation when an automated system fails, caused by their having been placed in the role of monitor instead of operator.

Share as a card

Definition#

The out-of-the-loop performance problem: the loss of a person's ability to take over manual operation when an automated system fails, caused by their having been placed in the role of monitor instead of operator. Named by Mica Endsley and Esin Kiris in 1995.

Share this definition as a card

The navigation task, and the five levels it was run at#

Endsley and Kiris automated an automobile navigation task using an expert system, and ran it at five separate levels of operator control: manual throughout; the system suggests and the person decides and acts; the system decides and acts with the person's consent required; the system decides and acts unless the person vetoes; and full automation with no operator interaction. Those five levels are Endsley's own scale, set out in her work on expert systems in cockpits in the late 1980s.

Situation awareness was lower under the automated conditions than under manual performance, and decision time after the expert system failed was longer where situation awareness was lower. The out-of-the-loop effect was significantly greater under full automation than under the intermediate levels. Keeping the person inside the decision, at a level below the top of the scale, preserved both their awareness and their ability to resume control.

They still saw the data and no longer knew what it meant#

Endsley's own account of the study, written up in a chapter for Parasuraman and Mouloua's Automation and Human Performance in 1996, reports the split that makes this finding useful outside aviation. Situation awareness in her framework has three levels: perceiving the elements in front of you, comprehending what they mean in relation to your goals, and projecting what the system will do next. Only the second was damaged. Participants remained aware of the low-level data, so they were monitoring the system perfectly well, and had "less comprehension of what the data meant" for the task they were there to accomplish.

Endsley attributes that specifically to passivity. Under the conditions of the experiment the information displayed to operators did not change between conditions, and vigilance and monitoring effects were too small to account for the decrement. Turning a performer into an observer damaged comprehension on its own, in people who were watching attentively and seeing everything they were shown.

Three routes out of the loop#

Endsley sets out three mechanisms by which automation carries a person out of the loop, and they remain the useful decomposition.

Monitoring. People are poor sustained monitors of reliable systems, which is the vigilance and complacency literature the estate covers under automation complacency. Attention drifts towards whatever else is competing for it, and a system that has never failed offers no reason to look.

Passive processing. Observing a decision is a different cognitive act from making one, and it lays down a weaker model of the situation. This is the route Endsley's own experiment implicates, and the one that survives any amount of diligence.

Feedback. Automating a function tends to remove the information that used to arrive as a by-product of doing it. Designers assume the operator no longer needs what the machine now handles, so the cues that would have supported a takeover are gone at the moment they are wanted.

What a 1995 navigation study cannot settle#

This is one laboratory task, run on an expert system that produced route recommendations, thirty years before the tools this site is otherwise about. A generative model differs from that system in the way that matters here: it produces fluent output across every task rather than one recommendation on a defined one, so the operator's comprehension problem may be larger or smaller and nobody has measured it.

Two further limits belong on the record. The paper's sample size is not given in its abstract and the full text sits behind a publisher paywall, so no participant count appears on this page. And the Level 1 against Level 2 split, which is the most quoted thing here, is read from Endsley's 1996 chapter summarising her own study rather than from the results section of the paper itself. Both are stated because the alternative is to imply a closer reading than was possible.

Comprehension is what the arrangement is quietly delegating#

Almost every oversight arrangement in commercial use is written as though the risk were that the person fails to look. Sign-offs, review steps and second pairs of eyes all address attention. Endsley and Kiris measured people who looked, who were correctly monitoring, and who could not act well when the system failed anyway. If that generalises, then the standard remedy treats the symptom the arrangement does not have.

This is the specific mechanism underneath the human-in-the-loop claim, and it gives Article 14 of the EU AI Act a harder edge than the drafting suggests. Article 14 requires that an overseer be enabled to understand the system's capacities and limitations and to intervene. The out-of-the-loop result says the capacity to intervene falls as the level of automation rises, in people who understand the system and are paying attention. No amount of training or instruction reaches it, because the cause is where the person sits in the decision.

The design lever Endsley identified is the level of control. That is the same lever as drift versus design, seen from the human-factors side. Nobody in an organisation chooses level five. Products arrive configured at the top of the scale, approval flows are added to make them feel supervised, and the arrangement that results asks a person to consent to decisions they did not participate in making. Consent and veto both leave the operator passive, which is the condition the experiment was measuring.

Designing the level down#

Name the level for every automated decision. Take the five-point scale above and place each one on it. Most organisations have never asked the question, and the answer is usually four or five for work nobody would have chosen to fully automate.

Give the person something to decide, not something to approve. A step that only permits yes or no keeps them out of the loop while producing a governance record that says otherwise. This is the difference between the second and the fourth level on the scale.

Test the takeover, not the output. Accuracy under normal running says nothing about the state this literature is concerned with. The measure is how long it takes someone to notice a failure and correct it, which almost nobody records and which is the number that would have shown the problem in advance.

Keep the feedback that automation makes redundant. Information that only mattered because a person was doing the work is exactly the information they will need in order to resume it.

On the monitoring route, automation complacency and automation bias. On the regulatory test, meaningful human oversight and why human in the loop is not a safeguard. On who is left holding the failure, the moral crumple zone. On the capability question underneath it, who supervises work they cannot do and what is deskilling.

Key sources

About this research#

Rahim Hirji is the author of SuperSkills (Kogan Page, 2026), keynote speaker on AI and human capability, and founder of The SuperSkills Intelligence Company. The out-of-the-loop performance problem is an established human-factors concept and no SuperSkills term. Findings are attributed to the studies that produced them and kept separate from the interpretation, which is the author's. The graded evidence, including what each study does not support, is in the evidence base. Last reviewed: 7 September 2026.

How this research works  ·  Reviewed quarterly  ·  Found an error? Tell me and it is corrected on the page.

Explainer · SS-2026-203 · Graded against the published rubric

Cite this page

Hirji, R. (2026). What is the out-of-the-loop performance problem?. The SuperSkills evidence base, SS-2026-203. https://thesuperskills.com/research/what-is-the-out-of-the-loop-performance-problem. Last reviewed 7 September 2026.

An evidence review by Rahim Hirji, not peer-reviewed research. For a material claim, cite the underlying study as well; every study here carries its own permanent link.

How citations and IDs work
Questions answered on this page

What is the out-of-the-loop performance problem?

It is the loss of a person's ability to take over manual operation when an automated system fails, caused by their having been placed in the role of monitor instead of operator. Mica Endsley and Esin Kiris named it in Human Factors in 1995. Someone in this state is slower to detect that a problem has occurred, and slower again to establish what it is once they have detected it. Endsley attributes it to three things: monitoring and vigilance problems, a shift from active to passive information processing, and changes in the feedback that reaches the operator.

What did Endsley and Kiris actually find?

They automated an automobile navigation task using an expert system, at five levels of operator control from fully manual to fully automatic. Situation awareness was lower under the automated conditions than under manual performance, and decision time after the expert system failed was longer where situation awareness was lower. The out-of-the-loop effect was significantly greater under full automation than under the intermediate levels, so keeping the operator inside the decision moderated the loss.

Does paying more attention fix the out-of-the-loop problem?

The evidence says not, and that is the most useful part of it. In Endsley's account of her own study, only the second level of situation awareness was damaged. Participants still perceived the low-level data, so they were monitoring the system effectively, and had less comprehension of what the data meant for the goals they were pursuing. The information displayed did not change between conditions and monitoring effects were too small to explain the decrement, so the cause was the shift from performing to observing rather than any failure of attention.

What are the five levels of automation on Endsley's scale?

One, manual, where the person decides and acts with no help. Two, decision support, where the system suggests and the person decides and acts. Three, consensual, where the system decides and acts but requires the person's consent. Four, monitored, where the system decides and acts unless the person vetoes. Five, full automation, with no operator interaction. The out-of-the-loop problem was significantly greater at level five than at the intermediate levels.

Does this apply to generative AI?

Not directly, and this page says so. The study is a single laboratory task from 1995, run on an expert system that produced route recommendations. A generative model produces fluent output across every task instead of one recommendation on a defined task, so the comprehension effect could be larger or smaller and nobody has measured it. What transfers is the mechanism and the design lever: the damage tracked the level of automation, and lowering that level protected the operator's ability to resume control.

In this hub

Definitions

The terms this field uses, defined against their primary sources.

Ask the evidence
What does the evidence actually show?What should our board be asking about this?Where does Rahim disagree with the consensus?
Bring this into your organisation

If this describes something happening in your teams, say so.

Keynotes, board sessions and advisory work, drawing on research across more than 200 organisations in 30 countries. Tell me the room, the date and the shift you need. A reply within 24 hours.

Start a conversation

Topics and audiences  ·  All research

Box of Amazing

Rahim’s free weekly letter on AI and human capability

If this was useful, the weekly letter is where the thinking happens first. Most of what ends up on this site starts there. Weekly essays on AI, capability and the future of work. Read by 25,000 people, every week since 2017. Free, and one click to stop.

Opens Substack to confirm. No pitch in it, unsubscribe in one click, and nobody follows up because you read something.

Running an event, or responsible for how AI arrives in your organisation? Keynotes  ·  Advisory for CEOs and boards  ·  Enquire