Appropriate reliance means taking the machine's answer when it is right and overriding it when it is wrong, and after two decades of research nobody has an intervention that reliably produces it. That absence is the most important thing a leadership team can know about human oversight, because every oversight policy in circulation assumes the opposite.
The answer, in one line
Accepting automated advice when it is correct and rejecting it when it is not. Lee and See drew the distinction in 2004: trust is the attitude, reliance is the behaviour, and calibration is the match between them and the system's actual reliability.
Definition#
Appropriate reliance: the behaviour of accepting automated advice when it is correct and rejecting it when it is not. Trust is the attitude, reliance is the behaviour, and calibration is the match between the two and the system's actual reliability. The distinction is Lee and See's, from 2004.
Why it is hard#
To rely appropriately, a person has to judge whether this particular output is right, which usually requires the capability the tool was brought in to supply. That is the circularity at the centre of human oversight, and it does not dissolve with training or with better interfaces.
The reliability of the system works against the human too. A tool that is right nine times in ten teaches the operator to accept the tenth, and the acceptance is rational given the base rate. This is the mechanism behind automation bias and complacency, reviewed by Parasuraman and Manzey in 2010, and it appears in novices and experts alike.
What has been tried, and what it produced#
Explanations. The intuitive fix, and the one regulation tends to assume. Bansal and colleagues, in a controlled study at CHI in 2021, found that explanations increased the rate at which people accepted the model's answer without improving the accuracy of the team, which is the definition of the wrong outcome.
Confidence scores. They help where the confidence is calibrated, and models are frequently confident in the wrong places, which moves the problem rather than solving it. The comparison this estate keeps is at calibration.
Cognitive forcing. Making the person commit to an answer before seeing the machine's. The most promising family, and the results remain mixed; it also costs the time saving that justified the tool.
Training. Warnings and instruction reduce automation bias somewhat and do not remove it, and the effect decays as the system proves reliable again.
What follows for oversight policy#
A policy that says a human will review the output is not a control until three things are true: the reviewer could produce or check the answer unaided, the review has time and authority to change the outcome, and somebody counts how often it does. The third is the one nobody instruments, and a review that never disagrees is not a review.
That is why this research treats override rate as the number worth watching. It is cheap to collect, hard to fake and it distinguishes an oversight arrangement that operates from one that exists on paper. The regulatory version of the same requirement is discussed at human in the loop is not a safeguard.
Key sources
- Lee, J. D. and See, K. A. (2004). Trust in automation: designing for appropriate reliance. Human Factors, 46(1), 50-80.
- Bansal, G. et al. (2021). Does the whole exceed its parts? The effect of AI explanations on complementary team performance. CHI 2021.
- Parasuraman, R. and Manzey, D. H. (2010). Complacency and bias in human use of automation. Human Factors, 52(3).
Explainer · SS-2026-315 · Graded against the published rubric
Hirji, R. (2026). What is appropriate reliance?. The SuperSkills evidence base, SS-2026-315. https://thesuperskills.com/research/what-is-appropriate-reliance. Last reviewed 26 September 2026.
An evidence review by Rahim Hirji, not peer-reviewed research. For a material claim, cite the underlying study as well; every study here carries its own permanent link.
How citations and IDs work