- What is the use-or-keep test?
- How do I decide whether to use AI on a piece of work?
- Is there a framework for when to use AI?
Six questions about the piece of work in front of you, asked before you use AI on it: can I check it, what happens if it is wrong and I do not notice, is this a repetition I need, would doing it myself teach me anything, has the model seen this before, and who answers for it. It is the task-level companion to the three terrains: the terrains decide the ground in advance, these six questions decide the piece of work in front of you.
The answer, in one line
Six questions asked about one specific piece of work before using AI on it: can I check it, what happens if it is wrong and I do not notice, is this a repetition I need, would doing it myself teach me anything, has the model seen this before, and who answers for it.
Definition#
The use-or-keep test: six questions applied to one piece of work before using AI on it, covering verifiability, consequence and reversibility, the repetitions that maintain capability, the quality of the feedback the task returns, novelty relative to the model's training, and accountability for the result.
Where this sits#
The parent of this test is the three terrains, from SuperSkills chapter seven, which maps where a machine should lead and where the call stays human: statistical terrain, bias terrain and override terrain. The terrains are decided in advance, for a class of work. This test is what you ask once you are standing on the ground with a specific piece of work in your hands, and it assumes the terrain map already exists.
The other published frameworks sort a different object again. Regulation sorts systems into risk tiers, which is a question for whoever procures the system. Human factors sorts functions across levels of automation, which is a question for an engineer at design time. The management frameworks sort organisational decisions, which is a question for the executive allocating them. Stanford's Human Agency Scale records what workers say they want automated, which is a measurement rather than an instruction. None of them addresses the person holding the task, which is where this test operates and why it is narrow on purpose.
The six questions#
1. Can I check it? If verifying the output would take longer than producing it yourself, you are not using AI, you are accepting it. This goes first because every question after it assumes you are in a position to evaluate what comes back. Two decades of work on appropriate reliance describe the same failure from the other end: people take answers they cannot assess, and no intervention has reliably stopped them.
2. What happens if it is wrong and I do not notice? Consequence and reversibility together. A wrong answer you catch is a cost in time. A wrong answer nobody catches is a different object, and the human factors literature on silent failure is largely about the second. Ask what recovery would involve and who would be doing it. The reversibility half of this question is what Amazon's one-way and two-way doors are for, and what the tiers in the EU AI Act are trying to express at the level of a system.
3. Is this a repetition I need? Some tasks are how a capability stays alive. Others were mastered years ago and will never be needed unaided again. Handing over the first kind is what this research calls custody, and this question is best asked before the tools arrive rather than after, because afterwards the answer tends to be whichever one is convenient.
4. Would doing it myself teach me anything? The one people skip, and the one with the most literature behind it. Hogarth's distinction between kind and wicked learning environments says that a task teaches you something only if it returns feedback that is quick, accurate and drawn from a complete sample. If the feedback from this task is delayed, partial or censored by your own choices, you were not learning much from it anyway, and the argument for keeping it is weaker than it feels. If the feedback is fast and honest, that is rare, and worth protecting.
5. Has the model seen this before? Novelty relative to the training distribution, in plain words. Confident wrong answers cluster where the situation is unusual, and the unusual part is often the context you hold and the model does not. If the thing that makes this case distinctive is the thing you would have to explain in a prompt, you are already doing the work that matters.
6. Who answers for it? If your name goes on the output, the checking has to match that rather than matching how finished the output looks. Fluency is not a signal of correctness and it reads as one. Where the accountability sits across a team, the question becomes which role holds it, and that is treated at how AI decision rights should be allocated.
How the answers combine#
They do not produce a score, and any version of this that produces a score is doing something the evidence does not support. The questions are ordered so that the first two settle most cases quickly. If you cannot check it and nobody would notice it was wrong, the answer is no whatever the other four say. If you can check it, the consequences are small and recoverable, and the task teaches you nothing, the answer is usually yes. The middle of the distribution is where the remaining questions do their work, and that middle is smaller than people expect once the first two have been asked honestly.
Where it sits against the published alternatives#
- The three terrains. The parent framework, from SuperSkills chapter seven. It sorts the ground rather than the task, and the sorting happens before the work starts. Where the terrain map and this test disagree, the terrain map wins, because it was made without the pressure of a deadline on it.
- Amazon's one-way and two-way doors. One criterion, reversibility, applied to decisions rather than to tasks. Question two is the same idea with consequence added.
- Parasuraman, Sheridan and Wickens, types and levels of automation. Richer than this test and aimed at a different reader. It asks how much of a function to automate at design time, and its evaluation criteria, including workload, situation awareness and skill degradation, are the ancestors of questions three and four. See levels of automation.
- Shrestha, Ben-Menahem and von Krogh. Five criteria for choosing a delegation structure at the level of an organisational decision. The most rigorous published criteria set, and addressed to the person allocating work rather than doing it.
- The MIT CISR decision matrix. Two axes, ambiguity and risk, sorting decisions into four types with a different human role in each.
- Stanford's Human Agency Scale. Five levels of desired human involvement, scored across 844 tasks. Descriptive: it records preference and capability, and does not say what should happen.
- Fitts's list. The ancestor of most two-by-two sorts of human work against machine work, abandoned in human factors for the reason Dekker and Woods set out in 2002: automation transforms the work that remains rather than removing a slice of it. Any framework that allocates tasks in advance by listing respective strengths inherits that fault, and this test avoids it by asking about one piece of work at a time rather than about categories of work.
What this has not been shown to do#
It has not been tested. No study has measured whether asking these questions changes what people do, or whether what they do afterwards is better. The six criteria are drawn from separate literatures which have never been combined or tested together, so the evidence behind each one is not evidence for the set. It produces no score, no band and no benchmark, and a person using it consistently would still be relying on their own judgement about the answers, which is the thing under discussion. It is a structure for thinking, offered because the alternative on the open web is a list of questions with nothing behind it at all.
Key sources
- Hogarth, R. M., Lejarraga, T. and Soyer, E. (2015). The two settings of kind and wicked learning environments. Current Directions in Psychological Science, 24(5).
- Shrestha, Y. R., Ben-Menahem, S. M. and von Krogh, G. (2019). Organizational decision-making structures in the age of artificial intelligence. California Management Review, 61(4).
- Shao, Y. et al. (2025). Future of work with AI agents: auditing automation and augmentation potential across the US workforce.
- Dekker, S. W. A. and Woods, D. D. (2002). MABA-MABA or abracadabra? Progress on human-automation co-ordination. Cognition, Technology and Work, 4(4).
- Sebastian, I. M., Weill, P., Haskamp, T. and vom Brocke, J. (2025). A framework for determining when AI can make decisions. MIT CISR.
Related SuperSkills research#
- The frameworks compared, which answers which question
- Keep, share, hand over
- What is custody
- How is judgement trained
Evidence review · SS-2026-320 · Graded against the published rubric
Hirji, R. (2026). The use-or-keep test. The SuperSkills evidence base, SS-2026-320. https://thesuperskills.com/research/the-use-or-keep-test. Last reviewed 26 September 2026.
An evidence review by Rahim Hirji, not peer-reviewed research. For a material claim, cite the underlying study as well; every study here carries its own permanent link.
How citations and IDs work