← Research
Research

The use-or-keep test

Six questions about the piece of work in front of you, and what to do with the answers.

Last reviewed: 26 September 2026 · Next review due: 26 September 2027

The three terrains decide the ground in advance, for a class of work. These six questions decide the piece of work in front of you, once you are standing on it. Everything else published on this subject addresses somebody else: the regulations classify systems, human factors classifies functions at design time, and the management frameworks address the executive allocating work.

Questions this page answersAll 996 questions this research covers

Six questions about the piece of work in front of you, asked before you use AI on it: can I check it, what happens if it is wrong and I do not notice, is this a repetition I need, would doing it myself teach me anything, has the model seen this before, and who answers for it. It is the task-level companion to the three terrains: the terrains decide the ground in advance, these six questions decide the piece of work in front of you.

The answer, in one line

Six questions asked about one specific piece of work before using AI on it: can I check it, what happens if it is wrong and I do not notice, is this a repetition I need, would doing it myself teach me anything, has the model seen this before, and who answers for it.

Share as a card

Definition#

The use-or-keep test: six questions applied to one piece of work before using AI on it, covering verifiability, consequence and reversibility, the repetitions that maintain capability, the quality of the feedback the task returns, novelty relative to the model's training, and accountability for the result.

Share this definition as a card

Where this sits#

The parent of this test is the three terrains, from SuperSkills chapter seven, which maps where a machine should lead and where the call stays human: statistical terrain, bias terrain and override terrain. The terrains are decided in advance, for a class of work. This test is what you ask once you are standing on the ground with a specific piece of work in your hands, and it assumes the terrain map already exists.

The other published frameworks sort a different object again. Regulation sorts systems into risk tiers, which is a question for whoever procures the system. Human factors sorts functions across levels of automation, which is a question for an engineer at design time. The management frameworks sort organisational decisions, which is a question for the executive allocating them. Stanford's Human Agency Scale records what workers say they want automated, which is a measurement rather than an instruction. None of them addresses the person holding the task, which is where this test operates and why it is narrow on purpose.

The six questions#

1. Can I check it? If verifying the output would take longer than producing it yourself, you are not using AI, you are accepting it. This goes first because every question after it assumes you are in a position to evaluate what comes back. Two decades of work on appropriate reliance describe the same failure from the other end: people take answers they cannot assess, and no intervention has reliably stopped them.

2. What happens if it is wrong and I do not notice? Consequence and reversibility together. A wrong answer you catch is a cost in time. A wrong answer nobody catches is a different object, and the human factors literature on silent failure is largely about the second. Ask what recovery would involve and who would be doing it. The reversibility half of this question is what Amazon's one-way and two-way doors are for, and what the tiers in the EU AI Act are trying to express at the level of a system.

3. Is this a repetition I need? Some tasks are how a capability stays alive. Others were mastered years ago and will never be needed unaided again. Handing over the first kind is what this research calls custody, and this question is best asked before the tools arrive rather than after, because afterwards the answer tends to be whichever one is convenient.

4. Would doing it myself teach me anything? The one people skip, and the one with the most literature behind it. Hogarth's distinction between kind and wicked learning environments says that a task teaches you something only if it returns feedback that is quick, accurate and drawn from a complete sample. If the feedback from this task is delayed, partial or censored by your own choices, you were not learning much from it anyway, and the argument for keeping it is weaker than it feels. If the feedback is fast and honest, that is rare, and worth protecting.

5. Has the model seen this before? Novelty relative to the training distribution, in plain words. Confident wrong answers cluster where the situation is unusual, and the unusual part is often the context you hold and the model does not. If the thing that makes this case distinctive is the thing you would have to explain in a prompt, you are already doing the work that matters.

6. Who answers for it? If your name goes on the output, the checking has to match that rather than matching how finished the output looks. Fluency is not a signal of correctness and it reads as one. Where the accountability sits across a team, the question becomes which role holds it, and that is treated at how AI decision rights should be allocated.

How the answers combine#

They do not produce a score, and any version of this that produces a score is doing something the evidence does not support. The questions are ordered so that the first two settle most cases quickly. If you cannot check it and nobody would notice it was wrong, the answer is no whatever the other four say. If you can check it, the consequences are small and recoverable, and the task teaches you nothing, the answer is usually yes. The middle of the distribution is where the remaining questions do their work, and that middle is smaller than people expect once the first two have been asked honestly.

Where it sits against the published alternatives#

What this has not been shown to do#

It has not been tested. No study has measured whether asking these questions changes what people do, or whether what they do afterwards is better. The six criteria are drawn from separate literatures which have never been combined or tested together, so the evidence behind each one is not evidence for the set. It produces no score, no band and no benchmark, and a person using it consistently would still be relying on their own judgement about the answers, which is the thing under discussion. It is a structure for thinking, offered because the alternative on the open web is a list of questions with nothing behind it at all.

Key sources

Evidence review · SS-2026-320 · Graded against the published rubric

Cite this page

Hirji, R. (2026). The use-or-keep test. The SuperSkills evidence base, SS-2026-320. https://thesuperskills.com/research/the-use-or-keep-test. Last reviewed 26 September 2026.

An evidence review by Rahim Hirji, not peer-reviewed research. For a material claim, cite the underlying study as well; every study here carries its own permanent link.

How citations and IDs work
Questions answered on this page

What is the use-or-keep test?

Six questions asked about one specific piece of work before using AI on it: can I check it, what happens if it is wrong and I do not notice, is this a repetition I need, would doing it myself teach me anything, has the model seen this before, and who answers for it. It is the task-level companion to the three terrains, the framework from SuperSkills chapter seven that maps where a machine should lead and where the call stays human.

Is there an existing framework for deciding when to use AI?

There are several, and they address different people. The EU AI Act and the NIST framework classify systems by risk. Parasuraman, Sheridan and Wickens classify functions inside a designed system. MIT CISR and Shrestha and colleagues classify organisational decisions for the executive allocating them. Stanford's Human Agency Scale records what workers want automated. None of them is a set of questions for the person holding a single piece of work.

Why is verifiability the first question?

Because every other question assumes it. If you cannot check the output faster than you could have produced it, you have not saved the work, you have moved it out of sight. The appropriate reliance literature has spent two decades showing that people accept answers they are not in a position to evaluate.

Has the use-or-keep test been validated?

No. It has not been tested against outcomes, no study has measured whether using it improves any decision, and the criteria are drawn from separate literatures that have never been combined or tested together. It is a structure for thinking, not an instrument.

In this hub

Definitions

The terms this field uses, defined against their primary sources.

Ask the evidence
What does the evidence actually show?What should our board be asking about this?Where does Rahim disagree with the consensus?
Bring this into your organisation

If this describes something happening in your teams, say so.

Keynotes, board sessions and advisory work, drawing on research across more than 200 organisations in 30 countries. Tell me the room, the date and the shift you need. A reply within 24 hours.

Start a conversation

Topics and audiences  ·  All research

Box of Amazing

Rahim’s free weekly letter on AI and human capability

If this was useful, the weekly letter is where the thinking happens first. Most of what ends up on this site starts there. Weekly essays on AI, capability and the future of work. Read by 25,000 people, every week since 2017. Free, and one click to stop.

Opens Substack to confirm. No pitch in it, unsubscribe in one click, and nobody follows up because you read something.

Running an event, or responsible for how AI arrives in your organisation? Keynotes  ·  Advisory  ·  Boards  ·  Enquire