← Research
Research · Question

How do you hire for judgement?

The capability everybody now says is scarce, and almost nobody selects for.

Last reviewed: 26 September 2026 · Next review due: 26 September 2027

Organisations have spent two years concluding that judgement is the thing that matters and have changed nothing about how they choose people. The methods that work are known, unglamorous, and mostly already in the psychometric literature.

Question this page answersAll 996 questions this research covers

Not by asking about it. Organisations have spent two years deciding that judgement is the capability that matters and have changed almost nothing about how they select people. The methods that work are known, unglamorous, and mostly sitting in the psychometric literature already.

The answer, in one line

Not by asking about it. Asking a candidate to describe a difficult decision tests narrative skill and rewards the rehearsed answer.

Share as a card

Why the usual question fails#

Tell me about a difficult decision you made tests narrative skill. It rewards the rehearsed answer, it is answered best by people who have practised answering it, and the account offered has been reorganised around how things turned out, because hindsight rewrites the memory rather than suppressing it. A candidate describing a decision is describing a reconstruction, sincerely.

The take-home exercise has failed in a newer way. Scoring the output measured capability when producing a good output was expensive. It is now cheap, so the exercise measures tool access and the time available to iterate.

What the evidence supports#

The psychometric literature is old and unfashionable and it is what there is. A structured interview, meaning the same questions in the same order scored against a rubric written in advance, carries a validity coefficient around .42. A revised work sample sits around .33. Unstructured interviews, which is what most organisations run, perform considerably worse and feel considerably better to the interviewer.

The other finding worth carrying is about quantity. Reliable assessment of a behavioural capability generally needs something like eight observations. One panel, however well structured, gives a noisy result, and a promotion decided on a single panel is as variable as the evidence suggests.

A work sample for judgement#

Take a real case from the role, with information deliberately missing, and ask four things:

The output is the reasoning rather than the answer. The confidence figure and the falsifier are the two entries that remain expensive to fake, and they are the two that a well-prepared candidate with a good tool will not produce unprompted. Score against a rubric written before anybody sat the exercise, and have more than one person score it.

Let candidates use whatever tools they use at work, and ask what they changed in the output and why. A person who cannot say has told you something. The same four things on the table are what the shared prompt review asks of existing staff, which makes the selection test and the internal practice the same artefact.

Promoting for it#

Internally the problem is easier and mostly ignored, because the eight observations already exist. A person who has kept the four receipts has a record: Hand, Head, Hours, Heart, decisions with their name on them, misses found and changed, the reasoning written before the outcome was known. That is a body of evidence about judgement accumulated over a year, and almost no organisation collects it or looks at it when deciding who to promote.

The negative signal is equally available. A candidate for promotion who cannot point to a decision they owned, or a miss they found, or anything they changed their mind about, has a career of approvals rather than judgements.

Which type of judgement you are hiring for#

The three kinds at types of judgement need different tests. Predictive judgement can be scored directly, with forecasting questions that resolve, which is the only part of this that produces a number. Evaluative judgement shows up in whether the candidate names the trade-off or takes the framing they were handed. Moral judgement shows up in whether they identify who carries the cost and who is not in the room, which is the same question as the fourth receipt.

What this does not establish#

The validity figures come from the general psychometric literature on job performance, not from studies of judgement specifically, and applying them here is an inference. No instrument has been validated for the kind of work sample described above, and the work sample itself is a proposal rather than a tested method. What the evidence supports firmly is narrower and still useful: structure beats intuition, multiple observations beat one, and scoring the output is now close to worthless.

Essay · SS-2026-343

Cite this page

Hirji, R. (2026). How do you hire for judgement?. The SuperSkills evidence base, SS-2026-343. https://thesuperskills.com/research/how-do-you-hire-for-judgement. Last reviewed 26 September 2026.

An evidence review by Rahim Hirji, not peer-reviewed research. For a material claim, cite the underlying study as well; every study here carries its own permanent link.

How citations and IDs work
Questions answered on this page

Can you test for judgement in an interview?

Not by asking about it. Asking a candidate to describe a difficult decision tests narrative skill and rewards the rehearsed answer. What has evidence behind it is a structured interview with the same questions, the same order and a scoring rubric, which in the psychometric literature carries a validity coefficient around .42, against roughly .33 for a revised work sample.

What does a work sample for judgement look like?

A real case from the role, with information missing, where the candidate has to say what they would do, what would change their mind, and what they would need to know. The output is the reasoning rather than the answer, and it should be scored against a rubric written before anybody sat the exercise.

Why do most processes fail at this?

Because they score the output. AI has made a good output cheap to produce, so a take-home exercise now measures tool access rather than capability. The reasoning, the confidence attached to it, and the falsifier the candidate names are the parts that remain expensive to fake.

How many observations do you need?

More than most processes allow. Reliable assessment of a behavioural capability generally requires around eight observations, so a single interview, however structured, gives a noisy result and why promotion decisions based on one panel are as variable as they are.

In this hub

Definitions

The terms this field uses, defined against their primary sources.

Ask the evidence
What does the evidence actually show?What should our board be asking about this?Where does Rahim disagree with the consensus?
Bring this into your organisation

If this describes something happening in your teams, say so.

Keynotes, board sessions and advisory work, drawing on research across more than 200 organisations in 30 countries. Tell me the room, the date and the shift you need. A reply within 24 hours.

Start a conversation

Topics and audiences  ·  All research

Box of Amazing

Rahim’s free weekly letter on AI and human capability

If this was useful, the weekly letter is where the thinking happens first. Most of what ends up on this site starts there. Weekly essays on AI, capability and the future of work. Read by 25,000 people, every week since 2017. Free, and one click to stop.

Opens Substack to confirm. No pitch in it, unsubscribe in one click, and nobody follows up because you read something.

Running an event, or responsible for how AI arrives in your organisation? Keynotes  ·  Advisory  ·  Boards  ·  Enquire