Not by asking about it. Organisations have spent two years deciding that judgement is the capability that matters and have changed almost nothing about how they select people. The methods that work are known, unglamorous, and mostly sitting in the psychometric literature already.
The answer, in one line
Not by asking about it. Asking a candidate to describe a difficult decision tests narrative skill and rewards the rehearsed answer.
Why the usual question fails#
Tell me about a difficult decision you made tests narrative skill. It rewards the rehearsed answer, it is answered best by people who have practised answering it, and the account offered has been reorganised around how things turned out, because hindsight rewrites the memory rather than suppressing it. A candidate describing a decision is describing a reconstruction, sincerely.
The take-home exercise has failed in a newer way. Scoring the output measured capability when producing a good output was expensive. It is now cheap, so the exercise measures tool access and the time available to iterate.
What the evidence supports#
The psychometric literature is old and unfashionable and it is what there is. A structured interview, meaning the same questions in the same order scored against a rubric written in advance, carries a validity coefficient around .42. A revised work sample sits around .33. Unstructured interviews, which is what most organisations run, perform considerably worse and feel considerably better to the interviewer.
The other finding worth carrying is about quantity. Reliable assessment of a behavioural capability generally needs something like eight observations. One panel, however well structured, gives a noisy result, and a promotion decided on a single panel is as variable as the evidence suggests.
A work sample for judgement#
Take a real case from the role, with information deliberately missing, and ask four things:
- What would you do, and why that rather than the alternative?
- How confident are you, as a number?
- What would change your mind?
- What would you need to know, and how would you get it?
The output is the reasoning rather than the answer. The confidence figure and the falsifier are the two entries that remain expensive to fake, and they are the two that a well-prepared candidate with a good tool will not produce unprompted. Score against a rubric written before anybody sat the exercise, and have more than one person score it.
Let candidates use whatever tools they use at work, and ask what they changed in the output and why. A person who cannot say has told you something. The same four things on the table are what the shared prompt review asks of existing staff, which makes the selection test and the internal practice the same artefact.
Promoting for it#
Internally the problem is easier and mostly ignored, because the eight observations already exist. A person who has kept the four receipts has a record: Hand, Head, Hours, Heart, decisions with their name on them, misses found and changed, the reasoning written before the outcome was known. That is a body of evidence about judgement accumulated over a year, and almost no organisation collects it or looks at it when deciding who to promote.
The negative signal is equally available. A candidate for promotion who cannot point to a decision they owned, or a miss they found, or anything they changed their mind about, has a career of approvals rather than judgements.
Which type of judgement you are hiring for#
The three kinds at types of judgement need different tests. Predictive judgement can be scored directly, with forecasting questions that resolve, which is the only part of this that produces a number. Evaluative judgement shows up in whether the candidate names the trade-off or takes the framing they were handed. Moral judgement shows up in whether they identify who carries the cost and who is not in the room, which is the same question as the fourth receipt.
What this does not establish#
The validity figures come from the general psychometric literature on job performance, not from studies of judgement specifically, and applying them here is an inference. No instrument has been validated for the kind of work sample described above, and the work sample itself is a proposal rather than a tested method. What the evidence supports firmly is narrower and still useful: structure beats intuition, multiple observations beat one, and scoring the output is now close to worthless.
Related SuperSkills research#
- How do you assess capability rather than output
- Types of judgement
- Hand, Head, Hours, Heart
- How do you build judgement in an organisation
Essay · SS-2026-343
Hirji, R. (2026). How do you hire for judgement?. The SuperSkills evidence base, SS-2026-343. https://thesuperskills.com/research/how-do-you-hire-for-judgement. Last reviewed 26 September 2026.
An evidence review by Rahim Hirji, not peer-reviewed research. For a material claim, cite the underlying study as well; every study here carries its own permanent link.
How citations and IDs work