← Research
Research

How do you assess capability rather than output?

Once the machine can produce the artefact, the artefact stops telling you about the person who submitted it.

Last reviewed: 27 August 2026

What the evidence supports for assessing the person instead, starting with a correction to the number the whole field has been quoting since 1998.

Once a machine can produce the artefact, the artefact stops telling you about the person who submitted it. Every assessment system in education and hiring was built on the opposite assumption. This page sets out what the evidence supports for assessing the person instead, and it opens with a correction, because the number most often quoted in support of the obvious answer turns out to have been wrong for twenty-five years.

The number everyone quotes has been revised

Ask anyone in selection what predicts job performance and you will be told work sample tests, validity .54, from Schmidt and Hunter's 1998 review. It is one of the most cited findings in organisational psychology.

In 2022 Sackett, Zhang, Berry and Lievens showed that the corrections applied for range restriction in that generation of meta-analyses were systematically too large. Re-corrected, the picture moves substantially:

In the authors' own words, validity in general is lower than we had thought, and for some predictors the difference was quite substantial.

Two things follow, and they pull in opposite directions. Assessment is a weaker instrument than the field has been claiming, so anyone promising to measure capability precisely is overselling. And the ranking still holds: structured interviews and work samples remain among the best things available. The correction is to the confidence, not to the choice.

One thing the correction does not cover. It offers no validity data at all for the formats now being sold into this gap, including gamified assessment and asynchronous video interviews. Those are unmeasured, not proven.

Structure is what carries the weight

The gap between structured and unstructured interviews, .42 against .19, is the most actionable finding in the whole literature, and it has been stable across three decades of meta-analysis. McDaniel and colleagues reached the same conclusion in 1994 from 245 validity coefficients across 86,311 people.

Estimates of the size of the gap vary considerably between reviews, so treat the ratio rather than any single pair of numbers as the finding. The direction has never been in doubt.

What "structured" means in practice is narrow and unglamorous: the same questions in the same order, scored against defined anchors, by people who agreed the anchors beforehand. Most organisations believe they do this. Very few do.

Observing someone work, and what it costs

Medicine has spent thirty years building the assessment method everyone else now needs, which is watching a person do real work and scoring it. The honest finding is that it works and that it is expensive.

Moonen-van Loon and colleagues ran a generalisability study over 12,779 workplace-based assessments from 953 residents. To reach a reliability coefficient of 0.80 took eight mini-CEX observations, or nine DOPS, or nine multi-source feedback rounds. Combined into a portfolio the requirement fell to seven, eight and one respectively.

The number to carry out of that is eight, not 0.80. A single observation of someone at work is not a reliable assessment of anything, and almost every organisation that has moved to "we will just watch them do it" is running one observation and treating it as evidence.

Note also what the study measures. Reliability is consistency, not validity. It does not establish that these scores predict later clinical performance or patient outcomes.

Oral examination, and the problem it imports

If the written artefact no longer certifies the person, the obvious move is to ask them about it. Denmark has already mandated oral defence of home-written exams. The method is old and its weakness is well documented.

A 2025 study of 151 medical students compared structured, traditional and hybrid viva formats. Traditional unstructured viva showed significant inter-examiner variability. Structuring it improved fairness and coverage. The best-performing format reached a reliability of 0.663, which is moderate rather than high, and this is a single institution.

So oral assessment solves authorship and imports an examiner problem. It is defensible where it is structured and where more than one examiner is involved. It is not defensible as an unstructured conversation, which is how most of it is actually conducted, and the fairness cost falls unevenly on anyone assessed in a second language or unused to being questioned by authority.

What is not evidence, however often it is repeated

There is a widely circulated claim that human and AI writing leave distinguishable traces in keystroke logs or version history, human editing appearing as many small changes over hours and machine text arriving in clean blocks.

There is real research on reconstructing writing processes from keystroke data, including work using it to distinguish patterns of AI reliance in collaborative writing. What does not exist, as far as this research can establish, is a validated instrument for determining authorship from process data. The specific signature claim traces to secondary and promotional sources rather than to a peer-reviewed test with known error rates.

Anyone selling process forensics as proof of authorship is currently selling something that has not been validated, and the parallel to AI detection is close enough to matter: the false-positive burden of an unvalidated instrument does not fall evenly.

What this adds up to

Related SuperSkills research

On the education version, how to assess students when AI can do the assignment and does AI detection work. On measuring organisations rather than individuals, how to measure AI adoption properly. On what is being assessed, capability debt and tacit knowledge. On the profession that made periodic testing a condition of practice, what professions can learn from aviation. On whether lost capability comes back, can you regain a skill you have lost.

Key research and primary sources

About this research

Rahim Hirji is the author of SuperSkills (Kogan Page, 2026) and founder of The SuperSkills Intelligence Company. The validity figures quoted are the re-corrected estimates from Sackett and colleagues rather than the older and higher figures still in wide circulation. No validity evidence exists for AI-era assessment formats, which the page states rather than fills in. Not employment or academic-regulation advice. Reviewed quarterly.

Cite this

Hirji, R. (2026). How do you assess capability rather than output? The SuperSkills Intelligence Company. Last reviewed 27 August 2026. thesuperskills.com/research/how-do-you-assess-capability-rather-than-output

In this hub

AI and Human Judgement

Does AI weaken judgement? The evidence, and what to do about it.

The work

Where the writing comes from.

These essays draw on research across more than 200 organisations in 30 countries. See the wider body of work, or bring it into your organisation.

All research →