← Research
Research

How do you assess capability rather than output?

Once the machine can produce the artefact, the artefact stops telling you about the person who submitted it.

Last reviewed: 27 August 2026

What the evidence supports for assessing the person instead, starting with a correction to the number the whole field has been quoting since 1998.

Questions this page answersAll 616 questions this research covers

Once a machine can produce the artefact, the artefact stops telling you about the person who submitted it. Every assessment system in education and hiring was built on the opposite assumption. This page sets out what the evidence supports for assessing the person instead, and it opens with a correction, because the number most often quoted in support of the obvious answer turns out to have been wrong for twenty-five years.

The number everyone quotes has been revised#

Ask anyone in selection what predicts job performance and you will be told work sample tests, validity .54, from Schmidt and Hunter's 1998 review. It is one of the most cited findings in organisational psychology.

In 2022 Sackett, Zhang, Berry and Lievens showed that the corrections applied for range restriction in that generation of meta-analyses were systematically too large. Re-corrected, the picture moves substantially:

In the authors' own words, validity in general is lower than we had thought, and for some predictors the difference was quite substantial.

Two things follow, and they pull in opposite directions. Assessment is a weaker instrument than the field has been claiming, so anyone promising to measure capability precisely is overselling. And the ranking still holds: structured interviews and work samples remain among the best things available. The correction is to the confidence, not to the choice.

One thing the correction does not cover. It offers no validity data at all for the formats now being sold into this gap, including gamified assessment and asynchronous video interviews. Those are unmeasured, not proven.

Structure is what carries the weight#

The gap between structured and unstructured interviews, .42 against .19, is the most actionable finding in the whole literature, and it has been stable across three decades of meta-analysis. McDaniel and colleagues reached the same conclusion in 1994 from 245 validity coefficients across 86,311 people.

Estimates of the size of the gap vary considerably between reviews, so treat the ratio rather than any single pair of numbers as the finding. The direction has never been in doubt.

What "structured" means in practice is narrow and unglamorous: the same questions in the same order, scored against defined anchors, by people who agreed the anchors beforehand. Most organisations believe they do this. Very few do.

Observing someone work, and what it costs#

Medicine has spent thirty years building the assessment method everyone else now needs, which is watching a person do real work and scoring it. The honest finding is that it works and that it is expensive.

Moonen-van Loon and colleagues ran a generalisability study over 12,779 workplace-based assessments from 953 residents. To reach a reliability coefficient of 0.80 took eight mini-CEX observations, or nine DOPS, or nine multi-source feedback rounds. Combined into a portfolio the requirement fell to seven, eight and one respectively.

The number to carry out of that is eight, not 0.80. A single observation of someone at work is not a reliable assessment of anything, and almost every organisation that has moved to "we will just watch them do it" is running one observation and treating it as evidence.

Note also what the study measures. Reliability is consistency, not validity. It does not establish that these scores predict later clinical performance or patient outcomes.

Oral examination, and the problem it imports#

If the written artefact no longer certifies the person, the obvious move is to ask them about it. Denmark has already mandated oral defence of home-written exams. The method is old and its weakness is well documented.

A 2025 study of 151 medical students compared structured, traditional and hybrid viva formats. Traditional unstructured viva showed significant inter-examiner variability. Structuring it improved fairness and coverage. The best-performing format reached a reliability of 0.663, which is moderate rather than high, and this is a single institution.

So oral assessment solves authorship and imports an examiner problem. It is defensible where it is structured and where more than one examiner is involved. It is not defensible as an unstructured conversation, which is how most of it is actually conducted, and the fairness cost falls unevenly on anyone assessed in a second language or unused to being questioned by authority.

What is not evidence, however often it is repeated#

There is a widely circulated claim that human and AI writing leave distinguishable traces in keystroke logs or version history, human editing appearing as many small changes over hours and machine text arriving in clean blocks.

There is real research on reconstructing writing processes from keystroke data, including work using it to distinguish patterns of AI reliance in collaborative writing. What does not exist, as far as this research can establish, is a validated instrument for determining authorship from process data. The specific signature claim traces to secondary and promotional sources rather than to a peer-reviewed test with known error rates.

Anyone selling process forensics as proof of authorship is currently selling something that has not been validated, and the parallel to AI detection is close enough to matter: the false-positive burden of an unvalidated instrument does not fall evenly.

What this adds up to#

On the education version, how to assess students when AI can do the assignment and does AI detection work. On measuring organisations rather than individuals, how to measure AI adoption properly. On what is being assessed, capability debt and tacit knowledge. On the profession that made periodic testing a condition of practice, what professions can learn from aviation. On whether lost capability comes back, can you regain a skill you have lost.

Key research and primary sources

About this research#

Rahim Hirji is the author of SuperSkills (Kogan Page, 2026), keynote speaker on AI and human capability, and founder of The SuperSkills Intelligence Company. The validity figures quoted are the re-corrected estimates from Sackett and colleagues rather than the older and higher figures still in wide circulation. No validity evidence exists for AI-era assessment formats, which the page states rather than fills in. Not employment or academic-regulation advice.

How this research works  ·  Reviewed quarterly  ·  Found an error? Tell me and it is corrected on the page.

Cite this

Hirji, R. (2026). How do you assess capability rather than output? The SuperSkills Intelligence Company. Last reviewed 27 August 2026. thesuperskills.com/research/how-do-you-assess-capability-rather-than-output

Questions answered on this page

How do you assess capability rather than output?

By structuring what you already do, and by accepting that it takes repetition. Structured interviews are now the strongest single predictor of job performance at a corrected validity of .42, against .19 for unstructured. Workplace observation is defensible but expensive: reaching a reliability of 0.80 took eight mini-CEX observations in a study of 12,779 assessments from 953 medical residents. A single observation is not a measurement. And performance with the tool must be measured separately from capability without it, because they are different quantities.

Is the work sample test validity of .54 still correct?

No. That figure comes from Schmidt and Hunter's 1998 review and has been widely quoted since. Sackett, Zhang, Berry and Lievens showed that the range restriction corrections applied in that generation of meta-analyses were systematically too large. Re-corrected, work sample validity falls from .54 to .33, structured interviews from .51 to .42, unstructured interviews from .38 to .19, and general cognitive ability from .51 to .31. The ranking largely holds but the confidence does not.

Can you tell whether work was written by AI from keystroke or version history?

There is no validated instrument for this. Research exists on reconstructing writing processes from keystroke data, including work distinguishing patterns of AI reliance in collaborative writing, but the specific claim that human and machine writing leave reliably distinguishable signatures traces to secondary and promotional sources rather than to a peer-reviewed test with known error rates. The parallel with AI text detection is close: an unvalidated instrument produces false positives, and they do not fall evenly.

Does oral examination solve the AI assessment problem?

It solves authorship and imports an examiner problem. A 2025 study of 151 medical students found traditional unstructured viva showed significant inter-examiner variability. Structuring improved fairness and coverage, and the best format reached a reliability of 0.663, which is moderate rather than high, at a single institution. Oral assessment is defensible where it is structured and uses more than one examiner. As an unstructured conversation, which is how most of it is run, it is not.

In this hub

Thinking, learning and capability

What sustained AI use does to thinking, and how capability is built and kept.

The work

Where the writing comes from.

These essays draw on research across more than 200 organisations in 30 countries. See the wider body of work, or bring it into your organisation.

All research →

Separating the work from the worker's capability is the harder half, and a performance process that already does it has solved most of this. Otherwise, building an assessment of what people can still do unaided is the engagement. Advisory and coaching.

This is the argument HR audiences push back on hardest, which is why it works on stage. There is the HR and CHRO version, and the full range of topics and audiences.

Box of Amazing

Rahim’s free weekly letter on AI and human capability

If this was useful, the weekly letter is where the thinking happens first. Most of what ends up on this site starts there. Weekly essays on AI, capability and the future of work. Read by 25,000 people, every week since 2017. Free, and one click to stop.

Opens Substack to confirm. No pitch in it, unsubscribe in one click, and nobody follows up because you read something.