← Research
Research

How do you measure AI adoption properly?

Licences, seats and prompts can all rise while nothing changes about what the organisation can do.

Last reviewed: 26 August 2026

The finding that should retire self-reported productivity measurement, four things worth measuring instead, two diagnostics that cost nothing, and why this became a compliance question in August 2026.

Questions this page answersAll 616 questions this research covers

Almost every organisation measuring AI adoption is measuring activity. Licences issued, seats active, prompts per user, percentage of staff who have logged in, hours reportedly saved. Every one of those numbers can rise while nothing whatsoever changes about what the organisation is capable of doing.

Worse, one of the most common measures is actively misleading. Self-reported time saved is not a weak measure of productivity. In the one randomised trial that checked, it pointed in the wrong direction.

The finding that should end self-reported measurement#

METR ran a randomised controlled trial with 16 experienced open-source developers across 246 real tasks on codebases they knew well. Each task was randomly assigned to permit or prohibit AI tools.

The developers forecast beforehand that AI would make them 24 per cent faster. They were measured as 19 per cent slower. And afterwards, having lived through the slowdown, they still estimated AI had sped them up by about 20 per cent.

A forty-point gap between belief and measurement, persisting after direct experience. The sample is small and specific, and it should not be generalised to all software work. But it is enough to retire one practice completely: asking people whether AI made them faster produces a number that may have the wrong sign. Most organisations are running exactly that survey and putting the result in a board pack.

Four things worth measuring instead#

1 · Unaided capability, sampled. Can a person still complete a representative task to standard without the system? Test a sample, twice a year, on real work. This is the only measure that detects the thing everyone claims to be worried about, and virtually nobody runs it. It is also the only one that would have caught the clinical deskilling finding in advance.

2 · The pairing against the better half. Does human plus system beat the better of human alone and system alone? The Vaccaro meta-analysis found combinations underperforming the stronger party on average, with losses concentrated in review-style arrangements. If you have not measured this, you do not know whether your deployment is adding or subtracting.

3 · Outcome quality at the individual level. Not the average. Yu and colleagues found the effect of AI assistance on radiologists running from strongly positive to strongly negative between individuals, unpredicted by experience or prior familiarity. An average conceals the fact that you are helping some people and harming others, and it conceals which is which.

4 · Where the time went. If AI saved hours, name what those hours became. If nobody can, the time was absorbed rather than redeployed, which is the ordinary outcome and the reason productivity gains so often fail to appear anywhere a finance team can see them. See the unclaimed hour.

Two diagnostics that cost nothing#

Count overrides. How many times has anyone disregarded the system this quarter? Zero evidences an untested right rather than a good system, or a workload that makes scrutiny impossible, or of a chain in which nobody is competent to object.

Ask who could detect an error. By name, per process. If the answer is a function rather than a person, or a person who could not have produced the work themselves, your verification is decorative. See who owns verification.

The limits of these four measures#

There is no validated instrument for organisational AI capability, and this page does not pretend otherwise. The four measures are constructed from what the evidence supports rather than drawn from a tested framework, and no study has compared organisations that measure this way against those that do not.

Unaided capability testing also carries a real cost that should be stated: it takes people off productive work, it can feel like an exam, and done badly it will be resented. It remains the only measure that answers the question, and it has to be designed carefully to be worth running at all.

Why the flattering metric survives#

Adoption metrics survive because they are easy to collect and flattering to report. Capability metrics are hard to collect and frequently unflattering. Those are the ones worth having. An organisation that only measures usage has bought a dashboard that cannot detect its most serious risk, and will keep reporting green until something breaks.

This matters more since 2 August 2026. Article 14 of the EU AI Act requires that people overseeing high-risk systems can detect anomalies and disregard output. Those are capability claims. An organisation whose only evidence is a completion rate and a seat count cannot substantiate them, which turns a measurement habit into a compliance exposure.

On the failure mode, usage theatre and the AI readiness lie. On the board version, what should a board ask about AI. On what erodes, capability debt and deskilling. On the oversight duty, meaningful human oversight.

Key sources

About this research#

Rahim Hirji is the author of SuperSkills (Kogan Page, 2026), keynote speaker on AI and human capability, and founder of The SuperSkills Intelligence Company. The four measures are a constructed framework rather than a validated instrument, which the page states. Findings are attributed to the studies that produced them. Not legal advice.

How this research works  ·  Reviewed quarterly  ·  Found an error? Tell me and it is corrected on the page.

Cite this

Hirji, R. (2026). How do you measure AI adoption properly? The SuperSkills Intelligence Company. Last reviewed 26 August 2026. thesuperskills.com/research/how-do-you-measure-ai-adoption-properly

Questions answered on this page

How should you measure AI adoption?

Not by activity. Licences issued, seats active, prompts per user and hours reportedly saved can all rise while nothing changes about organisational capability. Four measures are worth having instead: unaided capability sampled twice a year on real work, whether the human and system pairing beats the better of either alone, outcome quality at individual rather than average level, and a named account of what the saved time actually became.

Why is self-reported time saved a bad measure?

Because in the one randomised trial that checked it, the number had the wrong sign. METR found 16 experienced developers across 246 tasks were measured 19 per cent slower with AI tools, having forecast a 24 per cent speed-up, and still estimated afterwards that AI had made them about 20 per cent faster. A forty-point gap between belief and measurement, persisting after direct experience. The sample is small and specific, but it is enough to retire the practice of asking people whether AI made them faster.

Why does the average conceal the problem?

Because effects differ in direction between individuals. Yu and colleagues found the impact of AI assistance on radiologists ran from strongly positive to strongly negative between readers, and was not predicted by experience or prior familiarity with AI. An average conceals the fact that you are helping some people and harming others, and conceals which is which.

What are the cheapest diagnostics?

Two. Count overrides: if nobody has disregarded the system this quarter, that is evidence of an untested right rather than a good system. And ask who could detect an error, by name and per process: if the answer is a function rather than a person, or a person who could not have produced the work themselves, the verification is decorative.

In this hub

Organisations and leadership

What a leadership team actually has to decide, and what to measure.

The work

Where the writing comes from.

These essays draw on research across more than 200 organisations in 30 countries. See the wider body of work, or bring it into your organisation.

All research →

If your dashboard already answers whether anyone got better at anything, keep it. If it counts logins, I build the other kind. Advisory and coaching.

This is the argument HR audiences push back on hardest, which is why it works on stage. There is the HR and CHRO version, and the full range of topics and audiences.

Box of Amazing

Rahim’s free weekly letter on AI and human capability

If this was useful, the weekly letter is where the thinking happens first. Most of what ends up on this site starts there. Weekly essays on AI, capability and the future of work. Read by 25,000 people, every week since 2017. Free, and one click to stop.

Opens Substack to confirm. No pitch in it, unsubscribe in one click, and nobody follows up because you read something.