Almost every organisation measuring AI adoption is measuring activity. Licences issued, seats active, prompts per user, percentage of staff who have logged in, hours reportedly saved. Every one of those numbers can rise while nothing whatsoever changes about what the organisation is capable of doing.
Worse, one of the most common measures is actively misleading. Self-reported time saved is not a weak measure of productivity. In the one randomised trial that checked, it pointed in the wrong direction.
The finding that should end self-reported measurement
METR ran a randomised controlled trial with 16 experienced open-source developers across 246 real tasks on codebases they knew well. Each task was randomly assigned to permit or prohibit AI tools.
The developers forecast beforehand that AI would make them 24 per cent faster. They were measured as 19 per cent slower. And afterwards, having lived through the slowdown, they still estimated AI had sped them up by about 20 per cent.
A forty-point gap between belief and measurement, persisting after direct experience. The sample is small and specific, and it should not be generalised to all software work. But it is enough to retire one practice completely: asking people whether AI made them faster produces a number that may have the wrong sign. Most organisations are running exactly that survey and putting the result in a board pack.
Four things worth measuring instead
1 · Unaided capability, sampled. Can a person still complete a representative task to standard without the system? Test a sample, twice a year, on real work. This is the only measure that detects the thing everyone claims to be worried about, and virtually nobody runs it. It is also the only one that would have caught the clinical deskilling finding in advance.
2 · The pairing against the better half. Does human plus system beat the better of human alone and system alone? The Vaccaro meta-analysis found combinations underperforming the stronger party on average, with losses concentrated in review-style arrangements. If you have not measured this, you do not know whether your deployment is adding or subtracting.
3 · Outcome quality at the individual level. Not the average. Yu and colleagues found the effect of AI assistance on radiologists running from strongly positive to strongly negative between individuals, unpredicted by experience or prior familiarity. An average conceals the fact that you are helping some people and harming others, and it conceals which is which.
4 · Where the time went. If AI saved hours, name what those hours became. If nobody can, the time was absorbed rather than redeployed, which is the ordinary outcome and the reason productivity gains so often fail to appear anywhere a finance team can see them. See the unclaimed hour.
Two diagnostics that cost nothing
Count overrides. How many times has anyone disregarded the system this quarter? Zero is not evidence of a good system. It is evidence of an untested right, or of a workload that makes scrutiny impossible, or of a chain in which nobody is competent to object.
Ask who could detect an error. By name, per process. If the answer is a function rather than a person, or a person who could not have produced the work themselves, your verification is decorative. See who owns verification.
Where this is uncertain
There is no validated instrument for organisational AI capability, and this page does not pretend otherwise. The four measures are constructed from what the evidence supports rather than drawn from a tested framework, and no study has compared organisations that measure this way against those that do not.
Unaided capability testing also carries a real cost that should be stated: it takes people off productive work, it can feel like an exam, and done badly it will be resented. The honest position is that it is the only measure that answers the question, and that it has to be designed carefully to be worth running at all.
The SuperSkills view
Adoption metrics survive because they are easy to collect and flattering to report. Capability metrics are hard to collect and frequently unflattering, which is precisely why they are the ones worth having. An organisation that only measures usage has bought a dashboard that cannot detect its most serious risk, and will keep reporting green until something breaks.
This matters more since 2 August 2026. Article 14 of the EU AI Act requires that people overseeing high-risk systems can detect anomalies and disregard output. Those are capability claims. An organisation whose only evidence is a completion rate and a seat count cannot substantiate them, which turns a measurement habit into a compliance exposure.
Related SuperSkills research
On the failure mode, usage theatre and the AI readiness lie. On the board version, what should a board ask about AI. On what erodes, capability debt and deskilling. On the oversight duty, meaningful human oversight.
Key sources
- METR (2025). Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity.
- Vaccaro, M., Almaatouq, A. and Malone, T. (2024). When combinations of humans and AI are useful. Nature Human Behaviour, 8.
- Yu, F. et al. (2024). Heterogeneity and predictors of the effects of AI assistance on radiologists. Nature Medicine, 30(3).
- Budzyń, K. et al. (2025). Endoscopist deskilling risk after exposure to artificial intelligence in colonoscopy.
- Article 14, Human Oversight, Regulation (EU) 2024/1689.
About this research
Rahim Hirji is the author of SuperSkills (Kogan Page, 2026) and founder of The SuperSkills Intelligence Company. The four measures are a constructed framework rather than a validated instrument, which the page states. Findings are attributed to the studies that produced them. Not legal advice. Reviewed quarterly.
Cite this
Hirji, R. (2026). How do you measure AI adoption properly? The SuperSkills Intelligence Company. Last reviewed 26 August 2026. thesuperskills.com/research/how-do-you-measure-ai-adoption-properly