- What happens when people stop practising?
- What is capability debt?
- What is missing from most AI strategies?
- How do you avoid capability debt?
- How do I know if my organisation is already in capability debt?
- What is capability masking?
- Can capability loss from AI actually be measured?
- Does automation always cause deskilling?
- Can an organisation improve AI productivity while losing the capability it needs to survive without AI?
- How would you know whether AI caused the capability loss?
Capability debt is what an organisation owes its own future when it automates the doing without redesigning the learning. Every time AI takes over a task people used to perform, the work still ships, but the practice that built the underlying judgement stops. Capability falls behind the tools, invisibly, because the outputs still look fine. Nobody sees a problem on any dashboard. The bill arrives later and all at once, when a decision turns up that the AI cannot make and no human in the room has been trained to make either. Like financial debt, it is borrowing against the future. Unlike financial debt, most organisations do not know they are taking it on.
The strongest evidence for it now comes from medicine. Nineteen endoscopists with an average of 27.6 years of experience each got worse at finding tumours without the machine, within months of routine exposure to a tool that had been helping them. Their unassisted detection rate fell by six percentage points. That is the shape of the whole argument, measured in people whose degradation has consequences.
Borrowed from the code, and worse in people#
The term borrows deliberately from technical debt, which Ward Cunningham coined in 1992 to explain why hastily shipped software would need rewriting. It is fine to borrow against the future, he argued, as long as you pay it back, and if you do not, you pay interest in the form of everything the shortcut makes harder later.
Moving the metaphor from the code to the people changes how the debt behaves, in three ways that all make it worse. Technical debt is visible to the engineers carrying it, and capability debt hides inside capable-looking output, so the people accruing it often cannot see it. Technical debt is repayable by refactoring, and capability debt takes years of the very practice that was automated away. Technical debt sits in an artefact the organisation owns, and capability debt sits in people who can resign.
Two people, one metaphor, two scopes#
Wolfgang Rohde uses capability debt in Short-Term Gain, Long-Term Fragility: AI Labor Substitution and the Erosion of Sustainable Capability (SSRN, written 20 April 2026, revised 27 April 2026; also arXiv 2605.27399), arrived at independently of this research. No claim of first use is made here. Checked at source on 28 August 2026: Rohde does not claim to have coined it either. A term is credited to Rahim Hirji on this site only where a dated first publication exists, and for this one it does not. The likeliest explanation is the dull one, that two people reached for the same metaphor because it is the obvious metaphor. The wider picture is at cognitive debt, capability debt, and the rest.
The two definitions differ in scope, and the difference matters before citing either. Rohde uses capability debt as one layer of three: technical debt in artifacts and systems, capability debt in the human layer that maintains them, and institutional debt in the wider structures that reproduce skill and resilience. His argument runs to the societal scale, and what concerns him is economic fragility, narrowing entry paths and the concentration of power. His paper is a conceptual synthesis rather than new empirical work, which he states himself.
This page uses the term more narrowly, for the organisational phenomenon on its own: what a single organisation owes its own future when it automates the doing without redesigning the learning. That is a scope a manager can act on inside one firm, within one year.
The two accounts of the mechanism converge more than they diverge. Rohde separates masking, where plausible output is mistaken for durable capability, from the erosion that follows it. That is the sequence described here, reached from software engineering rather than from organisational research. Independent arrival at the same structure is a reason to take the phenomenon seriously rather than a reason to argue about the name.
A third use of the phrase, in a different field#
Anyone searching the phrase will meet a third use before either of the two above. It is not about people at all. Jeremy Jarrell, writing for software delivery teams, uses capability debt for friction in the development process itself: tracking time across three separate systems, or writing detailed documentation after a feature is finished to satisfy a reporting rule nobody has revisited. His definition is “any point of friction in your team’s software development flow”. He notes in passing that debt can happen to skills too, but the term as he defines it sits in the process rather than in the person.
The two are solved differently and on different timescales, which is the reason to separate them before citing either. Jarrell’s capability debt is an organisation performing below its potential because nobody has time to improve how the work flows. Count it in wasted hours, remove it in a quarter. The capability debt this page describes is an organisation losing the human judgement it will need later, because the practice that built that judgement was automated away. No hour count reaches it, repayment takes years, and nothing shows on any dashboard while the output still looks fine.
One cause, three mechanisms#
One thing causes capability debt: automating work faster than you redesign how people learn. Underneath that sit three mechanisms, each with its own page in this research. The missing rungs are the junior tasks that used to carry people up to senior judgement, removed by automation before anyone noticed they were load-bearing. The missed reps are the repetitions handed to the machine, so the work is done but the practice never happens. And synthetic seniority is the surface effect: output that looks like the product of judgement the person has not built.
This is a drift problem before it is anything else. No leadership team decides to hollow out its own capability. It accumulates by default, one automated task at a time, because the tool arrives faster than the redesign of how people develop. Capability debt is what drift costs, counted in people.
Twenty-seven years of experience, six percentage points#
Until 2025 the strongest objection to this argument was that nobody had measured it. Somebody now has. Budzyń and colleagues, publishing in The Lancet Gastroenterology and Hepatology, examined 1,443 colonoscopies performed without AI assistance across four Polish centres, by nineteen endoscopists averaging 27.6 years of experience, with a range from eight to thirty-nine. They compared the period before an AI detection tool was introduced with the period after.
Unassisted adenoma detection fell from 28.4 per cent to 22.4 per cent, a drop of six percentage points in the doctors' own unaided performance, within months of routine exposure to a tool that had been helping them. Graded entry.
Read what that is carefully. These were not trainees. They were among the most experienced practitioners in their field, and their capability without the machine degraded measurably while their capability with it improved. The study is observational rather than randomised, it covers one procedure in one country, and detection rate is a proxy for skill rather than skill itself. It is still the closest thing to direct measurement this argument has, and it came from medicine, where the consequences of a degraded professional are not a worse deck.
Fragile experts, and the blackout test#
The Polish study measures the erosion. A 2026 experiment measures what the erosion leaves behind. Sankaranarayanan gave 78 participants a programming task through a custom development environment, in three conditions: manual control, unrestricted AI, and a scaffolded version designed to make the user do some of the thinking. Both AI groups beat the manual control on the work itself, and they did not differ from each other.
Then the AI was taken away and participants had to maintain what they had built. The unrestricted AI group failed at 77 per cent, against 39 per cent for the scaffolded group. The author calls them fragile experts: people who produce expert-looking work and cannot support it once the support is removed. Graded entry.
The paper reaches for its own term, epistemic debt, and puts it in quotation marks. That is now a third researcher arriving independently at a debt metaphor for the same phenomenon, which says more about how obvious the metaphor is than about anyone's originality. It is one session, one blackout task, in novice programming, with no longitudinal follow-up. What it demonstrates is that the gap between assisted and unassisted capability can be produced experimentally and is large.
Grades that rose, then fell below where they started#
Bastani and colleagues ran nearly a thousand high-school students through three arms: unrestricted GPT-4, a hints-only tutor with guardrails, and a control. While the tool was present, grades rose 48 per cent with unrestricted access and 127 per cent with the tutor. Then access was removed. The unrestricted group scored 17 per cent lower than students who had never had the tool at all. The guardrailed tutor largely removed that harm. Graded entry.
This is the single most useful result in the whole literature for anyone deciding what to do, because it separates two things that ordinary measurement cannot. Performance with the tool went up in both AI arms. Capability without it went in opposite directions depending on interface design. An organisation watching output alone would have seen two successes.
Just past the frontier, the help becomes harm#
Dell'Acqua and colleagues gave 758 BCG consultants tasks inside and just outside GPT-4's competence. Inside the frontier, the AI-assisted consultants were dramatically better and faster. Outside it, they performed worse than consultants working with no AI at all. Graded entry. The frontier is jagged and its edge is not visible from inside the task, Carelessness has nothing to do with the failure.
Brynjolfsson, Li and Raymond studied 5,172 customer-support agents through a staged rollout. Productivity rose 15 per cent on average, 30 per cent for the newest and least experienced staff, and barely at all for the most skilled, because the tool transfers expert patterns to novices. Graded entry. That is a real gain and it is also the exact moment capability debt is created: the novice ships expert-looking work without the experience that expert-looking work used to require. The study measures output rather than development, over months rather than years, so what happens to those novices afterwards falls outside what it can tell us.
Autor and Thompson supply the reason this distinction has economic weight. Across four decades of task data covering 303 US occupations, automation that removed the less expert tasks from a job raised wages, and automation that removed the expert tasks lowered them. Graded entry. Their data ends in 2018, so it is a lens rather than a forecast. It does establish that which tasks get automated matters more than how many.
Aviation wrote this down in 1983#
Lisanne Bainbridge described the mechanism four decades before anyone applied it to knowledge work. Automating the routine parts of a task leaves the human with the hardest residue, monitoring and exception handling, while removing the routine practice that built the competence to handle it. Her sentence is the one to remember: By taking away the easy parts of his task, automation can make the difficult parts of the human operator's task more difficult.
Graded entry.
The regulators have been recording the consequences ever since. Canada's Transportation Safety Board, in its report on a 2019 floatplane crash off Addenbroke Island, quotes air-taxi operators surveyed in a separate safety issue investigation: concern was expressed that dependence on technology was causal in degradation of basic piloting skills
, and that over-reliance on GPS navigation may contribute to the decision to fly into adverse weather conditions
. That passage reports industry-wide concern and is not among this accident's findings as to causes, which cite weather, terrain-alerting ambiguity and fatigue. Cited here for what practitioners told their regulator, and not as a causal finding about that crash. Graded entry.
The US National Transportation Safety Board went further in March 2026, investigating two fatal crashes in which Ford BlueCruise-equipped vehicles struck stationary cars at highway speed, in San Antonio and Philadelphia in early 2024. Overreliance on the automation appears in the probable cause for both, rather than in the discussion, which is a meaningful distinction in an investigator's report. The recommendation is the interesting part: it asks for warnings about accumulated short distractions
over a prolonged period, meaning the monitoring system was defeated by ordinary human attention behaving ordinarily rather than by anyone circumventing it. Graded entry.
Driving is a continuous manual-control task with a machine watching the human, which is not the shape of AI-assisted professional judgement. What transfers is not the setting but the finding that competence decays quietly under supervision that feels adequate. More at what professions can learn from aviation.
What the boardroom says it can already see#
BCG surveyed 70 C-suite leaders and senior executives in June 2026. Half report already observing deskilling in their organisations, and more than 60 per cent expect it to become a material threat within three to five years. Graded entry. The base is seventy people. Anyone quoting the fifty per cent without the seventy is overstating it, and executive perception is not a measurement of anyone's skills. What it marks is a change in the executive agenda.
EY's 2025 Work Reimagined survey, covering 15,000 employees and 1,500 employers across 29 countries, found 88 per cent of employees using AI at work but only 5 per cent using it in ways that transform how they work, while 37 per cent worry that overreliance could erode their skills and expertise, rising to 43 per cent in the UK. Organisations pursuing AI gains on weak talent foundations saw those gains lag by over 40 per cent. Graded entry. Every one of those figures is self-reported perception at a single point in time, from a sample that is not random, and the productivity comparison is EY's own modelling rather than an experiment. Taken for what it is, it says near-universal adoption sits alongside very shallow use, and the people doing the work are already worried.
The World Economic Forum's Future of Jobs Report 2025 names skill gaps the primary barrier to business transformation for 2025 to 2030, cited by 63 per cent of surveyed employers. Graded entry. Stated employer preference and actual hiring behaviour diverge routinely, so this is a statement of what employers say constrains them.
The symptoms no dashboard reports#
Capability debt rarely announces itself, so it has to be looked for. The output is consistently good, and fewer people can explain how it was produced or defend it under questioning. Juniors cannot do unaided the thing the AI now does for them, and are not expected to try. Important decisions have no clearly accountable human owner, because the recommendation came from a system. Verification is treated as a rubber stamp rather than skilled work. And the organisation has lost the ability to say which of its capabilities live in its people and which now live only in its tools.
The common feature is that every one of these is invisible in output metrics and visible only when someone asks a person to work without the tool. That is why the diagnostic question is not how much AI you use. It is what happens when it is removed, which is the question assessing capability rather than output is built around.
What nobody has measured#
There is no validated index of capability debt, and no study measures the thing itself at organisational scale. The evidence above establishes the components: measured deskilling in one clinical setting, experimentally produced fragility in one programming task, a reversal in one school trial, the novice-expert transfer in one support centre, and executive perception at a base of seventy. Nobody has aggregated those into an organisational measurement, and this page does not pretend otherwise. That is why the SuperSkills work is building a Capability Debt Index rather than asserting a figure.
Four things cut against the argument and belong here rather than in a footnote. The first is the most direct test yet run, and it came back against the prediction.
A randomised trial removed the tool afterwards and found no deficit. Cruces and colleagues gave 1,174 adults a workplace-style problem-solving task with or without a generative AI assistant, then an unassisted module. Graded entry. Treated participants did not perform worse once the assistant was taken away. They also found AI compressing an education gap while it was present, from 0.548 standard deviations to 0.139, closing about three-quarters of it.
That is the shape of experiment this page's argument predicts a deficit from, and the deficit did not appear. It should be read as a genuine challenge rather than explained away.
Two things limit how far it reaches, and both are the authors' own. A sizeable gap re-emerged once the assistant was removed, so the equity gain is partly transient. And a single session with an immediate unassisted module tests transfer within a sitting, where the studies that did find post-removal deficits taught a body of knowledge over time and removed the tool afterwards. Those may be different constructs rather than contradictory findings, and nobody has run the experiment that would settle it.
Automation does not always deskill. Lee's panel of Japanese nursing homes found robot adoption raised employment, improved retention, shifted worker effort towards direct care, and improved quality on hard measures: less physical restraint, fewer pressure ulcers. Graded entry. Japanese long-term care faces an acute labour shortage, so the robots substituted for vacancies rather than for people, which is a particular condition rather than a general one. It remains a real case where the machine arrived and capability improved.
Interface design changes the outcome more than adoption levels do. Bastani's guardrailed arm and Sankaranarayanan's scaffolded arm both largely removed the harm while keeping most of the gain. If that finding generalises, capability debt is a design failure rather than an inevitability, and the pessimistic reading of the other studies is too strong.
Shen and Tamkin reached the same place from the other direction, studying developers learning an unfamiliar programming library. Graded entry. Those with an AI assistant scored 17 per cent lower on a comprehension quiz than those who coded by hand, while finishing only marginally faster, so the trade was a poor one on average. The part worth carrying is what they found underneath the average: six distinct interaction patterns, three of which preserved learning even with the assistant present. The variable was how people used it, not whether.
The timescale is unproven. Every study here runs over months. Capability debt is a claim about years. Whether six percentage points of endoscopist skill return with practice, plateau, or compound is not known. Nobody has measured how long recovery takes, or whether it happens at all. Can you regain a skill you have lost sets out what little is known.
Paying it down#
You reduce capability debt the way you handle any debt: stop taking it on without deciding to, then pay down what you have. Four things do most of the work.
Decide in advance where human judgement must remain, rather than letting automation settle it case by case. Protect the practice that builds judgement, which means keeping some work unaided and building new development rungs on purpose to replace the ones automation removed. Treat verification as real, skilled work and staff it accordingly, because in an AI-assisted organisation the checking is where the judgement now sits. And measure capability directly, through live decisions, simulation and the ability to explain and defend work, rather than trusting that good output implies a capable person.
The evidence points at one design principle above the others. In both experiments where an interface was deliberately built to make the user do part of the thinking, the capability harm largely disappeared and most of the performance gain survived. The choice is not whether to use the tool. It is whether the version you deploy leaves the human any reps.
This is the practical content of drift versus design, applied to the one asset that does not appear on a balance sheet.
Key sources
- Cruces, G., Fernandez Meijide, D., Galiani, S., Galvez, R. and Lombardi, M. (2026). Does generative AI narrow education-based productivity gaps? Evidence from a randomized experiment. NBER Working Paper 34851. Graded entry.
- Shen, J. H. and Tamkin, A. (2026). How AI Impacts Skill Formation. arXiv:2601.20245. Graded entry.
- Budzyn, K., Roman'czyk, M., Kitala, D. et al. (2025). Endoscopist deskilling risk after exposure to artificial intelligence in colonoscopy: a multicentre, observational study. The Lancet Gastroenterology and Hepatology, 10(10), 896-903. DOI 10.1016/S2468-1253(25)00133-5. Graded entry.
- Sankaranarayanan, S. (2026). Mitigating "Epistemic Debt" in Generative AI-Scaffolded Novice Programming. Graded entry.
- Bastani, H., Bastani, O., Sungu, A. et al. (2025). Generative AI Without Guardrails Can Harm Learning. Graded entry.
- Brynjolfsson, E., Li, D. and Raymond, L. (2025). Generative AI at Work. Quarterly Journal of Economics, 140(2), 889-942. Earlier version NBER Working Paper 31161. Graded entry.
- Dell'Acqua, F. et al. (2023). Navigating the Jagged Technological Frontier. Harvard Business School and BCG. Graded entry.
- Autor, D. and Thompson, N. (2025). Expertise. Graded entry.
- Bainbridge, L. (1983). Ironies of Automation. Automatica, 19(6), 775-779. Graded entry.
- Lee, Y. (2024). Robots and Labor in Nursing Homes. NBER Working Paper 33116. Graded entry.
- Transportation Safety Board of Canada (2019). Aviation Investigation Report A19P0112, Seair Seaplanes, Addenbroke Island. Graded entry.
- National Transportation Safety Board (2026). Highway Investigation Report HIR-26-02, Ford BlueCruise collisions, 31 March 2026. Graded entry.
- BCG (2026). When Everyone Uses AI, Companies Risk Losing Critical Skills, 10 June 2026. Base: 70 C-suite and senior leaders. Graded entry.
- EY (2025). Work Reimagined Survey 2025. 15,000 employees and 1,500 employers, 29 countries. Graded entry.
- World Economic Forum (2025). The Future of Jobs Report 2025. Graded entry.
- Rohde, W. (2026). Short-Term Gain, Long-Term Fragility. SSRN; also arXiv 2605.27399. Graded entry.
- Carr, N. (2014). The Glass Cage: Automation and Us. W. W. Norton.
- Cunningham, W. (1992), on the origin of the technical-debt metaphor: Introduction to the Technical Debt Concept, Agile Alliance.
Related SuperSkills research#
Capability debt is the organisational accumulation of effects developed elsewhere: AI and human judgement, AI and critical thinking, the missing rungs, the missed reps, synthetic seniority and drift versus design. On the learning mechanism underneath it, see how humans learn with AI, desirable difficulty and deskilling. On measurement, assessing capability rather than output and can you regain a skill you have lost. On the individual version, using AI without dependency and staying valuable in the age of AI. Lisanne Bainbridge made this argument about process control in 1983; see the essential works. For the dated record of how this argument developed, see the timeline. The withdrawn attribution claim for this term is logged in corrections.
About this research#
Rahim Hirji is the author of SuperSkills (Kogan Page, 2026), keynote speaker on AI and human capability, and founder of The SuperSkills Intelligence Company. This work draws on research across more than 200 organisations in 30 countries over seven years. Findings are attributed to the studies that produced them and kept separate from the interpretation and the named concepts. Capability debt is part of the SuperSkills lexicon, used without a claim of first use. This is a living reference, reviewed and updated as significant new evidence appears.
Corrected 30 August 2026. Three changes after every source on this page was re-read at its issuing body. The endoscopists' average experience was given as 28 years and is 27.6, ranging from eight to thirty-nine. The NTSB recommendation was quoted as "accumulated short glances"; the report says "accumulated short distractions". And the Transportation Safety Board passage on piloting skills was presented in a way that implied a causal finding about the Addenbroke Island crash, when it reports concerns raised by air-taxi operators in a separate safety issue investigation; that is now stated on the page. A sentence claiming the term as a coinage was also removed, which the standfirst had already contradicted.
Cite this
Hirji, R. (2026). Capability debt. The SuperSkills Intelligence Company. Last reviewed 4 September 2026. thesuperskills.com/research/capability-debt
