The professions where the machine takes the judgement rather than the preparation, where the automated steps are the ones people used to climb to competence, where nobody ever measures what the professional can still do alone, and where the feedback on being wrong arrives too late to correct anything. Those four conditions travel together, and none of them appears on an exposure ranking.
This page exists so that the profession-by-profession pages on this site inherit an argument instead of repeating one. It is a method rather than a league table, and the working is shown so that anyone who disagrees can disagree with something specific.
Exposure rankings answer a different question
Almost every published ranking of professions at risk measures exposure: what share of a job's tasks a model could touch. Eloundou and colleagues estimated that around 80 per cent of US workers could have at least 10 per cent of tasks affected, and about 19 per cent could see at least half affected. Their paper says plainly that this is exposure and not displacement. It is quoted the other way round more often than any other number in the field.
The genre has form. Frey and Osborne's 2013 estimate that around 47 per cent of US employment sits at risk was a measure of technical susceptibility across whole occupations. Arntz, Gregory and Zierahn re-estimated the same question task by task using PIAAC data and got 9 per cent, noting that occupations labelled high-risk usually contain a substantial share of tasks that are hard to automate. Neither figure has been scored against what actually happened. The honest reading is that a decade of exposure modelling has produced a range of five to one and no verdict.
Deskilling is a different quantity again. A profession can have most of its tasks touched and lose nothing, because the touched tasks were never where the capability lived. A profession can have one task automated and lose a great deal, if that task was the one where the judgement got built.
The distinction that does the work: substitution against complementarity
Autor and Thompson analysed four decades of task data across 303 US occupations from 1980 to 2018, with a content-agnostic measure of how expert each task is. Their result reframes the whole argument. Automation that removed the less expert tasks raised wages and reduced employment. Automation that removed the expert tasks lowered wages and increased employment. The same volume of automation, applied to different parts of the same job, produced opposite outcomes.
Their data ends in 2018, so this is a lens rather than a forecast, and they say so. The lens has since been held up to generative AI. Brynjolfsson, Chandar and Chen, using ADP payroll microdata covering millions of US workers, find no economy-wide displacement but employment among 22 to 25 year olds in highly AI-exposed occupations running about 19 per cent below where it would sit had it tracked similarly aged workers in less-exposed occupations. The decline runs through reduced hiring rather than separations, and it concentrates in occupations where AI substitutes for human tasks. Where it complements, employment is flat or rising, particularly for experienced workers.
Two studies, two methods, forty years apart in their data. Both say the direction of the effect is set by which tasks the machine takes.
Then the case that cuts against all of it, which belongs on the page rather than in a footnote. Kanazawa and colleagues studied a Japanese taxi fleet through the rollout of an AI demand-prediction system and found the gains going almost entirely to the low-skilled drivers, narrowing the gap between best and worst by 14 per cent. Anyone arguing that AI reliably erodes expertise has to account for that result, which runs the other way.
Four conditions, and why all four matter
Deskilling risk is high where these hold together. Each is drawn from a specific finding rather than from intuition. The underlying substitution distinction belongs to Autor and Thompson; the assembly into a working test is this research's, and the test itself has never been validated against outcomes.
- One. The tool substitutes for the judgement, not the preparation. A model that assembles the material and leaves the decision is a complement. A model that produces the decision and leaves the professional to agree is a substitute wearing a supervisor's badge. Autor and Thompson supply the direction; the review-only configuration is where this fails most often.
- Two. The automated steps are the ones people climbed to competence. Brynjolfsson, Li and Raymond found a 14 per cent average productivity gain among 5,179 support agents, 34 per cent for the newest and least experienced and close to zero for the most skilled. Output rose. Whether those novices became experts was not measured, and could not be over months. This is the argument set out at missing rungs.
- Three. Nobody tests unassisted performance. Budzyn and colleagues found unassisted adenoma detection falling from 28.4 to 22.4 per cent in endoscopists averaging 28 years of experience. That number exists only because somebody measured colonoscopies performed without the tool. In almost every profession, nobody does. Aviation is the exception, through recurrent proficiency checks that can be failed.
- Four. Error feedback is delayed, diffuse or absent. Arthur and colleagues' meta-analysis of 189 data points found skill loss running from d of -0.01 immediately after training to d of -1.4 after more than a year of non-use, with cognitive and accuracy-based tasks decaying faster than physical and speed-based ones. A professional who never learns they were wrong cannot recalibrate, and the skill decays on the untested side.
Condition four explains why aviation and surgery, both heavily automated and both intensely safety-critical, do not look the same. A pilot flying a badly configured approach finds out within minutes. A radiologist who misses a nodule may never find out at all.
The cognitive half goes first
Casner and colleagues put 16 airline pilots into a Boeing 747-400 simulator with automation varied across routine and non-routine scenarios. Instrument scanning and manual control held up, even among pilots reporting little recent hand-flying. What degraded was the cognitive layer: tracking position without a map, deciding the next navigational step, recognising instrument failures.
Sixteen pilots in a simulator is not a general law, and the paper does not claim to be one. But the shape recurs. The visible, practised, physical part of a professional skill survives disuse better than the invisible part that decides what to do. Deskilling audits that test whether someone can still perform the procedure are testing the half that lasts.
Eight professions, scored, with the reasoning visible
Assessed against the four conditions on evidence available in August 2026. These are judgements, not measurements, and the reasoning is shown so it can be argued with.
- Diagnostic imaging and endoscopy. Highest risk, and the only one with direct evidence. All four conditions hold. The tool produces the finding; unassisted rates are almost never measured; feedback on a miss is delayed by months or never arrives. Budzyn measured the loss. Yu and colleagues found the effect on individual radiologists ranging from strongly positive to strongly negative with no usable predictor of which.
- Early-career legal and audit work. High risk, by a different route. Condition two dominates. The document review, first-draft and citation-checking work that built professional judgement is the most automatable part of the job. The capability is not lost by anyone; it is never acquired. See how AI changes law.
- Management consulting. High on condition one, low on condition four. Dell'Acqua and colleagues found 758 consultants performing dramatically better inside the model's competence and worse than consultants with no AI at all outside it. Feedback in consulting is comparatively quick and commercially brutal, which cuts the risk. See how AI changes consulting.
- Software engineering. Medium, and the loudest disagreement. METR randomised 16 experienced open-source developers across 246 real tasks and measured them 19 per cent slower with AI tools permitted, while they believed themselves about 20 per cent faster afterwards. Tests, compilers and production incidents give fast feedback, which weakens condition four. Condition two is the live worry.
- Teaching. Medium, and asymmetric. Preparation and marking are complements; the diagnostic judgement of what a particular pupil has misunderstood is not something current tools substitute for well. The risk sits in the marking, where the feedback loop is weak.
- Clinical general practice. Medium. Documentation is a complement with measured wellbeing benefit. Diagnosis is where condition one bites, and the one randomised trial of physicians using a large language model for diagnostic reasoning found no improvement over conventional resources.
- Journalism. Medium to low on capability, high on economics. The threat is to the business model rather than to the skill, and the two are constantly conflated. Interviewing, source cultivation and knowing what is being hidden are not the automated parts.
- Skilled trades and frontline physical work. Lowest, and consistently under-studied. Condition three inverts: performance is visible, immediate and inspected. Arthur's meta-analysis finds physical and speed-based skills decaying more slowly. This research covers this territory less well than it should, and that is a gap rather than a finding.
What this test cannot do
- It has never been validated. No study has taken a set of professions, scored them on conditions like these in advance, and checked afterwards whether the scores predicted anything. Until one does, this is a structured way of arguing rather than a measurement.
- There is one direct deskilling measurement in the whole literature. Budzyn, observational, one procedure, one country. A single study carrying this much weight is a reason for humility, not confidence.
- Nothing here predicts individual risk. Yu and colleagues looked for predictors of who benefits from AI assistance and found that experience, subspecialty and prior AI familiarity all failed. Profession-level reasoning cannot be pushed down to a person.
- No countermeasure has been tested. Recurrent unassisted assessment is an inference from aviation regulation, not a clinical or professional intervention anyone has trialled.
- Deskilling can be the right trade. Nobody mourns the loss of mental arithmetic at the till. The question is whether the capability being surrendered is the one the profession is paid for and accountable for, and that is a judgement about the profession rather than about the technology.
What a profession can do with this
Four conditions, four counter-moves, in rough order of how cheap they are.
- Record an unassisted baseline before deployment. Budzyn's finding was only possible because unassisted procedures kept happening. A profession that goes fully assisted on day one destroys its own ability to detect the problem.
- Protect the reps that build the judgement, not the ones that fill the time. The test is whether a task is where a junior learns to be wrong safely. If it is, automating it is a training decision rather than an efficiency one.
- Shorten the feedback loop deliberately. Where outcomes arrive late, manufacture earlier signals: blind second reads, sampled audit, calibration exercises against known answers.
- Make the assisted and unassisted gap a reported number. A widening gap is information about the service. Nobody currently reports it anywhere.
Key research and primary sources
- Autor, D. and Thompson, N. (2025). Expertise. NBER Working Paper 33941; Journal of the European Economic Association, 23(4).
- Brynjolfsson, E., Chandar, B. and Chen, R. (2026). Canaries in the Coal Mine? Six Facts about the Recent Employment Effects of Artificial Intelligence. Stanford Digital Economy Lab.
- Budzyn, K., Roman'czyk, M., Kitala, D. et al. (2025). Endoscopist deskilling risk after exposure to artificial intelligence in colonoscopy. The Lancet Gastroenterology and Hepatology, 10(10).
- Arthur, W., Bennett, W., Stanush, P. L. and McNelly, T. L. (1998). Factors that influence skill decay and retention. Human Performance, 11(1).
- Casner, S. M., Geven, R. W., Recker, M. P. and Schooler, J. W. (2014). The Retention of Manual Flying Skills in the Automated Cockpit. Human Factors, 56(8).
- Eloundou, T., Manning, S., Mishkin, P. and Rock, D. (2024). GPTs are GPTs, and Arntz, M., Gregory, T. and Zierahn, U. (2016). The Risk of Automation for Jobs in OECD Countries.
- Kanazawa, K., Kawaguchi, D., Shigeoka, H. and Watanabe, Y. (2022). AI, Skill, and Productivity: The Case of Taxi Drivers. NBER Working Paper 30612.
- Yu, F., Moehring, A., Banerjee, O. et al. (2024). Heterogeneity and predictors of the effects of AI assistance on radiologists. Nature Medicine, 30(3).
- Model Evaluation and Threat Research (2025). Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity.
Related SuperSkills research
On the concept, deskilling and capability debt. On the professions, medicine, law and consulting. On the structural answer, what professions can learn from aviation and whether a lost skill comes back. On the junior end, missing rungs, missed reps and entry-level jobs. On measurement, assessing capability rather than output.
About this research
Rahim Hirji is the author of SuperSkills (Kogan Page, 2026) and founder of The SuperSkills Intelligence Company. The substitution and complementarity distinction is Autor and Thompson's and is credited to them throughout. The four conditions are this research's assembly of separate findings into a working test, and no claim is made that the assembly is novel or that it has been validated; the page says so in its own words. Every figure was checked against the primary source. Reviewed quarterly.
Cite this
Hirji, R. (2026). Which professions face the greatest deskilling risk? The SuperSkills Intelligence Company. Last reviewed 28 August 2026. thesuperskills.com/research/which-professions-face-the-greatest-deskilling-risk