Most AI training teaches the tool: where the buttons are and how to write a prompt. That is the part people pick up on their own, and the part that goes out of date fastest. The evidence points at three other things. People need to learn where the tool is weak in their own work, because outside that boundary AI-assisted professionals have done worse than colleagues with no AI at all. They need a habit of checking that survives a busy week, because confident machine output gets followed when it is wrong. And they need to keep doing some of the work the tool now does, because skills that go unused decay. Train the judgement around the tool, not the interface.
The answer, in one line
Train the judgement around the tool rather than the tool itself: where it is weak in your own team's work, how to check its output, and how to keep the skills it takes over.
Where training stands now#
Training is catching up with use rather than introducing it. In Deloitte's 2026 survey of UK working adults, run by Ipsos, 63 per cent said they use generative AI for work, about half said they have had no formal training in using it safely and effectively, and 65 per cent reported a lack of convincing leadership guidance. Of the users, 31 per cent use it without their employer's knowledge. Microsoft and LinkedIn's 2024 Work Trend Index found 78 per cent of AI users bringing their own tools to work. Both are self-reported and both sponsors sell into the adoption they measure, so treat the figures as a direction rather than a census. The direction is clear enough: by the time a programme starts, people have already formed habits with the tool, and some of those habits are the ones the programme needs to change.
That changes what a course is for. An introduction to the interface assumes a beginner. The person in the room is more often a confident user who has never been shown where the tool fails in their own job, and who has never been asked to check its work against anything.
Teach the boundary, not the buttons#
The most useful single finding for anyone designing training comes from Dell'Acqua and colleagues' field experiment with consultants at a global firm. Inside what the authors call the jagged frontier, the tasks the model handles well, consultants using AI were much better and faster. Outside it, on a task built to look similar, they performed worse than consultants with no AI at all. The frontier is local. It runs through different tasks in a finance team, a legal team and a marketing team, and it moves each time the tool changes.
So the core of a good session is the team's own work, with known answers. Give people real tasks from their own week where the right output is already settled, let them use the tool, and compare. What you want from it is a shared, specific list of the places where this tool, in this team, produces something plausible and wrong. That list is worth more than any generic guide. It is the start of knowing when AI is wrong.
Train the check, and expect people to dislike it#
Automated aids produce two kinds of error, as Skitka, Mosier and Burdick showed in 1999: missing what the system failed to flag, and following advice that was wrong. Generative AI adds a reason both happen more. Zhou and colleagues found that only about 5 per cent of generated answers carry any marker of uncertainty, that confidently expressed answers were wrong on average 47 per cent of the time in their tests, and that people relied on plain unmarked statements nearly as often as on confident ones. The human side of that study is small and US-only, but the mechanism is the one described on automation bias.
Checking can be trained, with a catch. Buçinca, Malaya and Gajos found that cognitive forcing designs, which make a person commit to a view before seeing the AI's answer, significantly reduced overreliance, and that participants rated those designs least favourably. The intervention that works is the one people like least. That is why a checking habit cannot be left to goodwill after the course ends. It has to be written into the work: a named rule for each type of task about what gets checked, against what, and by whom. The US government reached the same conclusion for its own agencies: OMB memorandum M-24-10 made sufficient training, assessment and oversight for the operators of AI a minimum practice, and named automation bias as the reason.
Keep the underlying skill alive#
Training that only teaches use will, in time, remove the ability to check. Arthur and colleagues' meta-analysis of skill decay found losses running from almost nothing immediately after training to a very large effect after more than a year of non-use, with cognitive and accuracy-based tasks decaying faster than physical ones. Budzyń and colleagues found that endoscopists' detection rate in unassisted colonoscopies fell from 28.4 to 22.4 per cent after routine exposure to AI assistance. The study is observational and covers one procedure, but it measured the thing training programmes usually assume away. Shen and Tamkin's small experiment pointed the same way for learning: people who learned a new coding library with AI help scored 50 per cent on a later quiz against 67 per cent for those who coded by hand, with the widest gap on debugging.
The practical answer is protected practice. Name the few tasks each role must still be able to do unaided, and schedule time to do them. The case for that, and its limits, is on AI-free periods at work and deskilling.
Train in the work, with feedback#
The older evidence on training judgement says the same three things each time. Calibration improves with intensive feedback, most of the gain comes early, and it stays mostly confined to the task it was trained on, according to Lichtenstein and Fischhoff. A single well-designed session can reduce specific named biases for eight to twelve weeks, according to Morewedge and colleagues, though their outcomes were bias tests rather than decisions at work. And structured debriefs improved team performance by around 20 to 25 per cent in Tannenbaum and Cerasoli's meta-analysis, when they reviewed process rather than outcome and let people find the lesson themselves.
Put together, that argues against the one-day course and for something smaller and recurring: practice on the actual task, quick feedback on whether the output was right, and a short monthly debrief on where the tool helped and where it misled. More on the method at how judgement is trained and calibration training, and on why most large programmes miss it at why reskilling programmes mostly fail.
What a programme looks like#
Six parts, in order. Sort each team's tasks into three groups: hand over to the tool, use the tool with a check, and keep human. Run a boundary session on real work with known answers, and keep the list it produces. Write a checking rule for each task type in the second group. Protect unassisted practice for the third. Hold a monthly debrief. And measure something other than usage, because licences and prompts can rise while nothing improves; how to measure adoption properly covers what to count instead.
Underneath the programme sit four decisions only leadership can make: which decisions a machine may make, who can stop each one, what people must remain able to do, and how anyone would know if it went wrong. Training follows from those answers. Without them it teaches people to use a tool for purposes nobody has agreed. Those four decisions are the subject of Rules Before Tools.
What this does not show#
It does not show that any training programme for generative AI has improved organisational outcomes, because no study in the evidence base measures one. The boundary finding comes from one firm's consultants and one set of tasks. The checking evidence is mostly laboratory work with small samples. The skill-decay figures come from older training research and from medicine, and transfer to office work is an inference. The survey figures are self-reported and sponsored. What the evidence supports is the shape of a programme: train the boundary, the check and the retained skill, in the work, with feedback. Whether a particular programme built that way pays back is something each organisation will have to measure for itself.
Essay · SS-2026-413 · 8 peer-reviewed studies, 2 working papers, 1 institutional survey and 2 of other kinds
Hirji, R. (2026). How do we train people to use AI?. The SuperSkills evidence base, SS-2026-413. https://thesuperskills.com/research/how-do-we-train-people-to-use-ai. Last reviewed 5 October 2026.
An evidence review by Rahim Hirji, not peer-reviewed research. For a material claim, cite the underlying study as well; every study here carries its own permanent link.
How citations and IDs work