← Research
Research

How leaders should respond to AI

The decisions that matter are not about tools. They are about what you will not delegate, and who is answerable when the machine is wrong.

Last reviewed: 26 August 2026

How should leaders respond to AI? Not by being enthusiastic or cautious, but by noticing which decisions are being settled without anyone deciding. This page sets out the evidence, the four postures, and six decisions worth making deliberately.

Questions this page answersAll 616 questions this research covers

The decisions that matter are not about tools. Buying the licences is the easy part and it is already done in most organisations; the hard part is deciding what you will not delegate, and who is answerable when the machine is wrong. The evidence supports an uncomfortable summary of where most leadership teams actually are. Adoption has run faster than the personal computer, measured effects on pay and hours are so far close to zero, and underneath that flat surface the structure of work is already being reorganised. In other words, the consequential choices are being made right now, mostly by default, by people well below the executive team, and the numbers that would tell you about it will not move for years. Leadership here means noticing which decisions are being settled without anyone deciding, and taking those back. Enthusiasm and caution about AI are both beside the point.

The pace, and what it has produced#

Start with the pace. Bick, Blandin and Deming found that by late 2024, nearly forty percent of US adults aged 18 to 64 used generative AI, twenty-three percent of employed respondents had used it for work in the previous week and nine percent used it every working day, with work adoption as rapid as the PC and overall adoption faster than the internet. And only one to five percent of total work hours were actually being assisted. Enormous reach, thin penetration into the hours: your people have the tool and your work has not been redesigned.

Then the outcomes. Humlum and Vestergaard linked adoption surveys to administrative labour records across roughly 25,000 Danish workers in 7,000 workplaces and eleven exposed occupations. Two years after ChatGPT launched they found precise null effects on earnings and hours, ruling out effects larger than two percent, while documenting substantial task reorganisation and new tasks in content generation, AI oversight and AI integration. Their phrase for it deserves to be read twice in a board meeting: technological change reshapes work well before it surfaces in earnings or hours.

On the design of oversight, the most important recent result is a warning against the reflex. Vaccaro, Almaatouq and Malone's 2024 meta-analysis of 106 experimental studies and 370 effect sizes found human-AI combinations performing significantly worse on average than the better of human or AI alone, with losses concentrated in decision-making and gains in content creation. Adding a person to a process is not a control. Dell'Acqua and colleagues showed the sharp edge of this with 758 consultants: outside the model's competence, those using GPT-4 did worse than those using none, because they trusted confident output they should have questioned.

On where value settles, Autor and Thompson analysed four decades of task data across 303 occupations and found that automation removing the less expert tasks raised wages, while automation removing the expert tasks lowered them. And on capability, Bastani and colleagues, in a 2025 PNAS field experiment, found that unrestricted access to a GPT-4 tutor left students performing seventeen percent worse than a control group once the tool was removed, while a version designed to give hints rather than answers largely eliminated the harm. Design of the tool, not presence of the tool, determined whether people developed. Meanwhile the World Economic Forum's 2025 Future of Jobs report names analytical thinking as the most valued core skill and skills gaps as the biggest barrier to transformation.

How far the Danish result travels#

Denmark is a high-trust, heavily unionised labour market with strong employment protection, and two years is early for a general-purpose technology; null wage effects there do not settle the question elsewhere. The meta-analysis covers studies published between 2020 and mid-2023, so it predates the current model generation, which probably shifts more tasks into the category where human intervention subtracts rather than adds. The Autor and Thompson data runs to 2018, making it a lens rather than a forecast. And the Bastani experiment was school mathematics, not professional work, so it establishes that the design variable exists and matters without telling you what to build.

The largest uncertainty is unresolvable today: nobody has measured what a decade of AI-assisted work does to organisational capability, because a decade has not passed. Anyone selling you certainty in either direction is selling something.

Nobody decided the current configuration#

Most organisations do not decide their way into their AI configuration. They arrive at it, through a thousand small choices nobody quite made, and then describe the result as a strategy. That is what I mean by drift versus design. It is the single most useful lens I can offer a leadership team, because it moves the conversation off whether AI is good or bad and onto a question with an answer: which of these decisions did we actually make?

In the CEOWORLD piece where I set the framework out, I described four postures organisations take. The Sleepwalkers are moving fast with no design, mistaking activity for transformation. The Programmed have adopted someone else's design, usually a vendor's, and are executing it without having chosen it. The Stuck have seen the risk clearly enough to freeze, and are losing ground while they deliberate. The Designers have decided in advance where human judgement has to remain and are building towards it. The postures are not a maturity model and you do not graduate through them; most large organisations contain all four simultaneously, function by function, which is itself the finding.

The leadership failure I see most often is the substitution of a phrase for a decision, rather than recklessness or timidity. "We keep a human in the loop" is the most common example, and the meta-analysis is now the empirical case against it: a person placed at the end of a process, with no time budget, no stated basis on which they would disagree, no authority to stop it and no consequence for approving, does not produce oversight. They produce a signature, and they can make the system worse than either party alone. Governance that cannot name who, at what point, with what authority to say no, is not governance.

The second failure is slower and more expensive. Redesigning the doing without redesigning the learning produces the productivity gain and the damage at the same time, and only one of them is visible this year. That accumulation is capability debt, and its distinguishing feature is that outputs look fine throughout, right up to the decision the AI cannot make and nobody has been trained to make either.

Three organisations that reversed course, and what they said#

The reversals are more instructive than the launches, because organisations explain themselves when they change their minds.

In August 2025 the Commonwealth Bank of Australia reversed forty-five call-centre redundancies attributed to an AI voice bot, after the Finance Sector Union took the matter to the Fair Work Commission. Call volumes had risen; overtime was being offered and team leaders put back on phones. The bank's explanation is the part to notice: it did not concede the technology had failed, but that its own process "did not adequately consider all relevant business considerations". The decision, not the tool, was the error.

Klarna's much-repeated reversal is routinely garbled, so here it is accurately. In May 2025 the company began rehiring human agents on a flexible model and guaranteed customers could always reach a person. Its chief executive said that cost "seems to have been a too predominant evaluation factor... what you end up having is lower quality". He did not say the AI had failed. He said the wrong thing had been optimised, which is a story about a decision criterion rather than a technology.

Most instructive of all: in December 2025 the Dutch national police decommissioned their Crime Anticipation System, running nationally since 2017. Campaigners had pressed a bias critique for a decade. The stated reason for shutting it down was different and, for any executive, more uncomfortable: the operational value was unclear "because there were no clear goals or measurable success criteria". It was never properly evaluated, so when it came under pressure nobody could defend it. That is the fate awaiting any AI programme measured by adoption rather than by outcome.

Two arguments a leadership team should hear before acting#

The first is a corrective to urgency. Daron Acemoglu's The Simple Macroeconomics of AI estimates the effect on total factor productivity over ten years as modest, roughly an order of magnitude smaller than the most quoted forecasts. Any board being shown a trillion-dollar opportunity slide should have this paper alongside it. It does not argue that nothing is happening; it argues that the scale being sold is not supported, and that the numbers depend heavily on which tasks prove genuinely automatable rather than merely exposed.

The second is a corrective to fatalism. Carl Benedikt Frey's The Technology Trap traces automation across several centuries and finds that periods of technological progress have frequently produced decades of falling wages and political backlash before broad gains arrived. The lesson is that things have often worked out eventually, while punishing a generation in between, and that the difference was made by choices about how technology was directed rather than by the technology itself. That is the strongest available reason not to leave the decisions on this page to resolve themselves.

Together they define the useful posture: sceptical about the promised scale, serious about the design choices. Microsoft's 2026 Work Trend Index, built on trillions of usage signals and 20,000 AI-using workers, is worth reading in the same session, and worth reading as a signal about where the field's attention has moved rather than as independent evidence: the company with the most usage data has organised its annual report around human agency and judgement allocation.

Six decisions a leadership team should actually make#

What to stop doing#

Stop reporting adoption as progress. Seat counts and prompt volumes measure activity, and people learn to perform whatever you measure. I call the result usage theatre.

Stop treating this as a technology decision with an HR appendix. The consequential choices are about work design, accountability and development, which means they belong to the executive team and the CHRO rather than to procurement. See the CHRO guide to AI.

Stop waiting for the productivity number. It is the slowest indicator available and it will move long after the positions have been taken.

Stop cutting the graduate intake on a productivity argument without pricing the pipeline. The saving is this year and visible. The cost is your senior population in seven years and invisible. See will AI replace entry-level jobs.

Development of the idea#

The four postures and the drift versus design framework are set out in CEOWORLD, Drift versus design: why most companies mistake activity for transformation (9 July 2026), and developed in the Box of Amazing essay The Architecture of Drift. The accountability argument is in the European Business Review, Why the Real AI Risk is Not Automation, but Accountability Gaps in Leadership Decisions (21 August 2026). On measurement, Irish Tech News, You're not adopting AI. You're paying for it. (July 2026). This is worked through properly in SuperSkills (Kogan Page, 2026).

Key research and primary sources

On the organisational plan, AI workforce strategy and the AI readiness lie. On oversight design, human and AI decision making, Human at the Start and AI agents and human judgement. On the capability consequences, capability debt and how humans learn with AI. The graded evidence is in the evidence base. On the literacy obligation now in force, what AI literacy means for leaders, and on the oversight duty, meaningful human oversight. For the board conversation, twelve questions a board should ask.

About this research#

Rahim Hirji is the author of SuperSkills (Kogan Page, 2026), keynote speaker on AI and human capability, and founder of The SuperSkills Intelligence Company. This work draws on research across more than 200 organisations in 30 countries over seven years. Findings are attributed to the studies that produced them and kept separate from the interpretation, which is the author's. Drift versus design, the four postures, capability debt and usage theatre are part of the SuperSkills lexicon; automation bias and cognitive offloading are established concepts from the research literature and are not his. This is a living reference, reviewed and updated as significant new evidence appears.

How this research works  ·  Reviewed quarterly  ·  Found an error? Tell me and it is corrected on the page.

Cite this

Hirji, R. (2026). How leaders should respond to AI. The SuperSkills Intelligence Company. Last reviewed 26 August 2026. thesuperskills.com/research/how-should-leaders-respond-to-ai

Questions answered on this page

How should leaders respond to AI?

By deciding what will not be delegated and who is answerable when the machine is wrong, rather than by being enthusiastic or cautious about the technology. Adoption has run faster than the personal computer while measured effects on pay and hours remain close to zero, which means the consequential choices are being made now, largely by default, and the numbers that would reveal it will not move for years. Rahim Hirji sets out six decisions a leadership team should make deliberately, beginning with a written map of where human judgement must remain.

What are the four postures organisations take towards AI?

Rahim Hirji's framework, set out in CEOWORLD in July 2026, names four. Sleepwalkers move fast with no design, mistaking activity for transformation. The Programmed have adopted someone else's design, usually a vendor's, without having chosen it. The Stuck have seen the risk clearly enough to freeze and are losing ground while they deliberate. Designers have decided in advance where human judgement has to remain. It is not a maturity model: most large organisations contain all four at once, function by function.

Why is keeping a human in the loop not enough?

Because it substitutes a phrase for a decision. A 2024 meta-analysis in Nature Human Behaviour covering 106 studies found human-AI combinations performed significantly worse than the better of human or AI alone, with losses concentrated in decision-making. A person placed at the end of a process with no time budget, no stated basis for disagreement and no authority to stop it produces a signature rather than oversight. Governance that cannot name who, at what point, with what authority to say no, is not governance.

What should leaders measure about AI?

Something that can fall. Adoption metrics such as seat counts and prompt volumes only rise. Executives like them for that reason, and people learn to perform them. Rahim Hirji calls that usage theatre. Add at least one measure of capability without the tool, and one of how often humans genuinely disagree with machine output. An approval rate near a hundred percent is a finding, not a success.

In this hub

Organisations and leadership

What a leadership team actually has to decide, and what to measure.

The work

Where the writing comes from.

These essays draw on research across more than 200 organisations in 30 countries. See the wider body of work, or bring it into your organisation.

All research →

Somewhere to think this through out loud is worth more than a consultant. Use that first. Failing that, turning a position into a working practice, privately, is what I do. Board advisory.

This argument is one a board usually meets for the first time in the room. There is the boards and leadership version, and the full range of topics and audiences.

Box of Amazing

Rahim’s free weekly letter on AI and human capability

If this was useful, the weekly letter is where the thinking happens first. Most of what ends up on this site starts there. Weekly essays on AI, capability and the future of work. Read by 25,000 people, every week since 2017. Free, and one click to stop.

Opens Substack to confirm. No pitch in it, unsubscribe in one click, and nobody follows up because you read something.