← Research
Research

How will AI change the public sector?

Fast on the administrative half, slow on the deciding half, and reported as one number. The sector already owns the deepest precedent in this debate, and it is a rule of evidence rather than a productivity study.

Last reviewed: 1 September 2026

The GDS cross-government experiment, the DWP evaluation and its own limitations chapter, what the National Audit Office found about ownership, the 143 transparency records, and the presumption that the computer is right.

Questions this page answersAll 435 questions this research covers

Quickly on the administrative half and slowly on the deciding half, and the two are being reported as one number. The UK has run the largest public sector trial of a generative assistant anywhere, 20,000 civil servants across twelve organisations, and its headline finding is 26 minutes saved a day. That figure is self-reported, calculated from the midpoints of tick-box ranges, and the report says it could not identify how the saved time was spent. A second departmental evaluation, using regression rather than tick boxes, put the figure at 19 minutes and published a page of reasons to treat it cautiously. Meanwhile the number of algorithmic tools government has formally disclosed, across every department and public-facing body, stands at 143.

Two flagship numbers, and both of them are people estimating#

The Government Digital Service ran a cross-government experiment on Microsoft 365 Copilot from 30 September to 31 December 2024, giving licences to 20,000 government employees across twelve organisations including DWP, HMRC, the Home Office, the Ministry of Justice and the Office for National Statistics. Every participating organisation committed at least a thousand licences; the previous largest deployment anywhere had been three hundred.

Adoption reached around 80 per cent and held. 7,115 users responded to the survey. The findings report gives an average saving of 26 minutes a day, with drafting documents at 24 minutes, creating presentations at 19 and scheduling meetings at 9. 17 per cent of users noticed no clear saving at all. 82 per cent said they would not want to return to working without it, with satisfaction at 7.7 out of 10 and recommendation at 8.2.

The report is candid about how the 26 minutes was produced. Participants picked a band, from "none or less than 5 minutes" to "more than an hour", and the average was calculated from the midpoint of each band with the largest savings estimated at 60 minutes. The conclusions state that experimental constraints made it impossible to identify how the saved time was spent. So the most-cited number in UK public sector AI policy is a self-report, bucketed, with its top tail capped by the analysts.

The Department for Work and Pensions published its own evaluation on 29 January 2026, covering a trial of 3,549 staff from October 2024 to March 2025, with 1,716 responses from users and 2,535 from a stratified comparison group of non-users. A seemingly unrelated regression estimated 19 minutes saved a day across eight routine tasks, with the largest effects on searching for information at 26 minutes, writing emails at 25 and summarising at 24. Job satisfaction rose 0.56 points and perceived work quality 0.49 points, both on seven-point scales. 73 per cent reported better quality outputs and 65 per cent felt more fulfilled.

What makes the DWP report worth reading is its own limitations chapter, which is more honest than most academic papers on this subject. There was no baseline, because the surveys ran after people had begun using the tool. Licences were allocated first come, first served, often on managerial discretion or peer nomination, so the treatment group is self-selected towards enthusiasts. The report states that this self-selection may lead to an overestimation of Copilot's benefits, and separately flags acquiescence bias on the time-saving question, since respondents who wanted to keep the tool had a reason to answer generously.

Two large, well-run, unusually transparent evaluations, and neither of them measured a minute with a clock. That is not a criticism of either team, both of which say so themselves. It is a warning about the second-hand versions, which quote the number and drop the method.

The productivity case underneath the policy was never costed#

The National Audit Office surveyed 87 government bodies for its March 2024 report on AI in government. 37 per cent had deployed AI, typically with one or two use cases. 70 per cent were piloting or planning, with a median of four use cases each. Only 21 per cent had an AI strategy for their organisation. Of the 32 bodies with deployed AI, 24 always or usually had a named accountable owner, and fewer than half, 15 of 32, said use cases were always or usually identified at organisational level before deployment. Across all respondents, 30 per cent had risk and quality assurance processes that explicitly incorporated AI risks. 70 per cent named difficulty recruiting or retaining AI skills as a barrier.

The finding that should have travelled furthest and did not is about the money. The Cabinet Office's Central Digital and Data Office carried out indicative analysis in 2023 identifying that almost a third of civil service tasks, those it defined as routine, could be automated. The NAO records that it did not examine the feasibility of delivering those gains, and made no assessment of cost. That estimate is the foundation of the productivity claim that has been repeated through budgets, speeches and departmental plans since. It was an indicative sizing exercise, and the auditor said so.

A second finding, smaller and sharper: fifteen of thirty-two bodies could not say that AI use cases were identified at organisational level before they went in. Which means the tools arrived below the line of sight of the people accountable for them. This research has a name for the gap between an organisation's stated AI position and what its people are actually doing with the tools, at usage theatre, and the public sector version has a constitutional edge to it, because a minister is answerable for things a departmental board never saw.

One hundred and forty-three records#

The Algorithmic Transparency Recording Standard is the UK's disclosure regime for algorithmic tools in public decision-making. It has been mandatory for all government departments, and for arm's-length bodies delivering public or frontline services, since a scope and exemptions policy published in December 2024. Anyone can read the register.

As at 1 September 2026 it holds 143 records, from a standard first published in January 2023. They are worth reading rather than counting. DWP has published a tool that flags Universal Credit journal messages indicating a risk of harm, and a scanner that reads around 25,000 documents and letters a day to flag citizens who may need urgent assistance. The Cabinet Office has recorded the verbal and numerical tests used to sift civil service applicants. Ofsted has recorded a tool that drafts sections of children's home inspection reports. Newcastle City Council has recorded a system that writes adult social care case notes.

Those are not chatbots on a website. Each of them sits between a citizen and a decision about them, at volume. The register is the most useful public document in this field and it is also the measure of how much of this is happening in the open: 143 records against a civil service of hundreds of thousands, in a sector where 37 per cent of surveyed bodies had already deployed something in 2023.

The precedent this sector owns is the presumption, not the productivity#

Every profession has an oversight precedent it has forgotten it owns. For banking it is model risk management, as set out at accounting and audit. For the public sector it is a rule of evidence.

Section 69 of the Police and Criminal Evidence Act 1984 required a party relying on computer-produced evidence to show the computer had been working properly. Following a Law Commission recommendation in 1997, it was repealed, and from 2000 English law has operated a common law rebuttable presumption that a computer was operating correctly at the material time. The Ministry of Justice's own foreword puts it as bluntly as anyone could want:

In simple terms, "the computer is always right", unless someone can show it is not.

On 21 January 2025 the Ministry of Justice opened a call for evidence on that presumption, running to 15 April 2025, prompted by the Post Office Horizon prosecutions. The document proposes that any reform cover evidence generated by software, including artificial intelligence and algorithms, naming accounting systems, automated fraud and plagiarism detection, and automated reporting from handheld devices, while excluding material merely captured by a device such as photographs, messages and breathalyser readouts.

This matters more than any productivity figure on this page. A legal presumption that machine output is correct until challenged is automation bias written into procedure, with the burden of proof pointing the same way the bias already points. It took hundreds of wrongful convictions to get it reopened. The generative AI debate in government is being conducted almost entirely without reference to it, which is a strange thing to watch, because the sector has already run the experiment on what happens when an institution trusts a system nobody outside the vendor can inspect.

Four things that make government different, and none of them are efficiency#

What has not been measured, in a sector that measures everything#

Six things a public body can do without waiting for a strategy#

Key sources

On oversight, meaningful human oversight, why human in the loop is not a safeguard and the invisible work of oversight. On accountability, who can override an AI system and auditing an AI-assisted decision. On the measurement problem, measuring adoption properly and the most quoted AI statistics, checked. On the neighbouring sectors, accounting and audit and law.

About this research#

Rahim Hirji is the author of SuperSkills (Kogan Page, 2026), keynote speaker on AI and human capability, and founder of The SuperSkills Intelligence Company. The GDS findings report, the DWP evaluation, the NAO summary, the Ministry of Justice call for evidence and the algorithmic transparency register were each read at source. The 143 figure is the register's own count on the date shown and will move. Figures circulating about public sector AI savings in billions are not used here, because the underlying Cabinet Office analysis was, on the auditor's account, never tested for feasibility or cost.

How this research works  ·  Reviewed quarterly  ·  Found an error? Tell me and it is corrected on the page.

Cite this

Hirji, R. (2026). How will AI change the public sector? The SuperSkills Intelligence Company. Last reviewed 1 September 2026. thesuperskills.com/research/how-will-ai-change-the-public-sector

In this hub

Professions and sectors

Where the pressure lands first, profession by profession.

The work

Where the writing comes from.

These essays draw on research across more than 200 organisations in 30 countries. See the wider body of work, or bring it into your organisation.

All research →
Box of Amazing

Rahim’s free weekly letter on AI and human capability

If this was useful, the weekly letter is where the thinking happens first. Most of what ends up on this site starts there. Weekly essays on AI, capability and the future of work. Read by 25,000 people, every week since 2017. Free, and one click to stop.

Opens Substack to confirm. No pitch in it, unsubscribe in one click, and nobody follows up because you read something.

Running an event, or responsible for how AI arrives in your organisation? Keynotes  ·  Advisory and coaching  ·  Enquire