← Research
Research

How will AI change accounting and audit?

The profession exists to check other people's numbers. It has not yet checked the tools it uses to do that, and its own regulator says so.

Last reviewed: 31 August 2026

What the replacement of SR 11-7 by SR 26-2 did to the definition of a model, why the Financial Reporting Council took the opposite approach, the sentence in its thematic review that nobody has quoted, and what effective challenge requires from a person who has never done the work by hand.

Questions this page answersAll 383 questions this research covers

By separating the production of a number from the ability to say why it is right, in a profession whose entire licence rests on the second. An auditor's signature is not a claim that the figure is correct. It is a claim that somebody competent looked, and could have found it if it were wrong. Every question about AI in this profession reduces to whether that second claim survives, and two regulators have now taken opposite views of how to keep it alive.

This page uses regulatory primary text rather than survey data, because for once the primary text is where the argument actually is, and because both documents say something more interesting than their press coverage did.

The precedent everybody cites has just declined the job#

When somebody in financial services argues that AI governance is a solved problem, they mean model risk management. The Federal Reserve and the OCC issued SR 11-7 in April 2011: an inventory of every model, independent validation, ongoing monitoring, outcomes analysis, and above all "effective challenge", defined there as critical analysis by objective, informed parties that can identify model limitations and produce appropriate changes. Fifteen years of practice, whole professions built around it. It is the most mature framework any industry has for governing decisions made with machine-produced numbers.

On 17 April 2026 it was withdrawn. Supervisory letter SR 26-2, issued jointly by the Federal Reserve, the OCC and the FDIC, supersedes and replaces both SR 11-7 and SR 21-8. Its attachment carries footnote 3, which reads:

Generative AI and agentic AI models are novel and rapidly evolving. As such, they are not within the scope of this guidance. Nonetheless, a banking organization's risk management and governance practices should guide the determination of appropriate governance and controls for any tools, processes, or systems not covered in this document. However, the principles described in this guidance apply to traditional statistical and quantitative models and non-generative, non-agentic AI models.

That is the whole of what the new guidance says about generative AI. The phrases "artificial intelligence" and "machine learning", written out, appear nowhere in the document. There is no AI section and no AI heading.

Two things follow. The narrow one: anybody presenting model risk management as the off-the-shelf answer for generative AI in banks is presenting a document that, on its own terms, declines to cover it. The broader one is more interesting. The regulator with the deepest institutional experience of validating machine-produced numbers looked at generative systems in 2026 and decided its framework did not yet fit them. That is a considered judgement from the best-placed judge available, and it deserves more attention than it has had.

What the replacement stopped covering#

The definition of a model was tightened at the same time. SR 11-7 defined one as "a quantitative method, system, or approach that applies statistical, economic, financial, or mathematical theories, techniques, and assumptions to process input data into quantitative estimates". SR 26-2 says:

For the purposes of this guidance, the term "model" refers to a complex quantitative method, system, or approach that applies statistical, economic, or financial theories to process input data into quantitative estimates. The term "model" in this guidance excludes simple arithmetic calculations, such as those found within spreadsheets, as well as deterministic rule-based processes and software where there are no statistical, economic, or financial theories underpinning their design or use.

"Complex" is new. "Mathematical" has gone, and so have "techniques, and assumptions". The exclusion sentence is new. So is what disappeared alongside it: SR 11-7 carried a footnote saying that qualitative approaches falling outside its definition "should also be subject to a rigorous control process". No equivalent sentence survives. The 2026 replacement leaves out-of-scope tools to the firm's own judgement about what governance is appropriate.

Read the two changes together and a gap opens with a shape. A tool that is complex enough to be hard to challenge but is not built on statistical, economic or financial theory now sits outside the definition, and generative systems are the obvious inhabitants of that space. They are also, unlike a spreadsheet formula, the category where an independent reviewer has the hardest time reconstructing how the answer was reached.

Effective challenge itself survives, restated and slightly sharpened:

Effective challenge is performed by individuals with the appropriate expertise to conduct a critical and objective challenge, sufficient independence to maintain objectivity, as well as the organizational standing and influence to effect any change.

Expertise, independence, standing. Hold that definition, because the rest of this page is about the first of the three.

One caveat, so that nobody quotes this page too hard. SR 26-2 is guidance, not rule. It says so: it "does not set forth enforceable standards or prescriptive requirements", though a footnote adds that supervisory action may still follow from unsafe or unsound practices. It is most relevant to banking organisations above $30 billion in assets. And a tool being out of scope does not make it ungoverned, because third-party risk expectations, consumer protection law and the firm's own controls all still apply.

The audit regulator went the other way#

Ten months earlier, on 26 June 2025, the UK Financial Reporting Council did the opposite. Its first guidance on AI in audit defines its scope to comprise "both traditional machine learning techniques and deep learning models, including generative AI". No carve-out.

Its position on explainability is the part worth studying, because it is where a regulator has actually thought about what a human can be expected to do:

Explainability is a measure of the extent to which the behaviour or decisions of a tool can be understood, rather than the extent to which the inner mechanics and processing of the model are transparent and understandable to humans. Appropriate explanations may, particularly in relation to tools that rely on neural networks, be approximate or post hoc explanations that seek to explain how inputs influence outputs rather than the internal features and workings of the model.

The FRC declines to set a threshold, holding that "what constitutes appropriate explainability will vary widely based on context". In its worked example, an unsupervised model flags journals as anomalous, and the standard the firm sets is that the team should be able to see which features of a transaction contributed most to the flag. That is a workable standard and an honest one. It is also, and the FRC does not hide this, a decision to accept an approximation of the reasoning in place of the reasoning.

What it then requires of the engagement team is the interesting half. They must consider why an item was identified in order to decide what work will settle it, and:

the methodology requires them to be alert for any information that indicates that the tool's assessment of items as high risk or not may be systemically flawed in the context of this engagement.

Alert to the possibility that the tool is wrong in a patterned way about this particular client. That is a demanding thing to ask, and it falls to the team member running the tool. Note also what the guidance does not contain. "Professional scepticism" appears nowhere in it. "Over-reliance" appears nowhere. "Automation bias" appears exactly once, in a documentation row noting that training material should include strategies to mitigate it. And the guidance is explicit that it creates nothing new: "the requirements against which firms will be assessed remain only those in the ISQMs and ISAs (UK)".

One sentence in the thematic review#

Published the same day, and much less quoted, was the FRC's thematic review of how the six largest UK audit firms certify automated tools before use. The firms are named: BDO, Deloitte, EY, Forvis Mazars, KPMG and PwC. The review covers process as at the second quarter of 2024.

The headline is reassuring. All six had certification processes, though "the maturity of these processes was found to vary and in some cases were not supported by formal documented policies". Underneath it, the counts are thinner than the headline suggests.

And then the sentence that ought to have been the story:

There was no formal monitoring performed by the firms to quantify the audit quality impact of using ATTs.

The profession that exists to give independent assurance over other people's numbers had deployed the tools that produce its own evidence without measuring their effect on the quality of that evidence. Not badly measured. Not measured. This is the regulator's finding, in its own words, about the six firms that audit almost everything of size in the United Kingdom.

The obvious defences hold, and should be stated. The review examined governance rather than outcomes, so it cannot show that audit quality has fallen. It is a snapshot of processes at a moment when generative tools in these firms were still, in the FRC's description, limited to productivity aids such as chatbots rather than tools producing audit evidence. Two years on, that is unlikely to still be true, and the review has not been repeated.

Effective challenge needs somebody who could have done it by hand#

Put the two documents side by side and the same requirement appears in both, worded differently. SR 26-2 asks for expertise, independence and standing. The FRC asks the engagement team to notice when a tool is systemically wrong about this client. Both are asking a person to hold a view about work they did not perform.

That is possible, and this profession has done it for a century. It is possible for a specific reason: the reviewer did the work by hand earlier in their career. A partner who can smell a wrong revenue recognition judgement built that in years of tying out balances, recalculating accruals and testing samples that turned out to be fine. Those steps produced little of independent value. They were how the pattern library was built.

Those steps are also the ones automated first, which is the pattern this research calls missing rungs. The tasks that make the reviewer are the tasks the tool removes, and the cost of removing them appears about a decade later, in the quality of people who were supposed to have become reviewers. Meanwhile the person signing today already has the pattern library and cannot easily tell that the next cohort is not building one. That gap between apparent and actual capability is what this research calls synthetic seniority.

There is a second-order problem specific to this profession. Audit's response to almost any risk is documentation, and the FRC's guidance is a documentation standard. A file can record that a tool was appropriately explainable, that training covered automation bias, and that the team considered why an item was flagged, while nobody involved could have found the misstatement unaided. The file would pass inspection. That is the general form of the argument on why human in the loop is not a safeguard: a control that records attention is not a control that produces judgement.

The uncomfortable version, and the one this page will not soften: no measurement exists either way. The FRC says the firms have not quantified the audit quality impact of their tools. Nobody else has either. So the argument above is a mechanism with strong support from other professions and no direct test in this one, and it should be read at that strength.

What a firm or an audit committee can actually do#

Key sources

Every quotation above was read in the primary document. The two FRC documents carry only "June 2025" on their covers; the 26 June date comes from the FRC's own announcement of them.

On the oversight question in general, meaningful human oversight, why human in the loop is not a safeguard and who can override an AI system. On who carries the checking, who owns verification and the verifier's discount. On the capability underneath it, synthetic seniority, missing rungs and the capability audit. On the same question in other professions, law, medicine, consulting and journalism, and on the method, deskilling risk by profession.

About this research#

Rahim Hirji is the author of SuperSkills (Kogan Page, 2026), keynote speaker on AI and human capability, and founder of The SuperSkills Intelligence Company. SR 26-2 and its attachment, SR 11-7, and both FRC publications were read in full at their primary sources, and every quotation is verbatim from them. No figure appears anywhere on this page for AI adoption rates in accounting firms, changes in graduate intake, or the proportion of audit work now performed by automated tools, because no source was found that could support one. This is commentary on the regulatory and evidential position, not accounting, audit or investment advice.

How this research works  ·  Reviewed quarterly  ·  Found an error? Tell me and it is corrected on the page.

Cite this

Hirji, R. (2026). How will AI change accounting and audit? The SuperSkills Intelligence Company. Last reviewed 31 August 2026. thesuperskills.com/research/how-will-ai-change-accounting-and-audit

In this hub

Professions and sectors

Where the pressure lands first, profession by profession.

The work

Where the writing comes from.

These essays draw on research across more than 200 organisations in 30 countries. See the wider body of work, or bring it into your organisation.

All research →
Box of Amazing

Rahim’s free weekly letter on AI and human capability

If this was useful, the weekly letter is where the thinking happens first. Most of what ends up on this site starts there. Weekly essays on AI, capability and the future of work. Read by 25,000 people, every week since 2017. Free, and one click to stop.

Opens Substack to confirm. No pitch in it, unsubscribe in one click, and nobody follows up because you read something.

Running an event, or responsible for how AI arrives in your organisation? Keynotes  ·  Advisory and coaching  ·  Enquire