← Research
Research

Can I trust an AI chatbot for financial advice?

One vendor's test of 18 chatbots on pensions, tax and debt, why the hard questions fail most, the accountability gap the figures point at, and five rules for a money question put to a machine.

Last reviewed: 21 September 2026

Not for a decision with money on it. A UK test of 18 models against 121 personal finance questions, published 14 September 2026, found the answers wrong 57 per cent of the time and 88 per cent on the harder ones. The tester sells AI to advisers. The finding that matters is that nobody is accountable for the answer. An evidence review by Rahim Hirji; every figure resolves to a graded entry in the evidence base that says what it does not show.

Questions this page answersAll 843 questions this research covers

Not for a decision with money on it. On 14 September 2026 Saturn, a UK company that sells AI software to advice firms, published a test of 18 popular models against 121 personal finance questions. Across more than 10,000 answers the models were wrong 57 per cent of the time, and 88 per cent of the time on the harder questions. The errors were confident figures, invented rules and the wrong order for paying debts. The tester has an interest in the result and one study settles nothing on its own. What it shows is the shape of the risk: a fluent answer with nobody accountable for it. The useful question is not whether to trust the chatbot but what you must still be able to check, and who carries the loss when nobody checked.

The answer, in one line

Not for a decision that moves money. A UK test of 18 models on 121 personal finance questions found the answers wrong 57 per cent of the time and 88 per cent on the harder ones, and no chatbot answer is regulated advice, so nobody is accountable when it is wrong.

Share as a card

What the test did, and who ran it#

The report is called Artificial Authority: Should you trust AI to deliver financial advice? Saturn describes itself as an AI platform for advice firms and says more than 500 firms and about 3,000 advisers use it, so a finding that consumer chatbots get money questions wrong supports the company’s own market. The method, as reported by Professional Adviser, IFA Magazine and Financial Reporter on 14 September, was 121 questions put to 18 free and paid models, each asked up to five times, each answer scored against Saturn’s criteria. BeInCrypto reported on 20 September that an answer failed if it contained a factual error, omitted material information or lacked a required warning. The Financial Times carried the study the same day; that report is behind a paywall and was not read here.

Free models were wrong 63 per cent of the time and paid models 49 per cent. The per-model range ran from 39 to 82 per cent wrong; the best and worst scores came from two models made by the same developer, so the spread is between models and modes of use rather than between companies. Saturn’s chief executive, Amal Jolly, said that “millions of people are trusting the AI models for money advice, but they are getting wrong answers that can lose them money”.

Why the hard questions failed most#

The examples share a shape. A pension answer miscalculated a tax charge in a way that would have exposed a saver to a potential £17,500 bill from HMRC. A debt answer said to clear the highest-interest debt first, ahead of the bills whose non-payment costs a home or a power supply. One answer invented a student loan rule. Another misdescribed what a mortgage payment holiday does to a credit record. None is a failure of arithmetic. Each is a failure to hold a rule that changes by year, by threshold and by sequence, and to know that it changes. A model produces the shape of a correct answer from the text it was trained on, and the shape of a pension answer is the same whether the allowance inside it is this year’s or one from three years ago. Confidence is a property of the writing, not the knowledge, so the wrong answer and the right one read identically. The harder the question, the more rules it chains together, and the more places the chain can break without the prose showing it.

So the reader cannot use the answer to check the answer. The reliable signal is knowing the failure patterns of the domain, and most people enter personal finance once a decade. The evidence on algorithm appreciation adds the uncomfortable half: given identical advice under a human label and a machine label, people took more of it from the machine.

The accountability gap is the finding that matters#

In the UK, telling a person which pension, investment or mortgage product to take is a regulated activity. A firm doing it must be authorised by the Financial Conduct Authority, must assess whether the advice suits that person, and answers to the Financial Ombudsman Service when a client complains. A chatbot’s answer carries none of that. It is not advice in the regulated sense, the provider’s terms say so, and no adviser signed it. The number to hold onto is not 57 per cent. It is zero: the number of people accountable for an answer that cost a saver £17,500. This is the missing record of where a decision came from: money moves, and no line anywhere says who decided, on what, and who checked.

What it means for an adviser and for an advice firm#

Clients now arrive with an answer. The workable stance is the one this estate applies to every AI output: treat it as a draft position the client has formed, and do the checking in front of them, so the client learns what checking looks like. The same test then turns on the firm. Saturn sells AI that drafts suitability letters and automates file checks, and a study showing consumer chatbots getting rules wrong is a study of the same kind of system, pointed at the same rules. Who owns verification when the machine does the work has to be answered in writing: which adviser reads the machine’s letter against the client file, what they check for, and whether they could still write the letter themselves in three years. A firm whose advisers can no longer produce the work they sign has bought the failure the study measured, at the professional’s desk rather than the kitchen table.

Five rules for a money question put to a machine#

Use it to learn the vocabulary and the questions, not the numbers; a model is good on what a payment holiday is and unreliable on what applies to you this tax year. Ask it for the rule, the year it applies to and the source, then read the source on gov.uk. Treat every figure as unverified until you have seen it somewhere a person is accountable for it. Do not act on the order in which to pay debts without a free regulated source; MoneyHelper, StepChange and Citizens Advice exist for this. And any decision with a penalty, a tax charge or a lost protection attached goes to a regulated adviser before the money moves, because the cost of being wrong is not the fee you saved.

Where this sits in the argument#

The risk that can be measured sits in the handover of decisions to machines and in what happens to human judgement afterwards. A chatbot answering a pension question is that handover at its most ordinary: nobody decided it, nobody monitors it, and the person asking cannot tell a right answer from a wrong one. The Rules Before Tools questions apply to a household as much as a board: which decisions a machine may make, who can stop each one, what the person must remain able to do, and how anyone would know it went wrong. For money the answers are: none that move funds; you; checking a rule against its source; and a letter from HMRC. The study will be superseded. The gap it points at will not be, until someone is accountable for the answer.

What this does not show#

It does not show that chatbots are wrong 57 per cent of the time about money in general. The 121 questions were chosen by a company with a commercial interest in the result, the scoring criteria have not been published, the report sits behind a form on the company’s site, nobody outside the company has re-run the test, and no peer reviewer has seen it. The questions are UK-specific. “Complex” is not defined in the coverage, and the 88 per cent depends on that definition. The models were tested in one window, so the per-model figures are not a ranking of companies. It measures single answers, not what happens when a person cross-checks or takes the answer to an adviser, which is what decides whether harm occurs. And the accountability point is an argument from the structure of UK regulation, not a count of people harmed; no regulator has published a figure for losses traced to chatbot answers, and this page will be updated when one does.

Essay · SS-2026-285

Cite this page

Hirji, R. (2026). Can I trust an AI chatbot for financial advice?. The SuperSkills evidence base, SS-2026-285. https://thesuperskills.com/research/can-i-trust-an-ai-chatbot-for-financial-advice. Last reviewed 21 September 2026.

An evidence review by Rahim Hirji, not peer-reviewed research. For a material claim, cite the underlying study as well; every study here carries its own permanent link.

How citations and IDs work
Questions answered on this page

Can I trust an AI chatbot for financial advice?

Not for a decision that moves money. A UK test of 18 models on 121 personal finance questions found the answers wrong 57 per cent of the time and 88 per cent on the harder ones, and no chatbot answer is regulated advice, so nobody is accountable when it is wrong.

How often do AI chatbots get financial questions wrong?

In Saturn's Artificial Authority study, published 14 September 2026, 18 models answered 121 UK questions on pensions, tax, debt and savings up to five times each. Answers were wrong 57 per cent of the time overall, 63 per cent for free models, 49 per cent for paid ones, and 88 per cent on average for the harder questions. Saturn sells AI to advice firms and its scoring criteria are not public, so the figures are one vendor's test and not a base rate.

Is a chatbot's answer regulated financial advice?

No. In the UK, recommending a pension, investment or mortgage product is a regulated activity that only an FCA-authorised firm may carry out, with a duty to assess suitability and a route to the Financial Ombudsman if it goes wrong. A chatbot's answer carries none of that, the provider's terms say so, and no adviser signed it.

How should I use AI for money questions?

Use it to learn the vocabulary and the questions, not the numbers. Ask it for the rule, the year it applies to and the source, then read the source on gov.uk. Treat every figure as unverified until a person accountable for it has confirmed it, use MoneyHelper, StepChange or Citizens Advice for debt order, and take any decision with a penalty or tax charge attached to a regulated adviser before the money moves.

In this hub

Everyday life

The same questions, asked about your own week rather than your organisation.

Ask the evidence
What does the evidence actually show?What should our board be asking about this?Where does Rahim disagree with the consensus?
Bring this into your organisation

If this describes something happening in your teams, say so.

Keynotes, board sessions and advisory work, drawing on research across more than 200 organisations in 30 countries. Tell me the room, the date and the shift you need. A reply within 24 hours.

Start a conversation

Topics and audiences  ·  All research

For every money question a client or an employee now puts to a chatbot, which answers the firm will check, who checks them against the rule and its year, and whether the advisers signing machine-drafted letters could still write them. Writing that verification line into the advice process, and naming who may halt on it, is the engagement. AI advisory for CEOs and boards.

This is the question underneath the rest of them, and the one a leadership team is least likely to have named out loud. There is the judgement version, and the full range of topics and audiences.

Box of Amazing

Rahim’s free weekly letter on AI and human capability

If this was useful, the weekly letter is where the thinking happens first. Most of what ends up on this site starts there. Weekly essays on AI, capability and the future of work. Read by 25,000 people, every week since 2017. Free, and one click to stop.

Opens Substack to confirm. No pitch in it, unsubscribe in one click, and nobody follows up because you read something.