- Why does AI get harder questions wrong more often than easy ones?
- Can I trust an AI chatbot for financial advice?
- Is a chatbot's answer regulated financial advice?
Not for a decision with money on it. On 14 September 2026 Saturn, a UK company that sells AI software to advice firms, published a test of 18 popular models against 121 personal finance questions. Across more than 10,000 answers the models were wrong 57 per cent of the time, and 88 per cent of the time on the harder questions. The errors were confident figures, invented rules and the wrong order for paying debts. The tester has an interest in the result and one study settles nothing on its own. What it shows is the shape of the risk: a fluent answer with nobody accountable for it. The useful question is not whether to trust the chatbot but what you must still be able to check, and who carries the loss when nobody checked.
The answer, in one line
Not for a decision that moves money. A UK test of 18 models on 121 personal finance questions found the answers wrong 57 per cent of the time and 88 per cent on the harder ones, and no chatbot answer is regulated advice, so nobody is accountable when it is wrong.
What the test did, and who ran it#
The report is called Artificial Authority: Should you trust AI to deliver financial advice? Saturn describes itself as an AI platform for advice firms and says more than 500 firms and about 3,000 advisers use it, so a finding that consumer chatbots get money questions wrong supports the company’s own market. The method, as reported by Professional Adviser, IFA Magazine and Financial Reporter on 14 September, was 121 questions put to 18 free and paid models, each asked up to five times, each answer scored against Saturn’s criteria. BeInCrypto reported on 20 September that an answer failed if it contained a factual error, omitted material information or lacked a required warning. The Financial Times carried the study the same day; that report is behind a paywall and was not read here.
Free models were wrong 63 per cent of the time and paid models 49 per cent. The per-model range ran from 39 to 82 per cent wrong; the best and worst scores came from two models made by the same developer, so the spread is between models and modes of use rather than between companies. Saturn’s chief executive, Amal Jolly, said that “millions of people are trusting the AI models for money advice, but they are getting wrong answers that can lose them money”.
Why the hard questions failed most#
The examples share a shape. A pension answer miscalculated a tax charge in a way that would have exposed a saver to a potential £17,500 bill from HMRC. A debt answer said to clear the highest-interest debt first, ahead of the bills whose non-payment costs a home or a power supply. One answer invented a student loan rule. Another misdescribed what a mortgage payment holiday does to a credit record. None is a failure of arithmetic. Each is a failure to hold a rule that changes by year, by threshold and by sequence, and to know that it changes. A model produces the shape of a correct answer from the text it was trained on, and the shape of a pension answer is the same whether the allowance inside it is this year’s or one from three years ago. Confidence is a property of the writing, not the knowledge, so the wrong answer and the right one read identically. The harder the question, the more rules it chains together, and the more places the chain can break without the prose showing it.
So the reader cannot use the answer to check the answer. The reliable signal is knowing the failure patterns of the domain, and most people enter personal finance once a decade. The evidence on algorithm appreciation adds the uncomfortable half: given identical advice under a human label and a machine label, people took more of it from the machine.
The accountability gap is the finding that matters#
In the UK, telling a person which pension, investment or mortgage product to take is a regulated activity. A firm doing it must be authorised by the Financial Conduct Authority, must assess whether the advice suits that person, and answers to the Financial Ombudsman Service when a client complains. A chatbot’s answer carries none of that. It is not advice in the regulated sense, the provider’s terms say so, and no adviser signed it. The number to hold onto is not 57 per cent. It is zero: the number of people accountable for an answer that cost a saver £17,500. This is the missing record of where a decision came from: money moves, and no line anywhere says who decided, on what, and who checked.
What it means for an adviser and for an advice firm#
Clients now arrive with an answer. The workable stance is the one this estate applies to every AI output: treat it as a draft position the client has formed, and do the checking in front of them, so the client learns what checking looks like. The same test then turns on the firm. Saturn sells AI that drafts suitability letters and automates file checks, and a study showing consumer chatbots getting rules wrong is a study of the same kind of system, pointed at the same rules. Who owns verification when the machine does the work has to be answered in writing: which adviser reads the machine’s letter against the client file, what they check for, and whether they could still write the letter themselves in three years. A firm whose advisers can no longer produce the work they sign has bought the failure the study measured, at the professional’s desk rather than the kitchen table.
Five rules for a money question put to a machine#
Use it to learn the vocabulary and the questions, not the numbers; a model is good on what a payment holiday is and unreliable on what applies to you this tax year. Ask it for the rule, the year it applies to and the source, then read the source on gov.uk. Treat every figure as unverified until you have seen it somewhere a person is accountable for it. Do not act on the order in which to pay debts without a free regulated source; MoneyHelper, StepChange and Citizens Advice exist for this. And any decision with a penalty, a tax charge or a lost protection attached goes to a regulated adviser before the money moves, because the cost of being wrong is not the fee you saved.
Where this sits in the argument#
The risk that can be measured sits in the handover of decisions to machines and in what happens to human judgement afterwards. A chatbot answering a pension question is that handover at its most ordinary: nobody decided it, nobody monitors it, and the person asking cannot tell a right answer from a wrong one. The Rules Before Tools questions apply to a household as much as a board: which decisions a machine may make, who can stop each one, what the person must remain able to do, and how anyone would know it went wrong. For money the answers are: none that move funds; you; checking a rule against its source; and a letter from HMRC. The study will be superseded. The gap it points at will not be, until someone is accountable for the answer.
What this does not show#
It does not show that chatbots are wrong 57 per cent of the time about money in general. The 121 questions were chosen by a company with a commercial interest in the result, the scoring criteria have not been published, the report sits behind a form on the company’s site, nobody outside the company has re-run the test, and no peer reviewer has seen it. The questions are UK-specific. “Complex” is not defined in the coverage, and the 88 per cent depends on that definition. The models were tested in one window, so the per-model figures are not a ranking of companies. It measures single answers, not what happens when a person cross-checks or takes the answer to an adviser, which is what decides whether harm occurs. And the accountability point is an argument from the structure of UK regulation, not a count of people harmed; no regulator has published a figure for losses traced to chatbot answers, and this page will be updated when one does.
Essay · SS-2026-285
Hirji, R. (2026). Can I trust an AI chatbot for financial advice?. The SuperSkills evidence base, SS-2026-285. https://thesuperskills.com/research/can-i-trust-an-ai-chatbot-for-financial-advice. Last reviewed 21 September 2026.
An evidence review by Rahim Hirji, not peer-reviewed research. For a material claim, cite the underlying study as well; every study here carries its own permanent link.
How citations and IDs work