Who validated the warning about AI and your money?

An evidence audit of a financial AI headline, its commercial context and the poll that followed
Steve Conley | Academy of Life Planning | 23 September 2026

I watched a poll ask people how much of their money they would let AI decide with. Before they answered, it told them that chatbots had got 57% of financial answers wrong, and 88% of harder ones. Sixty-nine per cent then chose “nothing”.

The poll asked four questions:

QuestionWhat it asked respondents to judge
1How much money they would trust an AI chatbot to decide with, after hearing the 57% figure.
2Which financial task they would trust AI to tackle first, after hearing the 88% figure. The choices included salary negotiation, taxes, investments and mortgages.
3Whether they would trust AI financial advice more if a qualified human adviser approved every answer.
4Whose answer they would trust most for a money decision that night: several sources, an adviser, official guidance, a knowledgeable friend, AI or a finance creator.

Here is what the screenshots actually show:

Question asked after respondents saw the 57% and 88% error claimsVisible result
How much money would you trust AI to decide with?69% chose “nothing”; 15% chose up to £10.
Which money job would you trust AI to tackle first?61% chose “none of these”; 16% chose salary negotiation.
Would a qualified human approving every AI answer increase trust?48% said “a little more” or “much more”; 21% said they still would not trust it with financial decisions.
Whose answer would you trust most for a money decision tonight?29% chose researching several sources; 23% a regulated bank adviser; 22% an independent adviser; 12% official guidance; 11% a financially savvy friend; 2% AI alone.

The result invites a simple headline: people do not trust AI with financial planning. But the poll asked whether they would let AI decide what to do with their money, after telling them it frequently gets financial answers wrong. It never asked whether they would use AI to understand their position, explore their options and make a better-informed decision themselves. And the figures used to frame those questions came from research whose methods still need independent scrutiny.

I have a stake in this question. At the Academy of Life Planning, we are building a client operating system: people hold their own information, use AI to investigate and plan, and call on human expertise when needed. Saturn, the company behind the research, sells AI software to advice firms. It says its technology helps those firms save time and grow assets under management. Both commercial positions belong in view. Neither determines whether a finding is true. Saturn’s account of its business

Here is the sequence as far as the available evidence permits us to trace it.

First, a test of chatbot answers

Saturn’s report page says its team tested 18 models on 121 questions and found mistakes in 57% of answers on average. The Financial Times report of 19 September says the questions were repeated up to five times, producing more than 10,000 responses; harder questions had an 88% error rate. Reported examples include pension tax, student loans, debt and mortgages. Some errors sound serious. They should be investigated, not dismissed.

But what exactly counts as an error? Was a false rule scored the same as an omitted warning? How many questions concerned a determinate fact, and how many required context or professional judgement? Were models allowed to ask for missing facts, use current sources or check calculations? Who wrote the model answers, and who independently reviewed disputed scores?

I could find the public summary and reporting, but not a publicly accessible complete question bank, answer key, response set and scoring protocol. Saturn offers the report through a form. That means this article cannot reproduce or independently validate the 57% or 88% figures. It also means I cannot say the figures are false. The proper status of the headline is a reported result from a company test, pending independent audit.

Then, a result across changing models

The FT reports that newer models did better than older models in Saturn’s test; it identifies Claude Haiku 4.5 in one adverse example and Claude Opus 5 in reasoning mode as the best performer, with a reported 39% error rate. That variation matters. An average across 18 models is not the error rate of the model a particular person uses today, still less a prediction about the model they will use next year. FT coverage

A reproducible benchmark would identify every product and exact version, access tier, mode, tools enabled, test date and settings, and publish results separately for each. It would also retest after significant releases. I have not verified that Saturn used obsolete models throughout. Without its full dated model inventory, that claim would outrun the evidence. The defensible point is that the pooled headline conceals differences between versions and cannot remain a timeless statement about “AI”.

Then, financial questions became “financial advice”

A wrong answer about student loan rules can harm someone. It does not follow that the model delivered regulated investment advice. The FCA’s PERG 8.26.2G even lists financial planning as a possible example of generic advice outside article 53(1). Different activities, including advice on particular investments and regulated mortgages, have different boundaries.

The point is not to minimise the error. It is to identify what was tested. A benchmark of answers about money cannot, without further classification, establish a failure rate for personalised regulated recommendations. Nor can it establish that an adviser-led service is the best remedy. Those are separate questions.

Then, the invisible comparator

The reported benchmark does not supply a matched group of advisers or relevant experts answering the same questions under comparable conditions. Without that comparator, it cannot tell us whether human answers would be more accurate, what they would omit, or how often they would ask for clarification.

That does not mean humans would perform worse. It means the study has not measured the comparison. It also has not measured an AI workflow that checks sources, records uncertainty, invites corrections and escalates consequential matters to a specialist. A one-shot chatbot answer is one system design, not the whole of AI-supported planning.

The relevant consumer question is: compared with the help a person could actually obtain, which arrangement leaves them better informed, safer and more able to choose?

Finally, a poll built on the headline

The screenshots sent to me show a poll headed “ChatG… Wrong”. Its introduction states the 57% and 88% figures, then asks how much money the respondent would put in AI’s hands. The first question repeats the 57% claim and asks how much money the respondent would trust AI to decide with. Sixty-nine per cent select “nothing”. The second repeats the 88% claim and asks which money job AI should tackle first; 61% select “none of these”.

Another question asks whether a qualified human approving each answer would increase trust. In the final question, after respondents have seen the error figures, 29% choose researching several sources, 23% a regulated bank adviser, 22% an independent adviser and 2% AI as the single answer they would trust most.

Those are the visible responses to those questions. The screenshots do not show the sample size, recruitment, weighting, field dates or whether the poll’s publisher has any relationship with Saturn. I therefore cannot call the poll representative of Britain, attribute it to Saturn, or measure how many opinions the introduction changed.

But the sequence within the poll is observable: the error claims were supplied before the trust questions. Survey researchers warn that framing, preceding text and question order can affect responses. To estimate that effect here, we would need a randomly assigned comparison group receiving a neutral introduction, plus a question about using AI to inform a decision while the person remains in charge. Pew on question wording and order · AAPOR on disclosing preceding text and methods

One result should give us pause before declaring people unwilling to use AI: the largest group in the final question wanted to consult several sources because none deserved complete trust alone. That is compatible with a person checking AI alongside other sources. The poll did not test that possibility.

What other evidence says

Different questions yield different findings. An FCA survey of 666 UK adults aged 18–40 who invest or might invest found that two-thirds expected to use AI more for investment decisions in the following year; 73% recognised that AI can be inaccurate and 86% understood the need to check its sources. This is a defined subgroup, not all UK adults. The FCA itself describes AI as useful for researching companies, understanding jargon and exploring options while the person retains judgement.

Separately, an FCA-commissioned survey of more than 5,000 retail financial services consumers found 20% likely to use AI that can act autonomously within preset goals. Its population and question differ from the smaller FCA study and the screenshots. We should not pretend these percentages are interchangeable. They do show why “no one trusts AI” is untenable.

The harm worth testing

An unreliable AI answer may cost someone money. A sweeping warning may also discourage someone from seeking information, checking an adviser’s assumptions, or organising their affairs when professional help is inaccessible. Neither harm should be assumed away. Neither has been quantified by this sequence of research and polling.

Saturn’s business interest is relevant: its product makes existing advice firms more efficient, and its website speaks of helping them grow AUM. The observed evidence does not establish that its researchers falsified scores or intended to make consumers dependent. The stronger concern is structural: research about unsupported AI answers is being carried into a public choice between trusting a chatbot with decisions and relying on an adviser, while a third arrangement—client-owned AI with proportionate human expertise—goes untested.

This is the question I would put to Saturn, journalists and anyone citing the poll as a verdict on the future of planning:

Publish the dated model list, all prompts and answers, the scoring rubric and adjudication record. Run the same tasks with relevant human experts. Then test a supported, client-owned workflow against adviser-led and unaided decisions. Measure errors, costs, access, understanding and who retains control.

I would welcome a correction to this audit if those materials are available. Good research should survive the same scrutiny it asks us to apply to AI.

The decision is too important to leave to either a frightening headline or a comforting counterclaim. We need to know not merely who gave the right answer, but who helped the person recognise, check and own a better decision.


Evidence note

This article was prepared on 23 September 2026 from Saturn’s public report page and product descriptions, FT coverage, FCA publications, AAPOR and Pew methodology guidance, and six screenshots of the poll supplied to the author. The full Saturn dataset and poll technical report were not available for independent inspection. Percentages from the screenshots are shown as displayed, without a known denominator or uncertainty estimate. Claims about causal effects and intention are identified as hypotheses requiring further evidence.

Leave a comment