
A striking headline appeared in Financial Reporter:
“AI models give wrong financial advice 57% of the time.”
The claim came from research commissioned by financial technology company Saturn. It reportedly tested 18 AI models using 121 financial questions, repeated five times, producing more than 10,000 responses.
Across those responses, 57% were classified as wrong. On harder questions, the reported error rate rose to 88%, with some models said to have failed 99% of complex queries.
Those figures deserve attention.
AI systems can make mistakes. They can misunderstand context, omit important warnings and state incorrect information with unwarranted confidence. When people are dealing with pensions, debt, tax or mortgages, those mistakes can cause serious harm.
But before accepting the headline—and Saturn’s proposed regulatory solution—we should ask a better question:
What, exactly, was being tested?
Financial information became “financial advice”
The reported examples of AI failure included:
- incorrect information about pension tax rules;
- recommending that borrowers repay the highest-interest debts before priority bills such as rent and council tax;
- inventing a rule about stopping student-loan repayments by moving abroad; and
- claiming that a mortgage payment holiday would not affect someone’s credit record.
These are potentially serious errors. But they are not examples of regulated financial advice in the conventional UK sense.
None appears to involve a personalised recommendation to buy, sell or hold a particular investment.
They are examples of answers to financial questions—covering tax, debt, student finance and mortgage administration.
That distinction matters.
By describing every incorrect answer to a financial question as “wrong financial advice”, the headline performs a subtle category shift:
Wrong answers about money become wrong financial advice.
The broader category makes the research sound more alarming. It also helps support Saturn’s argument that the FCA should regulate AI-generated financial advice.
But the examples reported do not establish that AI systems are routinely making unsuitable regulated investment recommendations.
Financial information, generic guidance, financial planning and regulated financial advice are related activities. They are not interchangeable.
A genuine risk does not prove the proposed remedy
Saturn’s central concern is credible.
People may rely upon general-purpose AI systems without understanding their limitations. They may provide incomplete information, ask poorly framed questions or treat a confident answer as authoritative.
That creates real risk.
But evidence of a risk is not proof that a particular regulatory remedy is correct.
The research may show that unsupported, one-shot AI answers can be unreliable. It does not automatically show that every AI response concerning money should be treated as regulated financial advice.
Nor does it establish that the best response is to preserve dependence upon conventional advisers.
There is a missing step in the argument:
- AI systems sometimes give wrong answers.
- Wrong answers can cause consumer harm.
- Therefore, AI-generated financial advice should be regulated by the FCA.
The first two propositions may be reasonable. The third requires separate evidence.
Regulation might form part of the answer. But so might better AI design, clearer boundaries, source verification, financial education, structured questioning and escalation to appropriately qualified humans when the consequences or regulatory perimeter require it.
AI making mistakes is evidence for better decision architecture—not automatically evidence for preserving adviser dependency.
Who checks the answer key?
Saturn reportedly marked responses as failures where they contained factual errors, missed key points or omitted important warnings.
That creates another question:
How do we know Saturn’s answers were right?
To assess the 57% claim properly, we would need to examine:
- all 121 questions;
- the wording of the prompts;
- any consumer circumstances supplied;
- Saturn’s model answers;
- how “easy” and “complex” questions were defined;
- how factual mistakes were distinguished from omissions;
- whether all failures carried equal weight;
- who adjudicated ambiguous answers;
- whether independent qualified reviewers validated the scoring;
- which versions of the AI models were used;
- whether reasoning or research functions were enabled; and
- whether models were permitted to ask clarifying questions.
A dangerously incorrect answer is not equivalent to an answer that is broadly accurate but omits a warning. Yet both may have been recorded within the same “wrong” category.
Without the complete methodology, the headline figure may partly mean:
The AI did not reproduce Saturn’s preferred answer.
That is not necessarily the same as proving that the answer was factually wrong, unsuitable or harmful.
This is the problem of benchmark authority.
Saturn is judging the AI. But who judged Saturn?
The study compares AI with an invisible ideal
The research appears to compare actual AI responses against a supposedly correct answer key.
It does not appear to compare AI with the real alternatives available to consumers.
Those alternatives may include:
- receiving no help;
- searching social media;
- following online misinformation;
- acting on product advertising;
- accepting conflicted sales guidance;
- asking friends or family;
- consulting a human adviser working from incomplete information; or
- making an unaided decision while under financial stress.
The relevant question is not merely:
Does AI make mistakes?
Of course it does.
The more useful question is:
Compared with what people would otherwise do, and under what conditions, does AI increase or reduce harm?
That requires a comparative study.
Give the same questions and circumstances to general-purpose AI models, specialist AI systems and human advisers. Then examine not only whether their answers match a predetermined script, but also whether they recognise uncertainty, identify missing information and preserve the person’s ability to decide.
We might discover that human advisers perform better.
We might also discover substantial disagreement between advisers, gaps in technical knowledge, commercial bias and equally important omissions.
Without that comparison, the research tells us how often AI departed from Saturn’s benchmark. It does not tell us how often humans give wrong financial advice.
What percentage of human advice is wrong?
This is the uncomfortable question missing from the debate.
How often do human advisers:
- misunderstand a client’s circumstances;
- rely on incomplete information;
- overlook tax or benefit interactions;
- fail to explore credible alternatives;
- recommend products that fit their business model;
- allow charging structures to influence their solution;
- confuse regulatory compliance with good planning;
- prioritise investable assets over the client’s wider life; or
- create long-term dependency where capability could have been built?
Human advisers are not truth machines either.
They possess judgment, empathy and contextual understanding that general-purpose AI may lack. But they are also affected by incentives, habits, limited knowledge, time constraints, institutional expectations and their own mental models.
A regulated recommendation can be compliant and still be wrong for the person.
It can be product-correct while being life-wrong.
Is the industry measuring the right wealth?
This leads to the largest omission.
The reported study evaluates answers within the conventional financial-services understanding of wealth. The focus is tax, debt, pensions, mortgages and financial products.
But financial capital is only one part of a person’s wealth.
Total Wealth Planning also recognises:
- human capital — skills, knowledge, health, energy and future earning capacity;
- social capital — relationships, family, trust, community and mutual support;
- environmental capital — the quality, security and sustainability of the places in which we live; and
- spiritual capital — meaning, purpose, values, identity and contribution.
For many people, particularly earlier in life, human capital is their greatest asset.
A recommendation to maximise pension contributions might be technically defensible while leaving someone without the liquidity to retrain, leave harmful employment or establish a business.
An investment recommendation might satisfy a risk questionnaire while conflicting with the person’s values.
An instruction to minimise expenditure might improve a cashflow projection while damaging health, relationships or quality of life.
A retirement calculation might identify the required financial capital while ignoring the loss of identity, structure, community and purpose that can accompany retirement.
If the benchmark measures only financial capital, an answer can score as correct while diminishing total wealth.
The answer may be financially right and humanly wrong.
An alternative test for good financial support
A more meaningful study would compare AI, human advisers and unaided consumer decision-making using independently validated criteria.
It would examine:
- Factual accuracy
Are the rules, calculations and material facts correct? - Regulatory accuracy
Does the response recognise when regulated advice or another qualified professional is required? - Completeness
Are important risks, options and consequences identified? - Contextual suitability
Does the answer reflect the person’s actual circumstances rather than an assumed average consumer? - Whole-person relevance
Does it consider human, social, environmental and spiritual wealth alongside financial capital? - Recognition of uncertainty
Does the system or adviser disclose what is unknown and avoid unjustified certainty? - Quality of questioning
Does it ask for the information needed before reaching a conclusion? - Conflicts of interest
Is the answer affected by products, fees, referrals, business models or regulatory positioning? - Consumer understanding
Can the person understand the reasoning and challenge its assumptions? - Preservation of agency
Does the support build the person’s capacity to choose—or encourage them to surrender the decision?
This would test more than whether an answer agrees with an answer key.
It would test whether the support helps someone make a better decision.
From answer accuracy to decision architecture
Consumers should not be encouraged to accept unverified AI output.
But neither should they be encouraged to assume that a regulated human automatically holds the correct answer.
The safer model is not blind faith in either humans or machines.
It is a decision architecture that combines:
- structured fact-finding;
- transparent assumptions;
- credible sources;
- alternative perspectives;
- explicit uncertainty;
- checks for commercial and institutional incentives;
- whole-person planning; and
- access to episodic human expertise when complexity, consequence or regulation requires it.
This is the role of Academy OS.
Its purpose is not to replace human judgment with machine authority. It is to help people understand their position, discover better questions, examine alternatives and recognise when expert help is needed.
The future is not simply human advice versus AI advice.
It is:
Continuous human agency, supported by AI capability and episodic human expertise.
BIG Checker is not a truth machine
There is an enjoyable irony in using AI to scrutinise research warning people about the unreliability of AI.
But that irony also reveals the intended role of BIG Checker.
BIG Checker does not declare Saturn’s research true or false. It does not claim that AI is safe, infallible or inherently superior to human advisers.
It helps the reader ask:
- Who produced this evidence?
- How was “wrong” defined?
- What was actually tested?
- What was placed inside the same category?
- What information is missing?
- What interests may shape the conclusion?
- What response is the reader being encouraged to support?
- Does the evidence justify that response?
- What alternatives have been excluded?
- What would change my mind?
BIG Checker is deliberately something different:
Not the truth machine. The better-question machine.
The wrong answer—or the wrong measure?
Saturn’s research may have identified real and important weaknesses in general-purpose AI.
But its findings do not appear to establish that 57% of regulated financial advice generated by AI is wrong. The examples reported are not regulated investment recommendations. The complete methodology and answer key require scrutiny. Human performance was not used as a comparator. And the conventional model of financial correctness omits much of what makes a life wealthy.
Before asking whether AI gives the wrong financial advice, we should therefore ask whether the financial-services industry has been measuring the right kind of wealth.
Because if financial capital is the only thing being measured, the industry may be producing technically correct answers to the wrong question.
And faster, regulated or more confidently delivered wrong questions do not restore human agency.
They merely make dependency more efficient.
