The Trust Deficit: Why AI-Enabled Financial Advice Needs Better Regulatory Oversight
Nearly half of consumers already ask AI what to do with their money, yet the supervisory toolkit still assumes a human is giving the advice. Here is where oversight has to go next.
Published
Nearly half of consumers already ask an AI what to do with their money. The oversight built to protect them still assumes a human is on the other end of the conversation. At Rayify we spend our days stress-testing how automated systems reach a recommendation, and the gap between how fast people trust AI advice and how well anyone can check it is the widest we have seen in a regulated market. This post is about that gap: why AI financial advice fails in ways that are hard to see, why the current supervisory toolkit cannot catch those failures, and where we think oversight has to go next.
The adoption is already here
The debate about whether people should use AI for financial decisions is over. They already do. EY's 2026 survey of more than 18,000 consumers across 23 countries found that nearly half of global consumers now use AI to guide savings and investment decisions, a figure that rises to 68% among Gen Z. The most-used channel is not a regulated robo-adviser but a general-purpose chatbot. OpenAI reports that more than 200 million people come to ChatGPT every month for budgeting help and questions about their investments.
So the question is not whether AI shapes financial decisions. It is whether we can trust how that advice is generated. Here the honest answer is uncomfortable: confidence is not competence. A model can produce a fluent, personalised pension recommendation that reads exactly like expert guidance and is wrong in ways neither the consumer nor the firm can detect. That is the trust deficit: adoption has raced ahead while the ability to verify the advice has stayed flat.
The reason the gap is so hard to close is structural. Traditional consumer-protection mechanisms assume the thing being supervised holds still long enough to inspect it. An AI advice system does not. It is a moving target running at a volume no supervisor has ever had to police, and its failures are engineered, by the very fluency that makes it useful, to look like successes.

The hard part is not hallucination in isolation
It is tempting to reduce the problem to "AI hallucinates" or "AI is biased." Both are true, but the harder issue is that these systems fail quietly. A loud failure, a garbled answer or an obvious factual error, is self-policing, because the consumer notices and discounts it. The failures that matter are the ones that survive inspection. Four of them matter more than raw error rates, because each is difficult to observe from the outside even when you are looking straight at the transcript.
They assume instead of asking. A suitability conversation is supposed to gather a client's circumstances before recommending anything. A model will often fill the gaps itself. Consider a concrete case: a user types "I have GBP 20,000, what should I do with it?" A compliant process stops and asks about time horizon, existing debt, emergency savings, and attitude to loss. A model under mild conversational pressure to be helpful instead infers a moderate risk appetite and a ten-year horizon the user never stated, then produces a growth-tilted portfolio that looks perfectly suitable, but only against facts it invented. The recommendation is internally coherent and externally unfounded, and nothing in the output flags that its premises were fabricated rather than elicited.
They skip compliance checks silently. Nothing in a fluent answer tells you whether a required step was actually performed. The output looks identical whether the model checked a rule, checked it wrongly, or ignored it entirely. Compliance, in a conversational system, is invisible unless you instrument for it, and the instrument does not exist by default.
They produce reasoning theatre. The explanation a model gives for its answer often does not reflect the computation that produced it. Turpin and colleagues showed that language models do not always say what they think, generating plausible chain-of-thought rationales that systematically omit the real influences on an answer. Anthropic's follow-up work found that even dedicated reasoning models frequently fail to mention the hints that actually changed their answer, and its earlier study on measuring faithfulness in chain-of-thought reasoning showed that a stated rationale can be post-hoc decoration rather than the cause of the output. For a regulator that relies on an audit trail, an explanation that is not faithful to the decision is worse than no explanation at all: it launders an unexamined process as a documented one.
They answer the same question differently on phrasing alone. Sclar and colleagues quantified how large language models are sensitive to spurious features in prompt design: a trivial change in formatting or wording can swing an answer across a wide range. Two clients with identical circumstances who phrase the same question slightly differently can receive materially different suitability assessments. The system is not just occasionally wrong. It is non-deterministic in a way that defeats the whole premise of "treat like cases alike."
None of these is visible in a single transcript that looks polished. That is precisely why they are dangerous at scale, and why sampling a few thousand pretty transcripts tells a supervisor almost nothing about the failures that are actually occurring.
The problem space regulators now face
LLM agents are no longer confined to FAQ chatbots. They handle product discovery, run suitability-style conversations, and field questions on pensions, insurance, and debt at a volume no human network could match. Meanwhile the unregulated foundation model is frequently the first place a consumer turns, before any authorised firm is in the loop at all.
Regulators have noticed. The FCA has published the Mills Review into the long-term impact of AI on retail financial services, the first review of its kind commissioned by a financial regulator, which frames AI as moving from a back-office efficiency tool into a more autonomous layer of regulated service delivery. In Europe, EIOPA has issued an opinion on AI governance and risk management for the insurance sector and flagged chatbots used in insurance distribution as a supervisory concern. The direction of travel is clear: the highest-volume advice channel now sits at, or beyond, the edge of the regulatory perimeter, and the perimeter itself is harder to draw than it has ever been.
The risks, concretely
Unreliable advice at scale. A human advice network fails one client at a time, and its failures are uncorrelated: a poor adviser in Leeds has no bearing on a good one in Bristol. A single model is a single point of failure that no human network ever had. One flawed system gives the same bad advice to millions of people simultaneously, and it does so without any of the local judgement that catches an outlier in a human process. The error is not distributed. It is broadcast.
Hallucinated product facts. When journalists at Sky News handed the same GBP 16,000 to ChatGPT, Copilot, and Gemini and asked each to build a portfolio, professional advisers reviewing the results found them heavily US-centric and poorly tailored to UK rules such as ISAs and gilts: plausible, confident, and wrong for the consumer in front of them. Broader testing points the same way. A study in the Journal of Financial Planning found that AI programs give inconsistent, inaccurate, and sometimes biased recommendations on emergency savings, asset allocation, and retirement withdrawals, in one case recommending different savings levels based on a user's inferred gender or race. That last finding is worth pausing on: the model was not asked about gender or race, and would deny using them if asked. The bias is real, unstated, and undetectable from the answer alone. It is the assume-not-ask failure and the reasoning-theatre failure compounding into a discrimination risk.
Advice-like outputs with no protections. When a general-purpose tool produces something that walks and talks like regulated advice, none of the regulated advice protections attach. The gap is not subtle. It is the entire consumer-protection stack:
| Protection | Regulated advice | AI answer that looks like advice |
|---|---|---|
| Suitability assessment | Required before a personal recommendation (COBS 9A.2) | None; the model infers what it needs |
| Suitability report | Firm must explain why the recommendation fits (COBS 9.4) | No durable record of the reasoning |
| Good-outcome duty | Consumer Duty obliges firms to act to deliver good outcomes | No obligation attaches to the output |
| Redress route | Financial Ombudsman Service and FSCS if advice causes harm | No eligible complaint, no compensation scheme |
| Accountable person | A named individual under SM&CR | Diffused across a model supply chain |
Accountability gaps and AI-washing. Modern agentic advice runs on a supply chain: a foundation model from one vendor, a retrieval layer from another, a fine-tune and an orchestration layer from the firm. The Senior Managers and Certification Regime pins accountability on a named individual, but that model strains when the decision-maker is a black box assembled from components no single person fully controls. When something goes wrong, each link in the chain has a defensible reason to point at another: the firm blames the base model, the model vendor points to the fine-tune, the integrator points to the prompt. Accountability that can be passed around a table indefinitely is not accountability.
Why the current toolkit falls short
Conduct regulators supervise advice through a small set of instruments: authorisation, suitability rules, disclosure, ongoing supervision, and enforcement. Every one of them was designed for human advice delivered by an authorised firm, on the assumption that the process being supervised is stationary - stable enough that a sample taken today still describes the system tomorrow. That assumption is the load-bearing beam, and a conversational agent snaps it.
A human advice process changes slowly: training, hiring, a policy update, all measured in months and all documented. An AI advice system can change its entire behavioural distribution overnight, with a prompt tweak, a fine-tune, or a silent vendor model upgrade that the firm did not initiate and may not even be told about. The thing you audited last quarter is not the thing serving customers this quarter. Against that non-stationarity, each classical instrument fails for a specific reason.
| Supervisory tool | How it works for human advice | Why it struggles with an AI agent |
|---|---|---|
| Thematic reviews and file sampling | Read a small sample of advice files against the suitability rules | A sample cannot represent an agent generating millions of unique conversations, and behaviour shifts the moment a prompt, fine-tune, or vendor model is updated |
| Mystery shopping | Run a scripted consumer journey through a few firms | Cannot probe the vast input space of a conversational model whose answer changes with phrasing, persona, and conversation history |
| Skilled-person reviews | Assess governance documents and model-risk frameworks | Reviews static documents while the live behaviour can change overnight with a new prompt or vendor upgrade |
The common thread is that these workflows are slow, backward-looking, and built to audit a stable process. A conversational agent is not a stable process. It is a system whose behaviour is a moving target across an input space too large to sample by hand and too volatile to freeze for inspection. You cannot mystery-shop your way across a space of every possible phrasing, and you cannot file-sample a distribution that resamples itself every deployment. The tools are not merely under-resourced. They are the wrong shape for the object.
Where oversight goes next: our thesis
The answer is not a smarter model. It is better mechanisms to test, audit, and explain how these systems behave. That is the problem Rayify works on, and it is why we think the next generation of supervisory tooling looks less like a rulebook and more like a test harness: something you point at a live system to characterise its behaviour empirically, continuously, and across an input space no human reviewer could walk by hand. A rulebook tells you what the system should do. A test harness tells you what it actually does, on this deployment, today.
Concretely, that means treating an advice system the way an engineer treats a service under test: probe it adversarially, assert on the properties that matter, and re-run the whole suite on every change. In pseudocode, an adversarial-persona red-team loop looks roughly like this:
for persona in persona_library: # includes vulnerable-customer profiles
for scenario in scenario_library: # pensions, ISAs, debt, insurance
for phrasing in paraphrase(scenario): # probe phrasing sensitivity
answer, stated_reason = advice_system.ask(persona, phrasing)
assert elicited_before_advising(answer) # did it ask, or assume?
assert compliance_steps_ran(answer) # suitability actually performed?
assert faithful(stated_reason, answer) # explanation reflects the decision?
assert refused_if_out_of_scope(persona, answer) # safe handoff to a human?
record(persona, phrasing, answer, failures) # build the distribution, not a sample
The point is not the specific asserts. It is the stance: you do not read a sample and hope it generalises, you generate the adversarial distribution on purpose and measure how often each property breaks. Five capabilities follow from that stance.
- Navigating the perimeter. Knowing where consumers actually receive AI advice - app stores, foundation models, embedded assistants, finfluencer channels - and estimating exposure inside and outside the regulatory boundary, so guidance and enforcement can target the highest-risk gaps rather than the most visible ones.
- Verifying authentic explainability. Distinguishing a faithful account of a model's decision from reasoning theatre. Given the faithfulness findings above, an explanation is only useful to a supervisor if it can be shown to reflect the actual computation, not a plausible story generated after the fact. This is the
faithful(stated_reason, answer)assert, and it is the hardest one to get right. - Testing for safe refusals. Checking that a model recognises the limits of its competence and declines, or hands off to a human, before it crosses from general information into regulated advice. The boundary is exactly where the assume-not-ask failure does its damage.
- Mapping supply-chain accountability. Tracing which vendor, layer, and configuration produced a given output, so that AI-washing has somewhere to stop and a named accountable person has something concrete to stand on rather than a diagram to point past.
- Identifying systemic concentration risk. Spotting when large parts of the market depend on the same underlying model, so that one model's flaw does not become everyone's flaw at once. Concentration turns an individual conduct problem into a financial-stability one.
We have been prototyping in each of these directions: an advice-market radar that maps where AI advice is actually consumed; a financial-advice red-teaming engine that probes advice systems with adversarial synthetic personas, including vulnerable-customer profiles, to surface suitability breaches and hallucinated product facts; a cross-market benchmark for agentic advice that scores any system, regulated robo-adviser or unregulated foundation model, against a shared scenario library; a consumer explainability translator that turns a model's decision into language a person can actually check; and a supervisory explainability toolkit in the spirit of the BIS Innovation Hub's Project Noor, which is prototyping XAI tools with the FCA and other supervisors to translate opaque model logic into human-readable explanations.
This is the same discipline we bring to Rayify's own product. We stress-test advice with synthetic personas before it reaches a person, and we translate a model's decision into an explanation that a regulator or a consumer can audit rather than merely admire. Trust in AI-enabled financial advice will not come from hoping the models are right. It will come from building the mechanisms that let us check.
Further reading
Adoption and market
- EY, 2026 - Nearly half of global consumers now use AI to guide savings and investment decisions
- OpenAI, 2026 - A new personal finance experience in ChatGPT
- Sky News test, write-up via Fast Data Science - AI for financial advice: the GBP 16,000 chatbot experiment
- CNBC, 2026 - Don't rely on AI for personal finance advice, study finds (Journal of Financial Planning)
Regulation and rules 5. FCA, 2026 - The Mills Review: AI and the future of retail financial services 6. FCA Handbook - COBS 9A.2: Assessing suitability 7. FCA Handbook - COBS 9.4: Suitability reports 8. FCA - Consumer Duty 9. FCA - Senior Managers and Certification Regime 10. EIOPA, 2025 - Opinion on AI governance and risk management
How models fail and how to explain them 11. Turpin et al., 2023 - Language Models Don't Always Say What They Think 12. Sclar et al., 2024 - Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design 13. Lanham et al., 2023 - Measuring Faithfulness in Chain-of-Thought Reasoning 14. Anthropic, 2025 - Reasoning Models Don't Always Say What They Think 15. BIS Innovation Hub - Project Noor: explaining AI models for financial supervision
Every link above was checked to resolve before publication. Where a primary source could not be linked directly (the Sky News portfolio test and the Lending Standards Board's chatbot research), we link the closest authoritative write-up and the peer-reviewed study that reaches the same finding.
Benchmarks and file samples describe a system that holds still. AI advice does not hold still. The way we earn trust in it is to build the harness that keeps testing it every time it moves.