Client communication
Incorrect product terms, disclosures, fees, suitability information, market facts, or account explanations can become representations by the institution.
Generative AI is entering financial workflows that depend on factual accuracy, traceable authority, and accountable judgment. These systems can produce confident falsehoods. Deploying them does not displace the institution's obligations governing the underlying activity.
NIST AI 600-1 uses the term confabulation for confidently presented erroneous or false content. In financial services, hallucination becomes institutional risk when unsupported output enters a consequential workflow.
A generative AI response can contradict its source material, invent a fact or citation, misstate a rule, perform a calculation incorrectly, or supply a plausible answer where the available evidence supports no answer. Fluency is not verification.
Under the Synthetic Outlaw framework, the institution can treat the output as routine software assistance even when it affects regulated communication, customer treatment, compliance interpretation, financial analysis, or operational judgment. The AI does not need intent or malice. It can produce a prohibited or harmful outcome while the surrounding process appears formally compliant.
The relevant question is not whether a chatbot occasionally produces a wrong answer. It is whether unsupported output can cross into a workflow that creates duties, decisions, records, or customer consequences.
Incorrect product terms, disclosures, fees, suitability information, market facts, or account explanations can become representations by the institution.
A false citation or misreading of a rule, policy, or supervisory requirement can redirect internal controls while appearing authoritative.
Invented facts, calculations, sources, or market claims can contaminate analysis before a reviewer recognizes that the supporting record does not exist.
Unsupported summaries or classifications can distort escalation, investigation, and case prioritization in high-volume review environments.
A vendor can supply the model, guardrail, or evaluation layer. Reliance on that provider does not eliminate the institution's obligations arising from deployment and use.
Selected U.S. authorities recognize the risk but address it through different institutional regimes. Hallucination monitoring is not a solved extension of traditional model validation.
FINRA reports that member firms have begun implementing generative AI, particularly for internal processes and information retrieval. It identifies hallucination, supervision, testing, monitoring, logging, human review, data sensitivity, and third-party use as considerations. The report describes observed practices and does not create a new cross-sector rule. Primary source →
The Federal Reserve, OCC, and FDIC apply the revised model-risk guidance to traditional statistical and quantitative models and non-generative, non-agentic AI. Generative and agentic AI are outside its scope; banking organizations are directed to use broader risk-management and governance practices to determine appropriate controls. This is a scope distinction, not an absence of obligations. Primary source →
Commercial and internal tools can test, observe, and flag generated output. Procurement should begin with the required control function, not a vendor claim. No monitoring product independently establishes accuracy, fitness for purpose, or compliance.
| CAPABILITY CATEGORY | WHAT IT SHOULD DO | WHAT THE INSTITUTION MUST REQUIRE | EVIDENCE BOUNDARY |
|---|---|---|---|
| Grounding and faithfulness evaluation | Compare generated claims with approved source material and identify unsupported, contradictory, or fabricated content before release and during production. | Testing against institution-specific documents, known failure cases, numerical and citation errors, conflicting sources, stale information, and appropriate abstention. | An evaluator is another model or ruleset with its own error rate. Its findings require validation, review, and an escalation path. |
| Production observability and alerting | Preserve traces, measure defined failure signals, detect changes over time, and alert accountable owners when thresholds are crossed. | Use-case-specific thresholds, retained evidence, reproducible incident review, access controls, version history, and a demonstrated stop mechanism. | A dashboard records behavior. It does not determine whether the resulting institutional action was lawful, fair, accurate, or appropriately authorized. |
| Domain-specific test harnesses | Evaluate outputs against financial terminology, calculations, policies, products, regulations, and realistic consequential workflows. | Representative internal cases, controlled reference answers, independent challenge, repeatable regression testing, and documented acceptance criteria. | A generic benchmark cannot establish fitness for a particular institution, jurisdiction, customer population, or deployment context. |
Procurement rule: require vendors and internal teams to demonstrate performance against the institution's own data, failure cases, thresholds, and escalation requirements. Purchasing a monitoring product does not transfer the institution's applicable legal, supervisory, contractual, or governance responsibilities.
Everyone can follow the approved process and the client can still receive an answer the institution cannot defend.
A relationship manager asks the firm's AI assistant whether a client qualifies for a product. The AI relies on an old policy, misses a key exception, and its answer is sent after a routine human review.
The employee asks the firm's approved AI whether the client qualifies.
The AI gives a confident answer based on an old policy and misses the current exception.
A reviewer checks whether the answer sounds plausible but cannot see which policy the AI used.
The answer is sent to the client. The firm cannot later show what supported it.
The Observatory tracks documented AI governance events in financial services. Open a record below, see the full sector archive, or enter the same evidence in Deep Field.
A credible control environment begins with evidence that the institution knows where generative output enters consequential work and can stop unsupported output before it becomes institutional action.
An enterprise inventory that lists tools but not decision pathways does not answer this question.
Logging a prompt and response is not equivalent to preserving the authority, context, model version, retrieval result, and human decision that supported the action.
A global accuracy target ignores material differences between drafting, research, client communication, and consequential decision support.
If the same model family or evaluation method produces and judges the answer, correlated failure must be treated as a control risk.
Distributed participation cannot become distributed non-accountability. Named ownership and escalation authority must exist before deployment.
Effective controls must reflect the institution, use case, jurisdiction, data, model, vendors, and consequences of failure.
Identify every point where generated content can influence a regulated communication, customer outcome, risk judgment, financial analysis, investigation, or institutional record.
Preserve the prompt, response, retrieved sources, applicable policy, model and evaluator versions, control result, human intervention, and final action where the use case is consequential.
Use approved data, realistic adversarial cases, calculation tests, stale-information tests, source-conflict tests, and known failure examples. A generic benchmark cannot establish fitness for a specific workflow.
Track unsupported claims, source adherence, abstention, overrides, user corrections, drift, exceptions, and downstream incidents. Define thresholds by use case and consequence.
Specify who can suspend a use case, what evidence triggers review, how customers and records are corrected, and how failures are reported across compliance, risk, legal, audit, and the board.
Document the senior owner accountable for permissible use, control performance, exceptions, and escalation alongside the distinct responsibilities of business, legal, compliance, risk, audit, technology, and board functions.
A focused institutional review maps where one deployed or proposed AI system can produce a prohibited or harmful outcome while appearing compliant or operating beyond effective institutional control. The work begins with the real decision pathway, evidence, authority, and consequence structure.
This brief relies on selected primary authorities. It does not provide a comprehensive statement of the law governing every financial institution or AI use case.
Synthetic Outlaw Research. “AI Hallucination Risk in Financial Services.” Executive Risk Brief 01, version 1.0. July 21, 2026. https://www.syntheticoutlaw.com/industries/financial-services.html.