Hallucination Check: a second compliance layer for relationship managers (September 2026)
10 min read
Key facts
Problem: about 20 in 1,000 agent answers contain a statement in the direct answer that contradicts the agent's own sources. We expect this for every product doing web and document search with the current generation of models, not only ours: our agent runs on the latest harness and frontier model (Claude Opus 5 on the Claude Code harness). Products without a check after the turn pass all of these answers to the user without a warning.
Solution: we built an LLM-as-a-judge for the financial industry. After every turn it checks the answer against every source the agent used, and it is tuned to the errors that matter in finance: figures, rates, dates and "as of" labels, reporting periods, regulatory status, attribution to the right institution or data provider.
Result: the judge catches nearly 9 in 10 of these hallucinations (GPT-6 Sol: 87%) and warns the relationship manager (RM) before the answer is used.
Two models evaluated: GPT-6 Sol (version of 22 September 2026) and GPT-5.1. GPT-5.1 runs fully in Switzerland with Swiss inference. RMs need to bring client identifying data (CID) into the agent, and for regulatory reasons the inference must then run fully in Switzerland. GPT-5.1 fulfils this, so we report it alongside GPT-6 Sol. GPT-5.1 catches about 8 in 10 (81%); the difference to GPT-6 Sol is not statistically significant.
False alarms: about 3 in 100 correct answers receive a warning (both models).
Grounding: only 5% [4–7%] of direct statements in real agent answers are not backed by any of the agent's references. This justifies checking the answer against its references.
Time: the check adds a median of 2–5 seconds with GPT-6 Sol and 17–32 seconds with GPT-5.1.
Benchmarking suite of 566 labelled agent answers: 345 real agent answers from web search, internal documents and nine tool integrations, and 221 answers with planted financial errors across 13 error types. Every answer is stored with its full sources and checked by every judge configuration, reported with confidence intervals and significance tests.
Next step: the agent uses the judge's findings to correct itself and proposes a corrected answer to the RM.
The problem: references alone do not protect the relationship manager
Relationship managers (RMs) use the agent to prepare client conversations: market moves, valuations, rates, regulation, fund due diligence. Every answer carries references to the sources it used (web pages, internal documents, tools such as Jira, Outlook or Slack). From a compliance perspective references are an important control: every statement can be traced back to a source.
In practice an RM does not have the time to open every reference and check every figure before a client call. A reference only protects the RM if someone checks it. This is why we add a second compliance layer after every turn: the hallucination check. It reads the finished answer together with every source the agent received and warns when the answer contradicts its own sources.
Our approach: an LLM-as-a-judge for the financial industry
An LLM-as-a-judge is a second language model that does not answer the question itself but reviews the answer of the first. General-purpose judges ask "is this answer supported?". That produces many false alarms on long research answers and misses the small errors that matter in finance. Our judge is built for the financial industry and looks for exactly those errors:
Figures and units: prices, yields, rates, AUM, weights, multiples, basis points vs percent.
Dates and periods: "as of" dates, fiscal vs calendar periods, forecast years, and relative dates ("last Friday", "yesterday") computed from the day the answer was written.
Regulatory status and rules: proposed vs adopted, allowed vs prohibited, effective dates.
Entities and attribution: similar names (Swiss Life vs Swiss Re), which broker, data provider or central bank said it.
Market claims: beat vs miss consensus, record highs, peaks and troughs, rankings.
It deliberately ignores what is not a compliance risk: rounding, wording, and statements the sources simply do not mention. Where two sources disagree, it reports a source conflict instead of blaming the answer.
How it works:
After the turn, the judge receives the answer, every source the agent used (web pages, documents, PDF pages as images, tool results) and the date the answer was written.
For every contradiction it claims, it must quote the sentence from the answer and the sentence from the source.
Each claim is verified: the quote must exist verbatim in the answer, and a second call rules on the claim.
Only a verified contradiction triggers a warning to the RM, showing the statement and the source it contradicts.
Quoting and verifying are what keep false alarms at about 3 in 100 correct answers: the judge cannot warn without evidence.
Scope and grounding
The check compares the answer with the references the agent received. This only protects the RM if the answer is actually grounded in those references, i.e. what the agent writes comes from its sources. We measured this.
Result: only 5% [4–7%] of direct statements are not backed by any reference. The remaining direct statements are backed by the references.
How we measured it:
We took the 310 real agent answers without a hallucination from the benchmarking suite (all tools: internal search, web search, Atlassian, Slack, public documentation, Microsoft Learn, Yahoo Finance) and split them into 7,869 statements (sentences and table rows).
GPT-6 Sol classified each statement as a direct statement or not. A direct statement gives the RM factual information they would rely on: a figure, date, name, price, rule, event or status. Headings, intros, offers, advice and remarks about the search itself are not direct statements. 6,682 statements (85%) are direct.
For each direct statement, GPT-6 Sol checked all references the agent received, regardless of which reference the statement points to, and decided whether a reference backs it. For every backed statement it had to quote the reference; the quote is verified verbatim.
95% interval by bootstrap over answers.
Whether a statement points to the right reference (citation correctness) remains out of scope; we focus on hallucinations.
Definition: what counts as a hallucination
An answer hallucinates when a statement contradicts a source the agent was given: the source states the same thing (same entity, metric, period, definition) differently. We focus on hallucinations in the direct answer, i.e. the part that answers what the RM asked. This is what the RM takes into the client conversation.
Counts as hallucination | Does NOT count as hallucination |
|---|---|
Wrong figure (price, rate, %, AUM, weight) | Statements the sources do not mention (unsupported) |
Wrong date, wrong period or "as of" label | Rounding, or a more precise figure that rounds to the source's value |
Wrong relative date or weekday ("last Friday", "yesterday") measured from the day the answer was written | Differences in wording, emphasis or tone |
Wrong rule, status or yes/no (allowed vs not allowed, proposed vs adopted) | Following one of two sources that disagree with each other (reported separately as a source conflict) |
Value on the wrong entity, wrong attribution (who said it) | The answer's own arithmetic, unless a source states the computed value |
Flipped comparison, wrong count, wrong record/high/low claim |
What to expect
Per 1,000 RM questions (95% intervals in brackets):
Per 1,000 answers | Products without a check after the turn | With check, GPT-6 Sol | With check, GPT-5.1 (Swiss inference) |
|---|---|---|---|
Answers with a hallucination in the direct answer | 20 [5–40] | 20 | 20 |
of which the RM is warned | 0 | about 17 (87% [79–93%]) | about 16 (81% [72–88%]) |
of which reach the RM without a warning | 20 | about 3 | about 4 |
Correct answers with a warning (false alarms) | 0 | about 3% of correct answers (9/310) [1.5–5.4%] | about 3% of correct answers (9/310) [1.5–5.4%] |
The rate of 20 per 1,000 comes from 200 human-labelled real agent chats (QA check of 9 September 2026). The catch rates come from the RM benchmark below; the false-alarm rate from 310 correct real answers.
Detection on the RM benchmark
We took real, correct agent answers to RM questions (market questions answered with web search, fund due-diligence questions answered from internal documents) and planted one small, plausible error in the direct answer of each: 94 answers.
GPT-6 Sol | GPT-5.1 (Swiss inference, reasoning effort medium) | |
|---|---|---|
Hallucinations in the direct answer caught | 87% (82/94) [79–93%] | 81% (76/94) [72–88%] |
Market questions (web search) | 88% (63/72) | 76% (55/72) |
Fund due diligence (internal documents) | 86% (19/22) | 95% (21/22) |
Correct answers with a warning | 2.9% (9/310) [1.5–5.4%] | 2.9% (9/310) [1.5–5.4%] |
The difference between GPT-6 Sol Sol and GPT-5.1 is t statistical significant (paired McNemar test, p = 0.26). GPT-5.1 must run at reasoning effort medium; at minim al effort it catches only 54% (51/94) of planted direct-answer hallucinations and flags 4.2% of correct answers. All evaluated configurations are listed on the sub-page Hallucination Check: judge model and reasoning effort comparison.
Side note on real cases: in 345 labelled real agent answers we found only 10 hallucinations in the direct answer (GPT-6 Sol caught 6, GPT-5.1 caught 5). Ten cases are too few for a reliable catch rate, and labelling enough real data to find more was out of scope for this work. We therefore modelled the benchmark's planted errors on the error types seen in these real cases and included them in the benchmark.
Time and cost
Median | 90th percentile | Cost per answer | |
|---|---|---|---|
Hallucination check, GPT-6 Sol | 2–5 s | 5–11 s | USD 0.04–0.07 |
Hallucination check, GPT-5.1 (effort medium) | 17–32 s | about 55 s | USD 0.03–0.06 |
The check runs after the turn, so the RM sees the answer first and the warning follows.
Examples: what is detected
Real answers where the RM was warned
Wrong date (web search). Question: when does the Glastonbury Festival take place?
Answer: "Wednesday 25 June to Sunday 27 June 2027"
Source (official site): "from Wednesday 23rd June to Sunday 27th June, 2027"
Caught by both models.
Wrong weekday (web search). Question: Apple's most recent closing price.
Answer (written Wednesday 30 September 2026): "at the 4:00 PM EDT close on Monday, 29 September 2026"
Source: market time 2026-09-29, which was a Tuesday. The price was correct; the day was not.
Error types in the RM benchmark (direct answer)
Error type | Example (source → edited answer) | GPT-6 Sol | GPT-5.1 |
|---|---|---|---|
Wrong figure | S&P 500 ten-year average forward P/E 19.0x → 18.4x | 11/12 | 11/12 |
Wrong rule or status | "national regulators may still impose inducement bans" → "may no longer impose" | 11/12 | 11/12 |
Wrong date | Swiss Federal Council decision 17 November 2025 → 19 November | 8/8 | 8/8 |
Wrong "as of" date | figure for May 2026 → labelled August 2026 | 7/8 | 6/8 |
Value on the wrong entity | Swiss Life 4.1% and Swiss Re 4.9% dividend yields swapped | 7/7 | 5/7 |
Flipped comparison | US core CPI 0.3% beat the 0.2% forecast → "came in below consensus" | 7/7 | 5/7 |
Wrong count | "two savings-credit rates instead of four" → "instead of five" | 6/7 | 6/7 |
Wrong attribution | consensus estimate from Visible Alpha → "from LSEG" | 6/7 | 5/7 |
Status or timing | Commission proposed ESMA as supervisor → presented as decided | 5/7 | 6/7 |
Wrong relative date | data "through 17 September" → "through last Friday" (= 18 September) | 3/5 | 0/5 |
Wrong period label | EPS growth forecast for 2026 → labelled 2027 | 4/5 | 5/5 |
Record or peak claim | "highest policy rate since April 1995" → "record high" | 4/5 | 4/5 |
Range endpoint | target 200–300 portfolio companies → 200–250 | 3/4 | 4/4 |
What is deliberately NOT flagged
Unsupported statements. An answer listing six investment committee members where the documents name three: the named three are correct and nothing contradicts the others.
Source conflicts. One source dated Apple's $338.98 close 18 September, four others 21 September. The answer followed the majority; this is reported as a source conflict, not a hallucination.
Rounding and wording. "4.24 trillion" where the source says "4.2 trillion".
Next step: from warning to correction
Today the check warns the RM. Next, we will work on a concept in which the agent improves automatically based on the judge's findings: the judge tells the agent which statement contradicts which source, and as a first step the agent proposes a corrected answer to the RM, who accepts it before using the answer.
The benchmarking suite
To build and tune the judge we created a benchmarking suite of 566 labelled agent answers: 345 real agent answers and 221 answers with planted errors. Each answer is stored with the full question, the answer, the date it was written and every source the agent received, so every judge configuration is tested on exactly the same evidence.
Part | Answers | What it tests |
|---|---|---|
Real agent chats, QA check of 9 September 2026: internal search, web search, Slack, Jira, Confluence, Outlook, public documentation, Microsoft Learn, Yahoo Finance | 200 | Production hallucination rate and false alarms across all tools |
Real web-search and internal-document answers with statement-level human labels | 145 | False alarms on long research answers, real hallucinations |
RM benchmark: real correct answers with one planted financial error (13 error types, market and fund due-diligence questions) | 199 (94 in the direct answer) | Catch rate on subtle, realistic errors |
Synthetic figure swaps | 22 | Catch rate on obvious errors |
Total | 566 |
The suite is used to compare judge models (GPT-5.1, GPT-6 Sol), reasoning efforts (none, minimal, medium) and prompt versions, always on the same cases:
every planted error is checked against the sources, and the cases a model misses plus a random sample of caught cases are independently audited;
310 real correct answers measure false alarms and grounding;
results are reported with 95% Wilson intervals and bootstrap intervals, and models are compared with paired exact McNemar tests;
latency and cost are logged per call.
Method notes
Models: GPT-6 Sol, version of 22 September 2026, reasoning effort none; GPT-5.1, version of 13 November 2025, reasoning effort medium, Swiss inference.
Production rate: 4 of the 200 real chats contain a hallucination in the direct answer, relabelled against the definition above with every contradiction quoted. With so few cases the interval is wide (5–40 per 1,000).
RM benchmark: the 13 error types are modelled on the real cases.
Grounding: 343 of 6,682 direct statements are not backed by any reference (5.1%, 95% bootstrap interval over the 310 answers 4–7%). Measured with GPT-6 Sol (reasoning effort none), one call per answer, USD 14.48 for all 310 answers.
Limitation: catch rates are measured on planted errors; see the side note on real cases above.