Hallucination Check: judge model and reasoning effort comparison

5 min read

This page lists every judge configuration we evaluated on the benchmarking suite. It complements the parent page, which reports only the two production configurations: GPT-6 Sol (reasoning effort none) and GPT-5.1 (reasoning effort medium, Swiss inference). All GPT-6 runs used GPT-6 Sol, version from 22nd of September 2026.

All GPT configurations use the same judge flow (contradiction prompt, verbatim quotes, verification call) and were run on the same 566 answers. TypeSafe Jev, an external decision model, is reported separately at the end because it needs a different flow and was run on public cases only.

Summary

  • GPT-6 Sol, effort none (production): best balance. It catches 87% of planted direct-answer hallucinations with 2.9% false alarms and is among the fastest.

  • GPT-6 Sol, effort medium: similar catch rate on planted errors (89%, not significantly different) but more than twice the false alarms (6.5% vs 2.9%, significant). Not chosen, because a warning that is often wrong erodes the relationship manager's trust in it.

  • GPT-5.1, effort medium (production, Swiss inference): 81% of planted direct-answer hallucinations, same false-alarm rate as GPT-6 Sol. Slower (median 29 s).

  • GPT-5.1, effort minimal: about five times faster than medium, but catches only 54% of planted direct-answer hallucinations, significantly fewer than both production configurations. Not recommended.

  • GPT-5.4-mini, effort none: the fastest and cheapest GPT option (median 2 s, USD 0.02) with the fewest false alarms (1.6%), but catches only 28% of planted direct-answer hallucinations. Not suitable as the judge.

  • GPT-4o on the same flow: catches only 13% of planted direct-answer hallucinations, with 3.9% false alarms. Six of the 566 answers exceed its context window. Not suitable as the judge.

  • TypeSafe Jev 1.13 (external, best of three flows): catches about 33% of planted direct-answer hallucinations at 2.6% false alarms, against 88% for GPT-6 Sol on the same public cases. Extremely cheap (about USD 0.003 per answer) but not suitable as the judge. See the sub-page Why TypeSafe Jev is not suitable as a hallucination judge.

Results

95% Wilson intervals in brackets. Planted cases: RM benchmark answers with one planted error. False alarms: 310 correct real answers.

GPT-6 Sol, none (production)

GPT-6 Sol, medium

GPT-5.1, medium (production, Swiss)

GPT-5.1, minimal

GPT-5.4-mini, none

GPT-4o

Planted, hallucination in the direct answer (94)

87% (82/94) [79–93%]

89% (84/94) [82–94%]

81% (76/94) [72–88%]

54% (51/94) [44–64%]

28% (26/94) [20–37%]

13% (12/94) [8–21%]

Planted, all (199)

90% (179/199) [85–93%]

90% (179/199) [85–93%]

84% (168/199) [79–89%]

56% (111/199) [49–63%]

26% (51/199) [20–32%]

8% (15/199) [5–12%]

False alarms (310 correct answers)

2.9% (9/310) [1.5–5.4%]

6.5% (20/310) [4.2–9.8%]

2.9% (9/310) [1.5–5.4%]

4.2% (13/310) [2.5–7.1%]

1.6% (5/310) [0.7–3.7%]

3.9% (12/310) [2.2–6.6%]

Answers flagged red in production (estimate)

4.9% [2.8–7.4%]

9.6% [6.5–12.8%]

4.9% [2.8–7.4%]

6.1% [3.7–8.9%]

2.1% [0.7–3.7%]

4.4% [2.3–6.6%]

Share of red answers that are real hallucinations (estimate)

45%

37%

45%

36%

27%

17%

Check time, median / 90th percentile

4 s / 9 s

6 s / 13 s

29 s / 59 s

6 s / 25 s

2 s / 6 s

10 s / 81 s

Cost per answer (mean)

USD 0.07

USD 0.08

USD 0.06

USD 0.04

USD 0.02

USD 0.06

For GPT-4o, the six answers without a verdict (context window exceeded) count as "no warning", as they would in production.

Per 1,000 RM questions

Expected: 20 answers per 1,000 contain a hallucination in the direct answer (200 real agent chats, QA check of 9 September 2026). Applying each configuration's catch rate on planted direct-answer errors:

Per 1,000 answers

GPT-6 Sol, none

GPT-6 Sol, medium

GPT-5.1, medium

GPT-5.1, minimal

GPT-5.4-mini, none

GPT-4o

Jev 1.13 (statement-level)

RM warned

about 17

about 18

about 16

about 11

about 6

about 3

about 7

Reach the RM without a warning

about 3

about 2

about 4

about 9

about 14

about 17

about 13

Correct answers with a warning

about 28

about 64

about 28

about 41

about 16

about 38

about 25

Statistical tests

Paired exact McNemar tests against GPT-6 Sol (effort none), on the same cases:

Compared with GPT-6 Sol, none

Planted direct-answer hallucinations caught

False alarms

GPT-6 Sol, medium

no significant difference (catches 5 that none misses, misses 3; p = 0.73)

significantly more (17 vs 6 discordant; p = 0.03)

GPT-5.1, medium

no significant difference (7 vs 13; p = 0.26)

no difference (6 vs 6; p = 1.0)

GPT-5.1, minimal

significantly fewer (3 vs 34; p < 0.001)

no significant difference (11 vs 7; p = 0.48)

GPT-5.4-mini, none

significantly fewer (0 vs 56; p < 0.001)

no significant difference (5 vs 9; p = 0.42)

GPT-4o

significantly fewer (1 vs 71; p < 0.001)

no significant difference (12 vs 9; p = 0.66)

GPT-5.1 minimal vs GPT-5.1 medium: medium catches 29 planted direct-answer hallucinations that minimal misses, minimal catches 4 that medium misses (p < 0.001). False alarms 13 vs 9 (p = 0.45). GPT-5.4-mini vs GPT-5.1 medium: medium catches 50 that mini misses, mini none that medium misses (p < 0.001).

External model: TypeSafe Jev 1.13

Jev is a decision model: it returns probabilities for typed yes/no, choice and score questions and cannot write text, so it cannot quote the contradicting sentences our flow relies on. We tested three flows; the best one (statement-level) is reported here. For data-protection reasons only public cases were sent (web search, public documentation, Microsoft Learn, Yahoo Finance): 72 planted direct-answer errors, 134 planted errors in total, and 151 correct real answers. Jev's flag threshold was set on one half of the correct answers (about 3% false alarms) and tested on the other half, in both directions; the two results are pooled below.

Public cases only

Jev 1.13, statement-level (best)

GPT-6 Sol, none (same cases)

Planted, hallucination in the direct answer (72)

33% (47/144 over both halves; 28% and 38%)

88% (63/72) [78–93%]

Planted, all (134)

29% (24% and 35%)

89% (119/134) [82–93%]

False alarms (correct answers, held-out half)

2.6% (4/151)

2.7% (4/150)

Ranking quality (AUC), planted direct errors vs correct answers

0.83 [0.77–0.88]

not applicable

Cost per answer

about USD 0.003

about USD 0.07

The other two Jev flows were weaker: whole answer in one request (sources trimmed to fit 32k tokens) 12–15% at about 3% false alarms, AUC 0.78–0.80; one request per source (no trimming) 6–18%, AUC 0.71–0.74.

Configurations

Configuration

Model deployment

Reasoning effort

GPT-6 Sol, none

GPT-6 Sol (2026-09-22)

none

GPT-6 Sol, medium

GPT-6 Sol (2026-09-22)

medium

GPT-5.1, medium

GPT-5.1 (2025-11-13), Swiss inference

medium

GPT-5.1, minimal

GPT-5.1 (2025-11-13), Swiss inference

minimal

GPT-5.4-mini, none

GPT-5.4-mini (2026-03-17)

none

GPT-4o

GPT-4o (2024-11-20)

not applicable (no reasoning)

TypeSafe Jev

jev-1.13.0 (external API, outside Switzerland)

not applicable (decision model)

Times and costs are measured per judged answer across all 566 answers (Jev: public cases), including the verification call when a GPT judge claims a contradiction.

Last updated