Why TypeSafe Jev is not suitable as a hallucination judge
7 min read
Summary
TypeSafe Jev 1.13 is an external decision model: it answers typed yes/no, choice and score questions with probabilities, very fast and very cheaply, but it cannot write text.
We tested it as a hallucination judge in four flows on the public part of our benchmarking suite. The best flow catches about 32% of planted hallucinations in the direct answer at 2% false alarms. GPT-6 Sol (version of 22 September 2026) catches 88% at 2.7% on the same cases.
Two things hold it back. Because of its 32k-token limit, each statement can only be checked against a few selected source passages, and that selection misses the correct value in about a third of cases. And even when the correct value is in its input, Jev catches only 38% (18 of 48), often judging the wrong statement as supported: it recognises that a statement and a passage are about the same topic, but not that a name, date or period differs.
It also cannot give the evidence a relationship manager (RM) needs: no quotes, no reason, and therefore no verification step, which is what keeps our false alarms low.
Operationally it is limited to 32k tokens and text only, and it runs outside Switzerland, so it cannot be used with client identifying data (CID).
What Jev is
Jev is the flagship model of TypeSafe AI and the first of what TypeSafe calls "System One" models: models built for fast, structured decisions inside software rather than for generating text. You send it a piece of content (the state) and a set of typed questions, and it returns one calibrated probability per question:
Noul: a yes/no question; returns the probability of yes.
Choice: pick one option from a set you define; returns a probability per option.
Score: rate the content on ordered levels you define; returns a probability-weighted score.
All questions are evaluated in parallel on the same state in one request. TypeSafe positions Jev for classification, routing, guardrails and checks such as citation checking; it is not a general reasoning or writing model.
| TypeSafe Jev 1.13 (models, API) | |
|---|---|
| Output | Probabilities only (Noul, Choice, Score). No text. |
| Input limit | 32k tokens for the state plus the longest question |
| Input type | Text only (no images, so PDF pages must be converted) |
| Language | English primary |
| Hosting | External API outside Switzerland; zero data retention only on enterprise plans |
| Price | USD 0.042 per million input tokens (output free) |
The flows we tested
All flows got the same information as our GPT judge: the question, the answer, the cited sources and the date the answer was written. The flag threshold was set on one half of the correct answers (about 3% false alarms) and tested on the other half, in both directions; results of both directions are pooled.
Whole answer: one request with the full answer and all sources, two yes/no questions (contradiction anywhere; contradiction in the direct answer). Sources had to be trimmed to fit 32k tokens for half of the answers.
Per source: one request per source (long sources split into overlapping pieces), no trimming. Answer score = highest score over its requests.
Per statement: the answer is split into statements in code (sentences; table rows with their header row). Every source is cut into overlapping 120-word passages, and for each statement a keyword ranking (BM25) picks the 6 best passages: 4 from the sources the statement cites, 2 from any source. One request per statement (about 1,000–1,500 tokens) asks whether the statement is contradicted, supported or not mentioned by the passages, and whether it is part of the direct answer. Answer score = highest statement score. About 35 requests per answer.
Per statement with source identity (best): as 3, but every passage is prefixed with its source's title, URL and site, so Jev can see who published it. Without this, a passage cut from the middle of a document does not say whose document it is.
| Public cases | Planted direct-answer errors caught (72) | False alarms | Ranking quality (AUC) |
|---|---|---|---|
| Jev, whole answer | 12–15% | about 3% | 0.78–0.80 |
| Jev, per source | 6–18% | 1–7% | 0.71–0.74 |
| Jev, per statement | 33% | 2.6% | 0.83 |
| Jev, per statement with source identity (best) | 32% (46/144 pooled) [25–40%] | 2.0% (3/151) | 0.84 [0.79–0.89] |
| GPT-6 Sol, our flow | 88% (63/72) | 2.7% (4/150) | not applicable |
AUC measures how well the scores separate answers with a planted error from correct answers (1.0 perfect, 0.5 chance). A score of 0.84 shows real signal, but the overlap is too large for a warning that is both sensitive and rarely wrong.
Why it fails
1. It confirms the topic, not the fact
In the statement-level flow we can look at the exact statement that carries the planted error. In 48 of 69 cases the correct value (the original name, date or figure) was in the passages sent with that statement. Jev gave the statement a contradiction probability above 0.95 (roughly the level needed to keep false alarms at 3%) in only 18 of these 48 (38%). In many misses it judged the wrong statement as supported:
| Planted error | Source says (in Jev's input) | Jev: contradicted | Jev: supported |
|---|---|---|---|
| ASML quote attributed to the CFO, Roger Dassen | the CEO, Christophe Fouquet, said it | 0.02 | 0.93 |
| SNB "conditional inflation forecast of March 2026" | forecast of 18 June 2026 | 0.11 | 0.85 |
| ICE CoCo Index spread "as of 31 August 2026" | as of 30 June 2026 | 0.18 | 0.67 |
| Law firm view attributed to Linklaters | Matheson wrote it | 0.20 | 0.16 |
| Market view attributed to Goldman Sachs | Gramercy's outlook (source URL gramercy.com) | 0.35 | 0.52 |
The quoted words and figures appear in the source, so the statement looks supported; what differs is who said it or which date it refers to. These are exactly the errors that matter to an RM.
By error type (planted direct-answer errors with a contradiction probability above 0.95):
| Caught well | Caught poorly |
|---|---|
| Flipped comparison 6/7 | Wrong attribution 0/7 |
| Range endpoint 2/2 | Wrong "as of" date 0/6 |
| Wrong rule or status 4/7 | Wrong relative date 0/5 |
| Status or timing 4/7 | Wrong date 1/5, wrong period 1/5, peak or trough 1/5 |
| Wrong count 2/5, wrong figure 3/8, wrong entity 1/3 |
2. The input limit forces a lossy evidence selection
Our GPT judge reads every source in full. Jev's 32k-token limit means each statement can only be checked against a handful of selected passages. The keyword selection missed the correct value in 21 of 69 planted cases (for example, the passage naming the Bank of England member who actually made a quoted statement), so Jev never saw the evidence. Better retrieval would help here, but even with the evidence present Jev catches only 38% (see 1).
3. It cannot show its evidence
Jev returns a number, not a quote or a reason. Our GPT judge must quote the answer sentence and the source sentence for every contradiction, and a second call verifies each claim. That evidence is what makes a warning useful to the RM and what keeps false alarms at about 3 in 100. With Jev the best we can do is point at the highest-scoring statement: that was the edited statement in 45 of 69 cases.
4. Confident false alarms on correct statements
The correct answers it flagged most strongly, with contradiction probabilities of 0.97 to 1.00, include plain correct statements such as "Over the past two years, clean energy equities have outperformed, not underperformed." Without a reason or quote there is no way to verify or filter such flags.
5. Operational limits
32k tokens: about 40% of research answers with their sources exceed it, so the answer must be split into many requests.
Text only: PDF pages that our judge reads as images must be converted to text first.
Hosting: the API runs outside Switzerland. It cannot be used where RMs bring client identifying data into the agent, which is why GPT-5.1 with Swiss inference is our second production judge.
Where Jev could still help
Jev is extremely cheap (about USD 0.003 per answer in the statement-level flow, against about USD 0.07 for GPT-6 Sol) and each request answers in under a second. With a ranking quality of 0.84 it could serve as a pre-filter that skips the GPT judge on answers that are clearly fine. We have not measured how many GPT calls that would save without losing catches; with the current overlap we expect the saving to be modest.
Method
Cases (public only): web-search answers from the earlier benchmark, planted web-search errors of the RM benchmark, and the 9 September QA chats from web search, public documentation, Microsoft Learn and Yahoo Finance. 326 answers: 134 planted errors (72 in the direct answer), 151 correct real answers, 26 real hallucinations. No internal documents or tool data were sent.
Model: jev-1.13.0 via the TypeSafe API, 30 September and 1 October 2026. Comparison: GPT-6 Sol, version of 22 September 2026, reasoning effort none.
Thresholds: set on one half of the correct answers at about 3% false alarms, tested on the other half, in both directions. AUC with bootstrap 95% intervals.
Evidence check: for each planted error we take the words of the original value that the edit replaced (e.g. "CEO Christophe Fouquet") and check whether all of them appear in the passages sent with the edited statement.