Updated

Desk

Jev's output stays inside the schema the caller declared. A value inside that schema can still be false.

A decision model returns a choice, a score, or a yes/no from a list the caller wrote. Jev is TypeSafe's hosted model and the one this desk has followed longest. What a call returns is on what Jev is. The failure list is on limits.

What does TypeSafe mean by can't hallucinate?

TypeSafe's September 15, 2026 launch post says Jev gives up string generation and can't hallucinate.

Opened October 9, 2026, the launch post puts that sentence next to structured outputs. The section titled Hallucination and Type-safety says a hallucinated tool call is a type error, and that schema matching is guaranteed. It says the 0% on that plot is not an empirical count. A separate line says a type error would be easy to falsify with one counter-example, and calls that impossible. The same post says a Choice stops at 255 options. The launch story has the workflow table. This page did not redraw the plot.

The System One docs, opened the same day, say the model returns typed decisions and probabilities. They say it does not write replies, produce code, or generate explanations. The caller names the answers through Choice, Score, and Noul. The earlier reading of that argument is what Jev returns.

Can the chosen option still be false?

The System One docs say calibration is measured across groups of predictions and does not guarantee that an individual answer is correct.

The jaggedness page for jev-1.13, last reviewed September 17, lists nine failure modes and a workaround for each. The two examples are different tickets. A Noul for "Is the customer asking for a refund?" returns 0.22, and a yes/no Choice on that question returns 0.01 yes and 0.99 no. On another ticket, a Noul for a refund request returns 0.72 and a Noul for "something other than a refund" returns 0.47. Those two sum to 1.19. The docs say not to carry a threshold from a Noul over to a Choice, and not to ask Jev to generate text. We did not reproduce the examples.

Which published tables name hallucination?

The tables that name hallucination are separate tests, and this desk did not rerun them.

Cloudflare's Clef card, an internal Decision Index 0.2.1 run rounded to one decimal, prints RAGTruth hallucination F1 as 79.4, 35.6, and 76.5 for Clef, Clef-flash, and Jev. The card does not print the 57.91 composite. The write-up is Clef. We did not call the endpoint.

Perplexity's September 2026 panel scores accuracy, not that F1. RAGTruth there is 1,500 rows: Jev 77.27%, Qwen3.8-27B 61.53%, pplx-decider 88.80%. TruthfulQA binary on the same panel is 500 rows: Jev 92.00%, Qwen3.8-27B 82.80%, pplx-decider 85.40%. The panel is on pplx-decider.

Vela 2.0's evaluation file has no Jev row. RAGTruth there is 0.774 for the 9B, 0.770 for the 4B, 0.652 for the 0.8B, and 0.706 for the 0.3B. The file re-scores LettuceDetect v2 mmBERT on the same 2,700 rows at 0.743. LettuceDetect v2 qwen-2b, a generative model that only does hallucination, is 0.817 on RAGTruth and 0.921 on a 10,698-example hallucination set. The 9B on that set is 0.885, against Vela 1.0 at 0.875. The file says the 10,698 numbers come from model cards, while the RAGTruth numbers were re-scored. The page is Vela 2.0. We did not run the weights.

RLCDAlignBench counts hallucination as 6 benchmarks and 1,164 items inside a 7,193-item suite. The models that produce the behavior are Qwen3.5-2B, Phi-4-mini, Gemma-2-2B, Llama-3.2-3B, and Olmo-3-7B. jev-1.13.0 is the monitor. A generic Noul on that paper has median AUROC 0.886 over 31 benchmarks. That figure is not an F1 on RAGTruth. The page is RLCDAlignBench.

What happens when the question is whether a passage supports a claim?

Jamie Watters's September 21 post says Jev cleared a quote that was not in the source.

The five-run correction is a September 23 page about 42 claims, each paired with the passage it cites. On the hard 18, Jev's mean is 73.3% (72.2 to 77.8) and GPT-5.4's is 70.0%. On all 42 the means are 83.8% and 87.1%. The easy 24 stays at 91.7% for Jev, missing the same two controls on every run. On five identical calls the label changed for 2 of 42 claims, both under confidence 0.47. The labels are his. The page is 42 claims. We did not rerun it.

ForkScape's review section is a different set. It says unbacked claims were 15 of 15. jev-latest scores 56.3% of 80 review items. At a confidence gate of 0.8, accuracy is 88.5% on the 26 items it still answers. The page is fsdecide.

What does a high probability leave open?

On the sets that print a gate, some high-confidence answers are still wrong.

Eikos's README puts the error, among decisions at confidence 0.90 or higher, at 5.9% for Jev and 2.4% for both Eikos sizes, on 7,140 items across 6 suites. That row is not a JevBench composite. The page is Eikos.

Kev's function-selection report, on the description view, says that at a top probability of at least 0.9 Jev covered 1,176 of 1,253 cases and was wrong 10 times (0.85%). Kev covered 933 and was wrong once. The report says this is not the overall BFCL leaderboard. The page is Kev.

Step-Jev's BrowseComp-Plus note is another test. The hosted arm scores 0% accuracy, with answers rated about 70% likely correct and 0.0 tool calls. The rollouts behind that line are not in the public tree, so this desk did not recompute it. On October 9 the repository's own no-step test state scored incorrect at 1.0 on jev-1.13.0. The page is Step-Jev.

What does the schema stop?

JevAdvBench's abstract says fields outside the schema never reach jev-1.13.0.

The same abstract says one unverified opinion appended to the state flips 12.1% of decisions, and that this rate is statistically tied with the strongest injected command, which it puts at 10.1%. We read the abstract, not its 19 tables. The page is JevAdvBench.

A Tetris note on this desk says the schema keeps the selected placement inside the legal list the program supplied. The pattern of ranking a list the host already filtered is on patterns. The launch post's wikiracing note says lists longer than 255 go through a score, then an explicit choice.

Evals and compare keep each of these rows with the caveat printed next to it. Limits keeps the jaggedness list and the later misses. We did not rerun the tables on this page.