Published
Updated
Five runs leave a 3.3-point lead on 18 claim checks
This topic was created 10 days ago, and the information it contains may have evolved or changed since then.
Jamie Watters reran a 42-claim support check five times on September 22. On the hard 18, Jev's mean is 73.3% and GPT-5.4's is 70.0%. Exact McNemar on the majority labels gives p = 1.000. The September 23 page says eighteen claims cannot show the models are equal. He calls the repo JevBench. The 534-task table on this desk is a different project.
Jamie Watters posted on September 22 that a scratchpad in front of the verdict did not make GPT-5.4, Claude Sonnet 5, or Gemini 3.1 Pro more reliable, and that the extra text cost 2.4 to 4.2 times as much per answer. A September 21 post in the same series says Jev cleared a quote that was not in the source. The five-run correction is a September 23 page. Neither of those posts links it. The repo README does. He calls the harness JevBench. On this desk that name is Florian’s 534-task table. This one is TheWayWithin/jev-bench: 42 claims, each paired with the passage it cites, and the question is whether the passage supports the sentence as written. The license is MIT.
The hard 18 are sentences from one of his published articles, labelled from audits of the papers they cite. The page says nine are wrong, six are right, and three cite a passage that does not cover them. The other 24 are controls he built from those audits: a restatement, the same sentence with one digit changed, and a claim the passage does not mention. One hard claim is worth 5.6 points. The decision rule was written down before the first run. Replacing the frontier model required recall on unsupported of at least 0.80 and precision of at least 0.60. The first single run landed at 78.9% recall, one item short of that bar, which the README calls a pre-filter.
On September 20, one run each, the hard 18 was Jev 77.8%, GPT-5.4 66.7%, Sonnet 5 66.7%, Gemini 3.1 Pro 55.6%. Letting the frontier models write a scratchpad first, once each, moved those hard scores to 61.1%, 44.4%, and 55.6%. Cost per claim went from $0.001289 to $0.005450, from $0.002228 to $0.006439, and from $0.004824 to $0.011545. Exact McNemar on those single runs is p = 1.000, 0.219, and 1.000. The README says “reasoning did not help” is supportable and “reasoning made them worse” is not. Those three arms were not repeated.
The five-run file is from September 22, 18:29 to 18:48 EDT: 38 runs, 1,596 calls, no failed calls. The scoring rules were committed as b4a7ce5 thirty seconds before the first call. On the hard 18, means and ranges: Jev jev-latest 73.3% (72.2 to 77.8), GPT-5.4 70.0% (66.7 to 72.2), Sonnet 5 66.7% (61.1 to 72.2), Gemini 3.1 Pro 63.3% (55.6 to 66.7). Four of Jev’s five runs are 72.2%. That equals the best run of GPT-5.4 and of Sonnet 5. The September 20 score, 77.8%, is the top of this range, and it appears once. On all 42 claims the means are 83.8%, 87.1%, 85.7%, and 84.3%, so GPT-5.4 leads once the easy controls are included. The easy 24 did not move. The frontier models scored 100% on all fifteen of their runs. Jev scored 91.7% on all ten, missing ctl-03 and ctl-17 every time.
Exact McNemar on majority-of-five labels, hard 18: Jev alone right on 3 claims and GPT-5.4 alone right on 2, p = 1.000. Against Sonnet 5, p = 0.500. Against Gemini, p = 0.250. The lead on means is 3.3 points, 0.6 of one claim. The September 23 page puts a rough 95% interval on that gap at about 19 points worse to 30 points better, and says the interval is too wide to call a winner. The page also counts 19 wrong claims in the whole set. By the same three-of-five rule, Jev caught 15 and GPT-5.4 caught 18. Sonnet 5 and Gemini caught 16 each. The page gives p = 0.25 for that catch-count gap.
Three changes to the question were run the same day. Pinning jev-1.13.0 matched jev-latest on majority labels for all 42 claims. Every jev-latest call that day reported jev-1.13.0 as the model that answered. Splitting unsupported into contradicted and altered dropped Jev 1.4 points on all 42, to 82.4%. Asking three Noul questions in one call, the shape the README says TypeSafe recommends, cost $0.000026 per claim against $0.000025 and scored 77.6% against 83.8%, p = 0.219. The spend for all 38 runs is $2.2510. Per claim on the schema-only arm: Jev $0.000025, GPT-5.4 $0.001289, Sonnet 5 $0.002259, Gemini 3.1 Pro $0.004940. Jev’s figure is input tokens at $0.042 per million.
On five identical calls, jev-latest changed its confidence on 28 of 42 claims and its label on 2. Both label flips stayed under confidence 0.47 on every run. The verdict file says no claim that reached 0.5 changed label, in 420 schema-only calls under both names. On sb-06 the page records one run with answer probability 0.53 and confidence 0.28. Given an “altered” option, every system called sb-06 altered on every split run. Watters had labelled it supported. The page says dropping that one claim puts Jev and GPT-5.4 both at 74.1% on the other 17.
The labels are his. The sentences are from his publishing pipeline. The five runs sat in one 19-minute window. The reasoning arms are still one run each. python3 tests34.py analyse rereads the JSONL files with no API key. We did not run it, and we did not call the models.
This site's reading
Editorial notes evaluating claims against primary sources, contextualizing findings alongside related implementations, and defining technical terms.
Verify
The September 22 post, @Jamie_within, 21:34 UTC, links the reasoning note and says a scratchpad did not make three models more reliable, at 2.4 to 4.2 times the cost. The September 21 post links the confidence-threshold note. Neither post links the September 23 page.
That page, and the repo README, carry the five-run table. Runs: September 22, 2026, 18:29 to 18:48 EDT, 38 runs, 1,596 calls, no failures. Pre-registration commit b4a7ce5, 30 seconds before the first call.
Hard 18 means, jev-latest: Jev 73.3% (72.2 to 77.8), GPT-5.4 70.0% (66.7 to 72.2), Sonnet 5 66.7% (61.1 to 72.2), Gemini 3.1 Pro 63.3% (55.6 to 66.7). Easy 24: frontier models 100% on every run, Jev 91.7% on every run. All 42: 83.8%, 87.1%, 85.7%, 84.3%.
Majority-of-five McNemar, hard 18, Jev against GPT-5.4: 3 and 2, p = 1.000. Sonnet p = 0.500. Gemini p = 0.250. Spend $2.2510. Per claim that day: Jev $0.000025, GPT-5.4 $0.001289, Sonnet 5 $0.002259, Gemini $0.004940.
Label flips on jev-latest: 2 of 42, each under confidence 0.47. Confidence moved on 28 of 42. Fan-out: 77.6% against 83.8% on all 42, p = 0.219, $0.000026 against $0.000025. MIT license.
We did not rerun tests34.py against the JSONL, and we did not call the models.
Compare
Florian's JevBench is 534 frozen decisions and a four-axis composite. This harness is 42 claims about whether a cited passage supports a sentence, labelled by the person who wrote the sentences. Open-Jev's public slice is 231 of those 534 tasks.
The McNemar results here do not transfer to either of those tables. The September 23 page says p = 1.000 leaves the models unseparated, and that eighteen claims cannot show they are equal.
Terms
- hard 18
- Real sentences from one published article, paired with the passage they cite. Nine are wrong, six are right, and three cite a passage that does not cover them. One claim is 5.6 points. The other 24 claims are controls built from the same audits. Headline comparisons use the 18.
- 0.47
- On five identical jev-latest calls, the Choice label changed for 2 of 42 claims. Both flips stayed under confidence 0.47 on every run. Confidence changed on 28 of 42. The verdict file says no claim that reached 0.5 on any schema-only run changed label, across 420 calls under both model names.