Updated

Published

JevBench scores typed decisions on four axes, and Jev's lead is a pricing story

Florian S published JevBench, a 534-decision suite for models that return a label and a probability. The composite is a geometric mean of intelligence, calibration, speed, and cost. Jev 1.13.0 sits at 75.3. A 4B open rebuild is 0.7 behind. GPT-5.6 Luna wins the hard items and loses the score on price.

Florian S posted JevBench on September 19 as a benchmark for models that take state plus a fixed label set and return a typed answer with probabilities. The interactive table is on Benchmark Heaven. The harness, tasks, and scoring live in fstandhartinger/jevbench. The README says the project is not affiliated with TypeSafe.

v1.2.2, the revision on the page we fetched, runs 534 decisions per system. Easy 72, standard 96, judge 146, hard 220. The published score is a geometric mean of four axes at 25% each: Intelligence (weighted accuracy), Calibration (hard-tier ECE and fidelity to gold distributions), Speed, and Cost. One weak axis pulls the number down.

The ranking Florian quoted in the first posts was Jev 1.13.0 at 75.3 and SemIf, a Qwen3.5-4B open rebuild, at 74.6. The current table still has those two rows. It also puts classifier.dev’s fast tier first at 84.8. The README says that tier is Jev behind another API: Intelligence 90.1 versus Jev’s 90.4, p50 0.39 s versus 0.65 s from the bench’s server, and a Pro plan the authors price at full use as about $0.0033 per 1,000 decisions against Jev’s measured $0.041. That is a hosting and billing row, not a second model.

GPT-5.6 Luna (low reasoning effort) has the highest Intelligence, 96.8, and the highest hard-tier accuracy among finished runs, 94.5% against Jev’s 74.1%. Its composite is 66.0, rank 12, because Cost is 28.2 at $0.247 per 1,000 decisions. DeepSeek V4.1 Flash is similar: Intelligence 96.1, Calibration 96.7, Cost 17.1.

Hard items were written by Claude Opus 5 and GPT-5.6 Sol, then cross-reviewed and hashed before any system ran. 111 of 220 are public. Self-hosted and demo endpoints have their latency multiplied by 2 and 0.15 s added; the README calls that an assumption, not a measurement. Production APIs, including Jev, are left raw.

A Benchmark Heaven reply on the same thread says TypeSafe called Master Customer Agreement section 2.3(f) an outdated pre-launch constraint they are fixing. We did not find a TypeSafe post that says that.

Open rebuilds fill most of the middle. SemIf 74.6, djev 74.3, Laya 70.1, jeff 66.9, Bespoke Nimble 9B 63.5. Several rows are estimated cost, not a bill. We did not rerun the suite.

This site's reading

Editorial notes evaluating claims against primary sources, contextualizing findings alongside related implementations, and defining technical terms.

Verify

The September 19 posts point at benchmarkheaven.com/jev-models and github.com/fstandhartinger/jevbench. The README calls JevBench a Benchmark Heaven project, not affiliated with TypeSafe. Current published revision is v1.2.2: 534 decisions per system (easy 72, standard 96, judge 146, hard 220). The JevBench Score is a geometric mean of Intelligence, Calibration, Speed, and Cost at 25% each. Intelligence is weighted accuracy (hard 30%, easy 14%, standard 28%, judge 28%). Hard items were written by Claude Opus 5 and GPT-5.6 Sol, cross-reviewed, frozen and hashed before any system ran; 111 of 220 are public. Headline ranking on the page we fetched: classifier.dev fast tier 84.8, Jev 1.13.0 75.3, SemIf (Qwen3.5-4B) 74.6, djev 74.3, Laya 70.1, open-alternative-jev 69.8, jeff 66.9, GPT-5.6 Luna (low) 66.0 at rank 12. The README states classifier.dev's fast tier is Jev behind that host's API: Intelligence matches (90.1 versus 90.4); it leads on measured p50 (0.39 s versus 0.65 s) and on a flat Pro plan priced at full use (~$0.0033 per 1,000 decisions versus Jev's $0.041). Luna Intelligence 96.8, hard-tier accuracy 94.5% versus Jev 74.1%, cost $0.247 per 1,000. Self-hosted and demo latency is adjusted ×2 plus 0.15 s; production APIs are not. We did not rerun the suite. TypeSafe's Master Customer Agreement section 2.3(f) forbids publishing benchmarks of the Services; this story reports the posts and repo as published. A Benchmark Heaven reply (2101385137128681713) says TypeSafe called that clause an outdated pre-launch constraint they are fixing. We did not find a TypeSafe post that says that.

Compare

Most measurements on this desk are one author's set against one or two models. JevBench is the first shared rubric we have that also prices latency and dollars, so Luna can win the hard items and still sit below a 4B open rebuild. Kumar and Poly report accuracy or F1 on public labels. jeff's own 1,600-item table is a different set; on this suite jeff is 66.9, run on the author's CPU with the latency adjustment. Nimble is 63.5 here versus 90.12% on its own 324 labels. LocalJev is not in the ranking. The geometric mean is a scoring choice, not a proof that Jev is generally stronger than Luna.

Terms

JevBench Score
Benchmark Heaven's composite for typed-decision systems: Intelligence, Calibration, Speed, and Cost at 25% each, combined with a geometric mean. v1.2.2 reports Jev 1.13.0 at 75.3.
Geometric mean
The four axis scores are multiplied and the fourth root is taken. A weak axis pulls the total down; Luna's cost score of 28.2 is why a 96.8 Intelligence row ranks 12th at 66.0.
classifier.dev fast tier
A hosted classification API whose own page, and JevBench's note, say the fast tier is Jev. It matches Jev on Intelligence and leads the ranking on a cheaper published plan and a faster measured p50 from the bench's server.

Sources

  1. Florian S, introducing JevBench
  2. Florian S, results post
  3. Benchmark Heaven, JevBench v1.2 recap
  4. JevBench interactive ranking
  5. fstandhartinger/jevbench
  6. Benchmark Heaven on MCA 2.3(f)
  7. Logan Markewich, jeff on JevBench