Updated

Desk

The public Jev benchmark is JevBench. The API board opened October 8, 2026 ranks Jev 1.13.0 fourth at composite 71.5.

Two different tests get called a Jev benchmark. JevBench is Florian Standhartinger's suite, published at benchmarkheaven.com. TypeSafe's launch post has a separate workflow table that the company ran on its own model. The second one is a vendor self-test, and this desk has not independently reproduced it. What the model returns is on what Jev is.

What does JevBench print for Jev?

Opened October 8, 2026, the API board is still release v1.6.1, revision v1.7.21, and ranks 26 hosted systems. Jev 1.13.0 is fourth at composite 71.5. Capability on that card is still 77.1, rank 2. Sage 1.3.0 is still 74.0. Liquid d1 is still 73.0. Mercury Decide is third at 72.4. We reopened the published page and did not rerun the suite.

Opened October 7, 2026, the API board is release v1.6.1, revision v1.7.12. That note says the headline scores are unchanged. Sage 1.3.0 is composite 74.0. Liquid d1 is 73.0. Open d1's 48.57 is the d1-3B card's self-run of Decision Index 0.2.1, a different test from this 73.0. Jev 1.13.0 is third at 71.5. Capability on that row is 77.1, from Intelligence 63.6 and Calibration 90.6, at $0.032 per 1,000 decisions and p50 0.24 s. On the open-weights board the same row is an unranked reference. 72.1 stays the v1.5 composite. We did not rerun the suite.

JevBench v1.5.0, posted September 28, scores 1,624 decisions. That v1.5.0 page ranks Cygnet 73.7, Winnow-12B Q8 73.2, and Jev 1.13.0 72.1, and it calls the top pair a statistical tie.

Jev's capability score on that page is 80.0, first among 56 Jev-class systems. The older v1.4.2.2 harmonic mean still has Jev at 63.29. That 63.29 is not 72.1, and it is not the older geometric mean of 75.4. The page we opened on September 30, and again on October 1, said v1.5.4 and 106 ranked systems. Those three scores stayed on top. The story, including the caveat, is JevBench v1.5.

imajev-4b is Mohit Garg's open 4B checkpoint. On the v1.5.6 page opened October 4 it is official rank 10 at 70.4. The older v1.4.2.2 composite still prints 67.37. On the v1.6.1 open-weights capability list opened October 7, the same checkpoint is rank 17 at 61.9, Intelligence 34.6 and Calibration 89.2, estimated $0.017 per 1,000 decisions, p50 0.29 s. Opened October 8, revision v1.7.21, that card is still 61.9, capability rank 31, official composite 47.

Which numbers are the vendor's own test?

TypeSafe's launch multiples are the company's own measurement, and this desk has not reproduced them.

The post cites about 20 to 200 times faster and about 40 to 400 times cheaper than language-model workflows. The high end it prints is 193.6 times and 444.6 times, on four workflows, and the company says that high end is likely the top of real use. On September 30 this desk scored a different set: 32 labeled tickets. jev-1.13.0 and gpt-6-luna were each right on 32 of 32. p50 was 913.5 ms and 2,698.5 ms. That run did not include the four workflows, so it does not confirm 193.6 times or 444.6 times. The notes are on the launch story. TypeSafe's Master Customer Agreement section 2.3(f) forbids customers from publishing benchmarks of the Services. Method says how this desk treats numbers people posted anyway.

What is a public-task count, not the composite?

Open-Jev scored 231 of JevBench's public tasks, which is not the full composite and not the sealed set.

On the September 21 audit the counts were Jev 200, Luna 206, and Astra 231. The hard tier was 81, 89, and 111 of 111. A later 27B v1.1 card scores 197 of those 231, and 80 of 111 on the hard tier, with Jev still at 200 and 81. The page is Open-Jev's public tasks. Laya's README cites an earlier JevBench row, Laya 70.1 against Jev 75.3, and says the Jev cell was not measured in that repo. That pair is on Jev vs Laya. A Snowflake package with one live account, and no accuracy table, is Jevflake. Who could sign up to call the hosted model is the signup thread.

Where are the other measurements?

Evals lists the measurements already in the stories, each with the caveat its author published.

Compare places similar tests next to each other and does not treat the rows as one ranking. The evals topic is every story tagged that way. We did not rerun the tables on either page.

Which Jev score should I quote?

Name the board. Opened October 8, 2026, the API board ranks Jev 1.13.0 fourth at composite 71.5, capability 77.1. On October 7 the same board ranked Jev third at that score. JevBench v1.5 prints Jev 1.13.0 at 72.1. The older v1.4.2.2 harmonic mean prints 63.29. Those are different formulas, and 63.29 is not the v1.5 score.

Did Jev News rerun JevBench?

No. The ranking is Florian Standhartinger's published page. This desk did not rerun the suite. TypeSafe's separate workflow multiples were not reproduced here either.

Are the 200x figures a JevBench score?

No. Roughly 20 to 200 times faster, and roughly 40 to 400 times cheaper, are TypeSafe's own test on four workflows. JevBench is a different suite, run by someone else.