Updated

Desk

Evals

Measurements this digest has already published, with the caveat that came with each one. Vendor tables sit next to independent write-ups and personal clips. We did not rerun any of them. TypeSafe's Master Customer Agreement section 2.3(f) forbids publishing benchmarks of the Services; the numbers below are what people posted. Compare notes on the articles go into more detail.

Decider 1, typed-decisions

0.768 · Jev 0.727 · KL 0.096 vs 1.442

meraGPT blog, September 22, no byline. Post is @latent_node. Model sd-1. The blog says 400 cases, four workflows, five questions each, 2,000 decisions, zero-shot, one run. Jev 1.13.0 through TypeSafe's API. Accuracy / KL / Brier: Decider 1 0.768 / 0.096 / 0.052, Jev 0.727 / 1.442 / 0.148, ModernBERT-base fitted 0.646 / 0.223 / 0.119, MiniLM-L6 fitted 0.587 / 0.262 / 0.143, prior 0.470 / 0.347 / 0.189. By type, noul 0.840 vs 0.775, choice 0.733 vs 0.720, score 0.739 vs 0.696. Labels are the average of three teacher samples. Two samples agree 73.5%. Hugging Face viewer for LocalLLaMA/typed-decisions shows 400 rows in each of four subsets and 1,600 combined. We did not download the files. Median about 0.5 s. Cap 4,096 tokens, at most ten choice labels. $0.03 per million input tokens against Jev at $0.042. Vendor-reported.

Decider 1

dejevu, Zurich to London

5.63 s · published Jev 7.09 s · $0.0092

Idov Mamane, idovmamane/dejevu README, measured 2026-09-22, MacBook Pro M3 Pro, headless Chrome 153, OpenRouter. Median of verified runs. llama-3.3-70b on Groq: 3 of 3, 5.63 s, 10 calls, 9 actions, 15,071 input tokens, $0.0092. Published jev-ultrafast cell, Jev 1.13 plus Mercury 2.5, not rerun here: 3 of 3, 7.09 s, 17 calls, 84,650 tokens, about $0.0036. DeepInfra, same Llama, 2 of 2 at 37.5 s and $0.0020. Five other models verified 0 flights runs. Wikipedia: Groq 3 of 3 at 1.31 s; published Jev cell 2.80 s with the verified count blank. Race overlay 7073 ms vs 6070 ms. GIF caption 6.07 s. Search date moved to Sunday October 18 2026. TypeSafe backend in the repo is untested. Three runs per cell. Independent harness, author-run. The Jev column is the other repo's published figure.

dejevu Flights

Banking feed, 612 private rows

36.1% · 6 s · $0.023

Peter, @LordMarket22, September 22 post. 612 sanitized rows from a harder subset of about 33,000. Jev 1.13: 36.1%, 6 s, $0.023. Gemini 2.5 Flash: 37.3%, 95 s, about $0.088. Gemini 3.7 Flash: 42.5%, 78 s, about $0.180. Gemini dollars are estimates from the same text-token budget at public rates, excluding thinking tokens. No category list, no repository, no row file. Independent harness, author-run. The rows are private.

Banking feed

poker-table, 46 handwritten spots

p50 ~1 s · live pill 545 ms

Jinzhengxu/poker-table README. Spot bank: 46 spots, three repeats, typesafe/jev-1.13 through OpenRouter. Range reads described as matching a reasoning model, 6 to 40 times faster. Median about one second, against 6 seconds single-shot and 35 seconds for the tool-loop agent. Three passes under a cent. No accuracy percentage printed. The September 22 post's live pill is a different clock: 545 ms average and $0.002 a hand. We did not open the live table. The bb/100 table in the same README is a rule-policy ablation, not this Jev run. Independent harness, author-run.

Poker

AgentOS Pilot Router, 121 labeled turns

0.901 · local 0.355 · ~$0.005

AgentOS, September 22 chart in the announcement thread. Corpus tests/data/router_eval/cases.jsonl. 121 turns, Vietnamese, Chinese, and English, gold labels R0 to R3, classifier only, two repeat runs identical. Jev accuracy 0.901, under-routing 0.066, over-routing 0.033, macro-F1 0.894, R3 recall 0.907, ECE 0.062. Local MiniLM 0.355 / 0.413 / 0.231, and the chart leaves macro-F1, R3 recall, and ECE blank. English 0.946 against 0.513. Vietnamese 0.938 against 0.312. Chinese 0.806 against 0.250. One call per turn, about 0.9 s, 121 cases about $0.005. Opt-in. Turn text goes to typesafe.ai. A same-day follow-up, pull 3317, says it sharpens the R2 and R3 criteria. The chart is the announcement image. Independent harness, author-run.

Pilot Router

laya-mlx, one short question on an M3 Max

13.42 ms · 7.39 ms · 63/63

mizorewww/laya-mlx README. M3 Max, 40 GPU cores, 128 GiB, macOS 27.2, MLX 0.32.2, FP16, model load excluded. Laya 421M P50 13.42 ms, P95 13.92 ms, 146.8 questions per second at batch 64, peak 943.6 MiB. Multilingual 322M P50 7.39 ms, P95 7.79 ms, 395.0 questions per second, peak 687.6 MiB. Selected answers matched upstream Laya on 63 of 63 validation questions in FP32 and FP16. Snake at --optimize --max-speed: 75.40 moves per second across 2,400 moves, zero deaths, two safety corrections. The tweet's 60 decisions per second is not that line. No new table against hosted Jev. Independent port, author-run.

laya-mlx

Spring AI TypeSafe, one laptop ticket call

310 ms · cookbook 111 ms

Christian Tzolov, Spring blog, September 21. Three-question support ticket, median 310 ms from his laptop. One question, median 275 ms, 73 output tokens on the three-question call. Blog sample: urgent 0.95, billing confidence 0.82, frustration Score 1.1. The README comments on the same ticket text show different values. He cites TypeSafe's self-consistency cookbook: 14 questions, $0.000043, 111 ms, against Haiku 4.5 at $0.0018 and 1.8 s. The blog says that comparison is TypeSafe's and still needs checking. Judge log on a scripted -455 C answer: plausibility 0.02, groundedness 0.89, helpfulness inconclusive at confidence 0.44. Author-run timings. The cookbook row is vendor-reported.

Spring AI

jev-rubiks, forty move picks

distance 22 to 22

maxli, maxlibin/jev-rubiks README. Jev asked for the best of 18 face turns, forty times. Kociemba distance 22 to 22. A random walk in the same note went 22 to 21. Kociemba in code is about 20 ms and at most 22 moves. The experiment was removed. The coach that shipped asks whether to speak. One screenshot caption: speak 0.83, warn_broke_progress 0.94. Nine live coach situations are described and not printed as a table. Independent harness, author-run.

jev-rubiks

MotherDuck prompt_jev, 100k AG News train rows

89% · 2,484 rows/s · $0.50

MotherDuck, blog dated 2026/09/21, no byline on the page we read. 100,000 articles, reservoir sample seed 43, from the AG News training parquet. Jev 2,484 rows/s, 89% accuracy, $0.50, 40s. gpt-5.6-terra 52 rows/s, 88%, $37.58, 31m 59s. gpt-4o-mini 84 / 80% / $1.93 / 19m 45s. gpt-5-nano 94 / 83% / $1.58 / 17m 49s. gpt-5.6-luna 61 / 84% / $3.53 / 27m 25s. Null choices are not scored as wrong. The null count is not printed. Training split, not a held-out test. Runs at 1 million and 10 million rows are described as similar, with no table. Vendor-reported.

prompt_jev

jev-gateway, 120 chess-agent sessions

-57% output · Opus feature +83% time

Vini Lana. vinilana/jev-gateway-bench, cited from the gateway README. 120 sessions, six models, two chess-engine tasks, five runs with routing and five without, no MCP and no plugins. Medians versus the same model with routing off, output / input / time. Bug-fix: Astra -57% / -7% / -39%, Sol -57% / -40% / -36%, Luna -12% / -10% / +10%, Fable 5.1 -13% / -19% / +6%, Opus 5 -7% / -22% / +2%, Sonnet 5 -41% / -48% / -25%. Feature: Astra 0% / +2% / +8%, Sol -9% / -39% / -16%, Luna -14% / -51% / -14% (3 of 5 solved, against 5 of 5 off), Fable 5.1 -24% / -27% / -26%, Opus 5 +22% / +61% / +83%, Sonnet 5 +9% / +16% / +37%. Claude Code rows are hint mode. Five runs per cell. The README says the bench repo records one run caught copying from another. Independent harness, author-run.

jev-gateway

llamacpp-jev, four questions on one image

526 ms · 32/32 · L4 236 ms

Chirag, NakliTechie/llamacpp-jev. Qwen3.5-2B-Q8_0, unmodified llama-server, 2026-09-22, M4 Pro 24 GB, Metal, GPU idle, four slots, cache-ram 0. Fresh 448 by 448 image, four typed questions, eight images: median 526 ms (525 to 546), 32/32. Same image 254 ms. Four text questions 238 ms. Idle NVIDIA L4, same commit: fresh image 236 ms (174 to 237), 32/32. September 21 on the same Mac with a busy GPU was 946 ms. Six hand-labelled photos, 22/22 scored, two unscored. 64-way questions on 0.8B and 2B both answered option 16. Probabilities are a raw label softmax. Not a JevBench row. Independent harness, author-run.

llamacpp-jev

Open-Jev, 231 public JevBench tasks

179/231 · Jev 200 · Astra 231

Zefan Cai. Audited JSON generated 2026-09-21T06:27:47Z. Public slice of JevBench: 231 of 534 tasks (original 72, easy 48, hard 111). Released Open-Jev 9B 179/231 (77.49%), hard 66/111. 2B 150/231, hard 46/111. Jev 1.13.0 200/231 (86.58%), hard 81/111, one probability vector renormalized. GPT-5.6 Luna 206/231, hard 89/111. GPT-6 Astra 231/231, hard 111/111. P50 189 ms / 291 ms / 954 ms / 2,206 ms for 9B, Jev, Luna, Astra. Open-Jev on one H100, loopback, cache off. Hosted calls are HTTPS. GPT probabilities are verbalized. Candidate order differs on 119 of 139 Choice tasks. Temporal/numeric family: Jev 4/15, Luna 5/15, Astra 15/15. Derived estimates about $0.038 / $0.179 / $8.82 per 1,000 decisions. Not the 534-task composite. Independent harness, author-run.

Open-Jev JevBench

jsort, CommonLit and 95 Fed statements

r 0.824 · $0.046 · rho +0.46

Khaled Eltokhy. CommonLit, 300 excerpts, Jev 1.13 through OpenRouter, 2026-09-19, uncached. jsort -k 10 Pearson r 0.824, Spearman 0.841, 1,500 calls, $0.046. One Noul per text r 0.754. SMOG 0.661. Teacher-scale ceiling near 0.88. Fed opening statements, 95 files, more hawkish about inflation: rank correlation +0.46 with the same-day funds-rate move, +0.37 with the next 180 days. Joe Weisenthal's September 22 post says he rescored 4,005 speeches. That count is not the 95-file script. Independent harness, author-run.

jsort

Security triage, 100 synthetic cases

65.3% · 75.1% · 91.2%

RZ, recorded run 2026-09-19-public-v5. One hundred synthetic TypeScript cases, five passes, 1,500 decisions, zero errors. Jev 1.13.0, gpt-5.6-terra in Codex, claude-opus-4-6 in Claude Code. Balanced accuracy 65.3% / 75.1% / 91.2%. Model-request p50 0.57 s / 10.23 s / 13.83 s. Same answer on all five passes: 95 / 76 / 91 of 100 cases. Jev class recall: insufficient context 93.9%, safe 45.5%, vulnerable 56.5%. Cost for 500 decisions $0.015 / $15.02 / $13.94. Labels by one author. Gap under about fifteen points is noise. Independent harness, author-run.

Security triage

jevsearch, TypeSafe docs ranking

83% Hit@1 · 41% keyword

Kyle McLaren. 109 TypeSafe documentation pages split into 443 documents, 41 labelled queries. jev-search 83% Hit@1, 83% Hit@3, 0.83 MRR@10, median 278 ms uncached. Keyword pass 41% / 59% / 0.51, median 8.7 ms, tied with Lunr. Intent Hit@3 73% versus 42%. 6,169 input tokens per uncached query, about $0.26 per thousand searches. Jev re-ranks a 20-hit pool; it does not retrieve. Independent harness, author-run.

jevsearch

duckdb-jev, live Choice throughput

2,049 rows · 0.887 s

Prasanth J. Live jev-1.13.0 Choice on DuckDB 1.5.5, macOS arm64. 1,000 rows at batch 100 and concurrency 10: median 0.515 s, 1,943 rows/s. Separate capture: 2,049 distinct serialized rows in 0.887 s (2,311 rows/s), cached replay 28.6 ms with zero API calls. Twelve support-ticket templates with unique IDs. Throughput and cache reuse, not general accuracy. Independent harness, author-run.

duckdb-jev

DocJev, 40 PDFs and eight packets

40/40 · 7/8 · 5.73× / 6.45×

Jerry Liu, real-small-v1-run01. Forty authentic public-sector PDFs, eight constructed packets, 116 pages, same LiteParse text for both engines. Classification 40/40 for Jev 1.13.0 and GPT-5.6 Luna; medians 138.6 ms versus 794.3 ms. Splitting 7/8 exact packets versus Luna 8/8; medians 209.6 ms versus 1,352.3 ms. Both found all 32 true boundaries. Jev added one extra cut on a Federal Reserve attachment. Estimated decision cost $0.011663 versus $0.046894. Human labels unreviewed. Small convenience sample. Independent harness, author-run.

DocJev

djev, JevBench hosted row

74.3 · hard 69.5%

Maisa hosted API on DiffusionGemma. JevBench v1.2.8: 74.3, third, Intelligence 88.4, Calibration 65.4, Speed 91.4, Cost 57.6. Hard-tier 69.5% against Jev 74.1%. p50 0.24 s on the production API. Announced $0.035 per million input, $0.026 per 1,000 decisions; nothing charged yet. Probabilities labelled experimental. Independent harness, author-run.

djev JevBench

SimpleJev Qwen3.8-27B, JevBench demo row

67.3 · hard 75.0%

Featherless public demo, logit read, no generation. v1.2.8 rank 13 (Florian's v1.2.6 post had it at rank 9). Intelligence 89.7, hard-tier 75.0% against Jev 74.1%. Cost 39.5, about $0.104 estimated per 1,000. Speed uses the public-demo x2 adjustment. Qwen3.6-35B-A3B 63.8, hard 66.4%. Independent harness, author-run.

SimpleJev JevBench

Kev family, author new-source table

79.6% vs 85.7% · JevBench 66.7

Jared Palmer. Out of domain, data Kev never trained on: Kev-8B 79.6%, Jev 85.7%. MMLU 70 versus 90, date arithmetic 60 versus 93 in the same thread. README later ships Qwen3.5 0.8B / 4B / 9B; Kev-9B new-source 0.812 / 0.837 against Jev 0.857. JevBench independent rows (Qwen3 checkpoints, commit 20fa626): kev 0.6B 66.7, kev 4B 62.2, kev 8B 58.3. Author suite is not JevBench.

Kev JevBench

AstroHan, 30-task context filter

25/30 vs 22/30

Same FrontierHarness 30, same agent, deepseek-flash, one run per arm. pass@1 25 with a Jev keep-or-drop on ~2k-character tool chunks (p > 0.5), 22 without. Cost per solved task $0.113 versus $0.127. meriyah: unfiltered 1.46 million characters, 0/49 tests; filtered 49/49. Input tokens 126 million versus 123 million. Two of the 30 the filter lost. Thread only, no repo.

30 tasks Compaction

JevBench v1.2.8

75.4 · 534 decisions

Florian S / Benchmark Heaven. Geometric mean of Intelligence, Calibration, Speed, Cost at 25% each. Jev 1.13.0 75.4; SemIf (Qwen3.5-4B) 74.7; djev 74.3; openJev Verdict 1.4 72.5; SimpleJev Qwen3.8-27B 67.3. GPT-5.6 Luna 66.2 at rank 19 with Intelligence 96.8 and hard-tier 94.5% versus Jev 74.1%. classifier.dev fast tier 84.8 is Jev behind that host's API, listed not ranked. Self-hosted latency is adjusted ×2 plus 0.15 s. Hard items written by Opus 5 and GPT-5.6 Sol, then frozen. Independent harness, author-run.

JevBench SemIf

LangChain, Jev as a judge

500/500 · $0.34 vs $28.17

Daniel Shea and Seán Roche: five frozen weather-agent traces, 100 repeats. does_pass versus one human oracle: Jev 100%, Terra 99.8%, Luna 96.4%, Claude Sonnet 4.6 80.0%. Quality-score variance: Jev 0.0000149; Luna 433×, Terra 913×, Claude 92×. Per-call cost Jev $0.00035 (0.44 s) versus Luna $0.00039 (2.50 s) and Claude $0.02811. Small corpus. Hosted Jev version not in the metadata.

Judge

Lemkin, SaaStr Connect

32% false admissions

Jason Lemkin: 600 judgments, $0.064 versus about $5 on Sonnet, about 90× cheaper. Agreement with Sonnet 59.5% on 200 borderline-weighted pairs, 67% on 200 random. Blind third-model referee on the random set: Jev 70.5%, Sonnet 77.5%. Jev 47 false admissions of 148 negatives (32%) and 2 false rejections of 52 positives. Thread only, no repo.

SaaStr

Aman Kumar, public sets

16k calls · filter

300 items each from Enron, SST-2, AG News, Banking77 against gpt-5.4-mini and gpt-5.6-luna. Jev leads the three short-text sets and loses Banking77 (76.0% versus 78.7% / 81.7%). Production page gates on 1,005 pages: 96 to 98% agreement, lossless reject filter, 50 to 67% fewer routing calls on replay. A 135-row extraction gate at 68.6%. Inbox 800 mails at 87.4%, no lossless threshold. Public cases in jev-eval. Pipeline rows are not.

Filter

Charly Poly, Banking77 encoders

F1 0.782 · ECE 0.105

77-way intent, human labels, three seeds. Macro-F1: Jev 0.782, ModernBERT-large-zeroshot 0.712, bart-large-mnli 0.453, gliclass-modern-base 0.381. ECE favors ModernBERT (0.081 versus 0.105). Jev assigned probability 0 to the correct label on 6.6% of decisions; the open models did that zero times in 1,155 each. At a 90% precision gate Jev auto-handled 67.5% of traffic versus 35.7%. Hosted BART p50 5,131 ms versus Jev 299 ms. Thread only, no repo.

Encoders

Mathore, 200-piece Tetris

9,200 · ~300 ms · 0 illegal

Same 200-piece sequence, legal landings only. Jev 9,200 points, 75 lines, about 300 ms per move. Claude Haiku 4.5: 8,900 points, 72 lines, 1.52 s, $0.48. Gemini 3.5 Flash-Lite: 9,000 points, 75 lines, 1.13 s, $0.02. One run. Steve's clip is a separate ~400 ms / one cent per 100 pieces note.

Tetris

TypeSafe workflow table

193.6x / 444.6x

Company evals on four internal workflows, the high end of the launch 20-200x / 40-400x range. Built by TypeSafe's model capabilities team. Vendor-reported.

Launch

Every, 777 judgments

0.7s · 6 of 7 defects

Mike Taylor ran 37 documents times 21 writing questions (777 answers) in under 0.7 seconds, about a quarter of a cent. Dan Shipper's later 12-passage defect check against Claude Fable 5.1: median 0.35s versus 8.83s, about 25 times faster and about 580 times cheaper, six of seven planted problems. Independent, small sample.

Every

Bryo, 1,565 emails

96.4% vs 97.5% / 98.5%

Nikhil Mudholkar routed 1,565 German and English supplier emails into 10 categories. Jev 96.4%, Gemini 3.5 Flash-Lite 97.5%, Gemini 3.8 Flash 98.5%. Per 1,000 emails: $0.08 / $0.80 / $1.79. 737 Jev answers at 99% confidence all matched; the most confident 85.5% had zero errors on this set. 364 mails were synthetic. No attachments.

Bryo

Vercel fx auto mode

5-18x vs Luna

Pranit Sharma: a command-safety classifier previously on gpt-5.6-luna was about 5 to 18 times faster and more accurate after swapping Jev. No published accuracy percentages or sample size.

TechCrunch

News Desk Dealer, 384 headlines

24.9s · $0.19

Elvis scored a morning wire for 15 brand desks. Jev finished 384 headlines in 24.9 seconds at $0.19. Opus 5 got through four and cost $0.77 before the run was stopped. Throughput and spend, not agreement with an editor.

384 headlines

SREGym-Lite

24/50 vs 20/50

Jackson Clark: Codex plus gpt-5.6-luna on 10 SREGym-Lite problems, five attempts each. 20 passes without Jev, 24 with jev_plan and jev_submit at a 0.70 threshold. Two problems got worse. Authors did not measure time to diagnosis.

SREGym-Lite

Laya, author table

70.1 JevBench · Banking77 0.425

Convai Innovations. JevBench independent row 70.1 versus Jev 75.3. Author T4: 32.8 ms multilingual, 39.5 ms English, one question. Author quality versus published Jev figures the README says were never measured in that repo: typed-decisions 0.766 vs 0.727 (fine-tuned), Banking77 0.425 vs 0.870 (77 vs 72 labels). Base checkpoints 0.362 / 0.342 against a 0.461 majority baseline.

Laya JevBench

Bespoke Nimble holdout

90.12% vs 93.21%

324 synthetic holdout examples. Bespoke-Nimble-9B matched 292 labels, Jev 1.13.0 matched 302, untuned Qwen3.5-9B matched 215. Labels unreviewed by people. Hemant's Verdict 151M encoder, on a different 337-case TypeSafe public set, scored 48.1% where Jev scored 90.8%.

Nimble

Postgres JOB overlay

12% geomean

Michael Malis: Jev picking join order was about 2x slower. Estimating filters helped some IMDB queries and blew one up. Postgres plans first, Jev overrides when confident: 12% geometric-mean improvement, no dramatic regressions, planner itself slower by hundreds of milliseconds per call. Thread only, no repo.

Postgres

Vogel, 1,500 emails

n = 1,500 · no accuracy

Ryan Vogel classified 1,500 personal emails on a livestream. Sample size is in the post. Precision, recall, and a comparator are not.

1,500 emails

jeff, 1,600 public items

75.5% vs 90.5% AG News

Logan Markewich: 200 items each from eight public datasets. Sequential p50 from a laptop 151 ms on an L4 versus Jev 129 ms. About $2.6 versus about $15.6 per million single-question requests. AG News 75.5% versus 90.5%. Averages: choice accuracy 0.61 versus 0.69, noul AUROC 0.84 versus 0.975. Author-run, public labels.

jeff

SemIf, 21-question 3090 note

1.023 s vs 5.332 s

Theodore Lee: frozen Qwen3.5-4B, 21 binary criteria, median of 3. Direct logit readout 1.023 s and 0 output tokens; compact JSON array 5.332 s and 111 tokens. Argmax agreed on 18/21. Authored balanced accuracy 0.813 on 144 rows. TypeSafe subset agreement 0.845 on 102 aligned rows against a published Jev 0.883; authors did not call live Jev. JevBench is the independent row.

SemIf

GitHub Next, LocalJev bake-off

76.7% short-input · 1,200 requests

Five 4-bit checkpoints on oMLX, 120 gold labels (AG News, BoolQ, SST-5), M5 Max. Short-input macro: Qwen3.6-35B-A3B 76.7%, Gemma 4 26B-A4B 75.0%, DiffusionGemma 26B-A4B 74.2%. The probabilities are prompted JSON. The bake-off does not include hosted Jev, and the authors say it is not a logit read. Long input drops every model's score.

LocalJev

Jev macOS Loop

6/6 · 7.39 s Finder

Josh C. Simmons: six native GUI tasks passed on Vercel and on OpenRouter. Finder batch of nine files in 7.39 seconds through Vercel, 22.69 seconds one file at a time. Median Jev round-trip about 300 ms. Small functional sample. An earlier Calculator run failed and is in the notes.

macOS Loop

term-llm, 30 prompts

n = 30 · anecdotal

Sam Saffron's PR: 30 real session-store prompts on jev-latest, no session context, all of which should reach the model. Zero false switch_session or new_session above threshold. Navigation and status examples at or above 0.88 with untuned criteria. The author calls it anecdotal. No comparator model.

term-llm