Updated
Desk
Limits
TypeSafe's jaggedness page for jev-1.13 lists nine failure modes. The blocks below keep that list, then add misses that already showed up in stories. We did not reproduce the examples. For similar tests placed next to each other, see compare.
TypeSafe's jev-1.13 list
Sasha Sheng posted that Jev is still too slow, too expensive, and too dumb, and pointed at the jaggedness docs, last reviewed September 17. The page names nine modes and a workaround for each. Numeric examples on that page: a Noul for a refund request at 0.22 next to a yes/no Choice at 0.01 / 0.99 on the same ticket; a Noul and its negation summing to 1.19 on another. Discord is the place the company named for new issues.
- Write the exact condition; put boundary cases in the criteria, or split the judgment.
- Keep arithmetic in code. jev-1.13 does not tally characters, term counts, or long lists.
- Extract dates, then compare them in code.
- Cut extra hops of indirection.
- Filter state before the call. Accuracy falls as unrelated material piles in.
- Test adversarial text; it can steer the answer.
- Align the instruction with the criteria.
- Do not assume two phrasings of the same question agree, and do not reuse a Noul threshold on a Choice.
- Do not ask Jev to generate text. Chaining choices to force generation is slow and weak.
Hard items versus Luna, and false admissions
JevBench's hard tier, 220 frozen items, is 74.1% for Jev 1.13.0 and 94.5% for GPT-5.6 Luna. The composite still ranks Jev higher because Luna's cost score is 28.2. Lemkin's SaaStr Connect trial is a different miss: 47 false admissions out of 148 negatives (32%) against Sonnet's 20 (13.5%), on a job that cannot take weak yeses. LangChain's five-trace judge eval did not show that split; Jev matched every binary label there. RZ's 100-case security demo is the first set on this desk where Jev trailed both frontier models on the headline score: 65.3% balanced accuracy against Terra 75.1% and Opus 91.2%. Safe-class recall was 45.5%. The same run repeated Jev's answer on 95 of 100 cases. On Zefan Cai's 231 public JevBench tasks, Jev 1.13.0 is 200/231 and the hard tier is 81/111, against Luna 89/111 and Astra 111/111. The 15 temporal and numeric tasks in that file are 4/15 for Jev and 5/15 for Luna. Astra is 15/15. That slice leaves out 303 private or judge tasks, so it is a different number from the 534-task composite.
Whole pages and 77-way intent
Kumar's 135-row extraction gate matched later outcomes 68.6%. Probability swung by 0.5 between runs on the same page; he treats that as the task, not the model. On Banking77, 300 items, Jev lost to gpt-5.4-mini and gpt-5.6-luna (76.0% versus 78.7% / 81.7%). Poly's full-set encoder run won macro-F1 (0.782 versus ModernBERT 0.712) and lost Expected Calibration Error (0.105 versus 0.081). Jev assigned probability 0 to the correct label on 6.6% of those decisions; the open encoders did that zero times in 1,155 each. DocJev's 40-PDF classify run was 40/40 for both engines; on eight packets Jev split 7/8 exactly and Luna 8/8, after Jev cut a Federal Reserve implementation attachment that the frozen rules keep with the statement.
Text only
Bryo's 1,565-email table excluded attachments because Jev does not take non-text input. Mudholkar called that the biggest blocker for production. TypeSafe's models page already says input is text. The 364 synthetic mails in that set are a separate limit on the label, not on the model.
Loops that got worse
Malis asked Jev to pick Postgres join order on JOB: about 2x slower. Blind filter estimates helped some IMDB queries and blew one up by an order of magnitude when Jev was unsure. The hybrid he kept (Postgres plans first, Jev overrides when confident) was a 12% geomean with no dramatic regressions; the planner itself got slower by hundreds of milliseconds per call. Clark's SREGym-Lite run went 20/50 to 24/50 with jev_plan and jev_submit; two of the ten problems got worse. Simmons's macOS Loop recorded an earlier Calculator result of 3,968 on 32 x 14, caught by the verifier. AstroHan's context filter lost two of 30 FrontierHarness tasks; meriyah is the one that timed out without it. Vini Lana's chess-engine bench, five runs per cell, cut bug-fix output tokens and then spent more on the feature task for Opus 5 (+22% output, +61% input, +83% time) and Sonnet 5. Luna solved that feature task 3 of 5 times with routing and 5 of 5 without. maxli asked Jev for the best of 18 cube turns, forty times. Kociemba distance stayed at 22. A random walk in the same note went from 22 to 21. The coach that shipped asks whether to speak, and code still solves.
A private feed, mid-30s accuracy
Peter's September 22 post categorizes 612 sanitized banking-feed rows. Jev 1.13 scores 36.1% in 6 seconds at $0.023. Gemini 2.5 Flash scores 37.3% in 95 seconds. Gemini 3.7 Flash scores 42.5% in 78 seconds. Gemini costs are estimates and exclude thinking tokens. The post prints no category list and no repository. Banking77 on this desk is a different set, with Jev in the 0.78 macro-F1 range.
Teacher-average labels
meraGPT's Decider 1 blog puts Jev 1.13.0 at accuracy 0.727 and KL 1.442, against Decider 1 at 0.768 and 0.096. The choice gap on that table is 0.013. The reference is the average of three teacher-model samples, and the blog says two samples agree 73.5% of the time. One evaluation run. The blog counts 400 cases and 2,000 decisions. The Hugging Face viewer lists 1,600 rows. We did not recompute either figure. Decider 1 is hosted, caps a request at 4,096 tokens, and allows at most ten choice labels.
Local copies trail hosted Jev
Jev's weights are not public. SemIf's Qwen3.5-4B row on JevBench is 74.7 against Jev 75.4, and 59.5% versus 74.1% on the hard tier. djev, a hosted DiffusionGemma API, is 74.3 with hard-tier 69.5% and Calibration 65.4. Laya is 70.1 on that suite. Laya's own table, against published Jev figures the README says were never measured in that repo, trails on Banking77 (0.425 versus 0.870, 77 labels versus 72) and scores 0.000 on Khmer at 0.952 confidence on the English checkpoint. Nimble matched 90.12% of 324 synthetic labels against Jev at 93.21%. jeff scored AG News 75.5% versus Jev 90.5% on 1,600 public items, and 66.9 on JevBench. Verdict scored 48.1% on 337 TypeSafe public cases where Jev scored 90.8%; the 1.4 engine is 72.5 on JevBench. Palmer's Kev-8B scored 79.6% on his new-source split against Jev 85.7%, with MMLU 70 versus 90 and date arithmetic 60 versus 93; JevBench listed kev 0.6B at 66.7. SimpleJev Qwen3.8-27B is 67.3 on the composite and 75.0% on the hard tier. Open-Jev's released 9B checkpoint, a different project from SemIf, is 179/231 on the public tasks against Jev 200/231, and 66/111 on that slice's hard tier. LocalJev's bake-off compares local checkpoints with each other, not with hosted Jev, and the authors say the JSON probabilities are not a logit read. llamacpp-jev is another /v1/systemone server on stock llama-server. Its probabilities are a raw label softmax, and on 64-way questions Qwen3.5-0.8B and 2B both answered option 16. The 32/32 figure is eight synthetic images, not JevBench. laya-mlx runs Laya's checkpoints in MLX. On 63 validation questions the selected answer matched upstream Laya. That check is not a score against hosted Jev, and the port does not publish one.