Updated

Desk

Decision models like Jev return a choice, a score, or a yes/no.

Jev is TypeSafe AI's System One model. A program sends unstructured text and questions whose legal answers were listed in advance. The model returns a Choice, a Score, or a Noul. It does not write the answer out as text. What Jev is is the page for that call. This page lists other hosted APIs and open checkpoints that publish the same kind of return. The figures are the ones already written up on this desk.

What is a decision model like Jev?

A decision model like Jev takes a state plus a fixed answer set and returns a probability for each allowed answer.

TypeSafe's three names are Choice, Score, and Noul. Choice picks among named options. Score is a number on a rubric. Noul is a yes/no probability from 0 to 1, with no separate confidence field. Choice and Score include a confidence computed from the distribution. Several questions can share one request. Open weights often serve that shape at POST /v1/systemone. A hosted product sometimes uses a different path and still documents the same three types.

What is OpenAI's Decisions API?

On October 6, 2026 OpenAI's Decisions API was in public beta on gpt-6-luna, and a POST from this desk returned HTTP 200.

The September 29 posts say the call can send text or an image. They are written up at OpenAI's Decisions API. A later page, OpenAI Jev, records Tibo Sottiaux on visual inputs and a chart The Decoder opened from DevDay, labeled 150 ms for the Decisions API and 1.6 s for the GPT-6 Luna API. That chart is The Decoder's. On October 6, 2026 OpenAI Developers posted public beta for all developers. The guide, opened that day, prices gpt-6-luna input at $0.10 per million tokens and lists no output charge. A POST to https://api.openai.com/v1/decisions from this desk returned HTTP 200, with 420 input tokens and 0 output tokens. The choice and score answers included confidence. An empty POST returned HTTP 400, missing model. Later the same day a cat sentence scored 0.98 on Decisions and on Jev, and a Stripe ticket was technical on both, with confidence 0.99 and 0.74. October 2 and October 4 calls had returned HTTP 403. The DevDay recap, opened October 6, still said limited preview. The reference page returned HTTP 200.

What is Cloudflare's Jev?

Cloudflare lists TypeSafe's Jev and Cloudflare's own Clef weights as two cards.

typesafe/jev on Workers AI is the hosted TypeSafe model. Access records that card at 32,000 tokens. Clef and Clef-flash are Cloudflare's Apache 2.0 decision models, posted October 1, 2026, written up at Clef. Workers AI prices them at $0.24 and $0.09 per million input tokens. The docs print a 65,536-token window and up to four images. The blog's median row is Clef 209.3 ms, Clef-flash 38.8 ms, and a printed Jev cell of 524.1 ms. Fastino's October 3 chart prints Clef at 61.2 and Clef-flash at 57.1 and captions those Cloudflare scores as self-reported. The model card does not print that composite.

Which other hosted APIs return a typed decision?

Perplexity, Venice, Databricks, Liquid, Nace, Inception, Together, Fastino, Microsoft, and TypeLLM each document a hosted decision call. Access lists doors onto TypeSafe's Jev. The rows below are not all doors onto Jev.

Call Who What the page prints
jev-1.13.0 TypeSafe, POST /v1/systemone $0.042 per million input tokens, output free, on the models page opened October 1, 2026. Pricing.
Decisions API OpenAI, POST /v1/decisions Public beta on October 6. This desk's POST returned HTTP 200. Guide price $0.10 per million input tokens, output uncharged. The recap, opened that day, still said limited preview. OpenAI Jev.
Clef, Clef-flash Cloudflare Workers AI $0.24 and $0.09 per million input tokens. Median 209.3 ms and 38.8 ms against a printed Jev cell of 524.1 ms. Clef.
pplx-decider-v1.1-27b Perplexity, POST /v1/decisions Opened October 7, 2026: $0.02 per million input tokens, output free. Decision Index 0.3 full score 62.8, rank 1, with Jev 1.13.0 as a reference at 60.1. pplx-decider.
jev-latest Venice, POST /api/v1/decisions Billed at TypeSafe's $0.042 rate on the September 18 post. Venice.
ai_decide Databricks SQL and REST Beta. Noul, choice, or score. The docs do not name the model and do not print a price. ai_decide.
d1:free Liquid, POST /decisions/v1/systemone Output tokens always 0. An internal chart prints 58.9 against 57.9. Liquid d1.
Mercury Decide Inception, on OpenRouter Opened October 8, 2026, the JevBench API board ranks this call third at composite 72.4. The free model page returned the same day, priced Free, with up to 14 decisions a second and no JevBench score. Output tokens free. The September 30 post prints no score. The free model URL returned Not Found. Mercury Decide.
Microsoft-Decision-1 Microsoft Foundry and OpenRouter Command Line, October 9, 2026: $0.042 per million input tokens, output free. The post says 4.5 times quicker than Quyet-1.0-Large. The catalog, opened October 10, prints a 32,768-token window and says the weights are not distributed. Microsoft-Decision-1.
typellm-latest TypeLLM, POST /v1/generate Homepage, opened October 10, 2026: input $0.05 per million tokens, thinking $0.50 per million, typed answers free. The September 23 file scores Qwen3.8-27B at 228 of 231 public JevBench tasks with thinking. That file is not typellm-latest. TypeLLM.
Tev1-4B-experimental Together, POST /v1/chat/completions Opened October 7, 2026: $0.04 per million input tokens, output free. The call returns one letter. The September 23 post said $0.042. Tev1.
Drex 1.5 Nace The September 30 post prints $0.04 per million input tokens. Opened October 8, 2026, the product page prints no price. The chart still prints 199 of 231 against Jev at 201 of 231. Drex.
fastino/glide Fastino, POST /v1/systemone Opened October 8, 2026, the pricing page lists $0.15 per million input tokens and $0 output. The earlier docs reading was $0.30. Self-run Decision Index 64.81 against a published Jev cell of 57.91. GLiDE.

Which open-weight models return Choice, Score, or Noul?

The open checkpoints below publish weights and a typed decision. Repos lists the rest of the projects these stories cite, including servers that only wrap someone else's model.

Checkpoint License and base One figure from its own page
Clef, Clef-flash Apache 2.0. Qwen3.8-27B and Qwen3.5-9B. Weights are public. The hosted prices are in the table above. Clef.
Open d1, d1-3B and d1-omni-600M LFM Open License v1.0. LFM2.5-VL-3B and LFM2.5-Encoder-350M. d1-3B card, Decision Index 0.2.1, 48.57, self-scored, not a board submission. The same table prints Winnow-12B at 50.02. Open d1.
Laya Apache 2.0. Encoder, three checkpoints and a router. JevBench 70.1. Banking77 0.425 against a published Jev 0.870. Laya.
Kev Apache 2.0. LoRA plus a pointer head on Qwen. Palmer's new-source table prints Kev-8B at 79.6% against Jev 85.7%. Kev.
Tev1 4B and 0.8B Qwen3.5. The fine-tune license is still being finalized. Training scripts are MIT on Ollama's page. Ollama /v1/systemone. Mean on 3,880 decisions: 73.3% and 63.5%, against Nimble 75.7% and Jev 76.0%. Tev1.
pplx-decider-v1.1-27b The card says a fine-tune of Qwen3.8-27B. Opened October 8, 2026, the README frontmatter lists license apache-2.0. The usage section still says the repository is private. The model API lists private false and gated false. The October 6 post says the weights are open. The card we read has no license line. Card table, overall 61.56 against Jev at 57.9. The space's edition 0.3 full score for this name is 62.8. The board labels it 28B. The card says 26B parameters. The name says 27B. pplx-decider.
Drex DLM CC BY-NC 4.0 weights. MIT code. Backbone NVIDIA Efficient-DLM-8B. POST /v1/systemone. Decision Index 0.2 cell 52.31 against the kit's Jev cell of 51.67. Drex DLM.
PolicyLM-1.7B Apache-2.0. Fine-tune of BidirLM-1.7B-Embedding. A score from 0 to 1 for each category in a policy, up to 16. The card does not serve Choice, Noul, or /v1/systemone. Custom-policy accuracy 0.842 at a 0.5 cutoff. The table has no Jev column. PolicyLM.
Bespoke Nimble 9B LoRA on Qwen. The 324-row holdout is the author's labels. 292 of 324 against Jev 1.13.0 at 302 of 324. Nimble.
imajev-4b Apache 2.0. LoRA on Qwen3.5-4B. Up to two photos. Opened October 8, 2026: JevBench v1.6.1 capability rank 31 at 61.9. Image JevBench v0.3.0 lists capability 68.0. imajev-4b.
Strands Decider 2B Apache 2.0. 1.9 billion parameters. 167 of 231 on the public tasks. Median 115 ms on an RTX 3090. Strands Decider.
NIRNAY 450M Apache 2.0. Laya 421M plus about 30M. phase_b is 0.8792 on 3,080 Banking77 rows. The Jev cell was not measured in that repo. NIRNAY.
WaterSheep Apache 2.0. Fine-tune of ModernBERT-base. 77.8% in distribution, 61.2% on held-out datasets. No Jev column. Adds multi-label. WaterSheep.
StartLux-Decision-27B Code Apache 2.0. Weights CC BY-NC 4.0. Five dense sizes, plus a 35B mixture. Self-run Decision Index 0.2.1 at 63.88, against the September 28 board's Jev cell of 57.91. The README says the run is not on the board. StartLux-Decision.
basal-1.5 Apache 2.0. Bielik base. 1.5B, 4.5B, and 11B. Werdykt v1 macro 72.12%, against Jev 1.13.0 at 81.64%. basal-1.5.
Decision 2.0 Vega Apache 2.0. Six sizes. Vega is an adapter on Qwen3.8-27B. Own Decision Index 0.2.1 run 56.5. The card does not print a Jev row. Decision 2.0.
Vela 2.0 Apache-2.0. Four sizes. The 9B is a fine-tune of Decision-2.0-Lux-9B. The 0.3B tokenizer carries the Gemma Terms of Use. 9B card, Decision Index 0.2.1, 41.63 with Noul calibration on. The default without that option is 41.09. The table does not print a Jev row. Vela 2.0.
fsdecide README says MIT. 149M ModernBERT, trained on five app questions. Review 91.3% of 80, against jev-latest at 56.3%. On the authors' general set, Jev is 90.5% and fsdecide is 60.5%. fsdecide.
JEV-27B-VL Apache 2.0. Adapter on Qwen3.8-27B. Images. POST /v1/decide. 15 of 20 pick-and-place scenes. Six-benchmark mean 84.07 against a hosted Jev cell of 83.85. That JevBench column is not the official board. JEV-27B-VL.
Open-Jev Qwen3.5 LoRA plus a decision head. Released 9B is 179 of 231 on the public tasks, against Jev at 200 of 231. Open-Jev.

Do these scores belong on one board?

No. A Banking77 accuracy, a JevBench composite, and a Decision Index self-run are different tests, and a copied Jev cell is not a new measurement.

Compare places similar tests next to each other and keeps the caveat on the row. Evals collects the numbers with the sample the author printed. Jev benchmark is the JevBench board for TypeSafe's model, which is a separate ledger from the vendor tables above. TypeSafe's Master Customer Agreement section 2.3(f) forbids customers from publishing benchmarks of the Services. This desk reports published figures as published and does not rerun them.