Desk
Jev is TypeSafe's System One model, and it returns a typed decision.
Jev is TypeSafe AI's System One model. A program sends unstructured text and questions whose legal answers were listed in advance. The model returns a Choice, a Score, or a Noul, each with a probability, and it does not write the answer out as text. TypeSafe announced it on September 15, 2026. This page is not a TypeSafe page.
Tev1, opened October 7, 2026, is $0.04 per million input tokens on Together's page. That chat call returns one letter. Ollama serves the same name as Choice, Score, or Noul. The catalog is decision models. Decision Index 0.3, opened the same day, ranks Perplexity Decider v1.1 first at 62.8 and lists Jev 1.13.0 as a reference at 60.1. Nace's open Drex DLM card scores edition 0.2 at 52.31 against the kit's Jev cell of 51.67.
The table puts that return next to an ordinary language-model call. Cells are claims already published on this desk. The speed multiples in the launch post are the vendor's own test, and this desk has not reproduced them.
| Question | Jev | A language model |
|---|---|---|
| Output | A Choice, a Score, or a Noul. No generated string. | A token stream. JSON mode is still tokens the caller parses. |
| Latency | TypeSafe cites 70 to 500 ms. On September 30 this desk measured p50 913.5 ms on 32 tickets. | The same 32 tickets: gpt-6-luna p50 was 2,698.5 ms. TypeSafe's 20 to 200 times line is the company's, on other workflows. |
| Cost | $0.042 per million input tokens, output free, on the models page opened October 1, 2026. The 32-ticket list cost was $0.000529. | The 40 to 400 times claim is TypeSafe's. The Luna side of that 32-ticket run returned token counts and no dollar amount. |
| Structured errors | The caller names the options. TypeSafe says an unlisted option cannot come back. | On one email in that run, three of eight serial Luna replies wrote a sentence instead of the requested label. |
| Confidence | Choice and Score return a confidence from the distribution. Noul is a probability from 0 to 1, with no separate confidence field. | A chat reply can print a number if the prompt asks. That number is generated text. |
What is Jev, in one sentence?
Jev is TypeSafe's System One model: state goes in, a typed decision comes out, and the model does not write text.
Diogo Almeida announced it on September 15, 2026, after two years in stealth. The name points at the Jevons paradox, the claim that a cheaper decision gets used more often. The weights stay with TypeSafe. Other hosted calls and open checkpoints are on decision models like Jev. StartLux-Decision-27B prints 63.88 on a self-run of Decision Index 0.2.1. basal-1.5 prints 72.12% on Werdykt v1, against Jev 1.13.0 at 81.64%. Decision 2.0 Vega prints 56.5 on its own Decision Index 0.2.1 run, and the card does not print a Jev row. Vela 2.0's 9B card prints 41.63 with Noul calibration on, and 41.09 with it off. The evaluation file prints Lux at 46.23 and does not print a Jev row. fsdecide scores 91.3% of 80 review items against jev-latest at 56.3%. JEV-27B-VL finishes 15 of 20 pick-and-place scenes. Open d1's d1-3B card prints 48.57 on a self-run of Decision Index 0.2.1. That table has no Jev row. Step-Jev trains its own judges on the agent's traces. On one BrowseComp-Plus seed, grading every step scores 19.6% and grading only the final answer scores 5.4%. A bag of Jev questions moves a published WANDS mean from 0.574 to 0.610 and leaves the median at 0.561. Microsoft-Decision-1 lists $0.042 per million input tokens and says it was 4.5 times quicker than Quyet-1.0-Large. The October 9 post does not print an accuracy, and the Foundry catalog says the weights are not distributed. TypeLLM scores Qwen3.8-27B at 228 of 231 public JevBench tasks with thinking, and 195 of 231 without it. The hard tier in that file moves from 76 of 111 to 109 of 111. The homepage prices input at $0.05 per million tokens. The hosted name is typellm-latest, and the 228 of 231 run is not that model. What Jev returns is the explainer. The launch story is Almeida's thread and the company's workflow table.
How Jev differs from an LLM
A language model writes the next token. Jev returns one of the outputs the caller already named.
TypeSafe says the questions in one request are sampled together, and it cites 70 to 500 ms for that. A chat model writes the reply one token at a time. On September 30, 2026 this desk scored 32 labeled tickets, eight each in billing, sales, spam, and technical. jev-1.13.0 and gpt-6-luna were each right on 32 of 32. p50 wall time was 913.5 ms for Jev and 2,698.5 ms for Luna. Jev has no question type that returns a sentence. The notes are on the launch story.
The three primitives
TypeSafe's primitives page names three returns: Choice, Score, and Noul.
Choice picks among up to 255 named options and returns a distribution plus a confidence. Score returns a number on a rubric the caller defines, with a confidence. Noul is a yes or no. It returns a probability between 0 and 1 and has no separate confidence field. Several questions can share one request. The confidence figure is computed from that distribution. It is not a second model checking the first. The homepage diagram links the same three names, and the explainer defines them.
What Jev is good at — and what it isn't
Jev fits a call whose legal answers are known before the request goes out.
Routing, labeling, and a yes or no inside a larger loop are the jobs in the stories. Patterns collects the shapes that show up more than once: rank a legal set, judge a pile of documents, act only when the confidence is high, send the expensive call elsewhere, or run a different model locally. Jev does not write prose, code, or a reason. A product that needs a sentence still asks a language model for the sentence. Limits keeps Sasha Sheng's jaggedness list: jev-1.13 misses arithmetic, date comparison, and two phrasings of the same question.
TypeSafe's speed and cost multiples are the vendor's own test, and this desk has not independently reproduced them. The launch post cites about 20 to 200 times faster and about 40 to 400 times cheaper than language-model workflows, with a high end of 193.6 times and 444.6 times on four workflows. Those four workflows are not in the 32-ticket set. JevBench, Laya's tables, and the other rows on evals were published by their authors. We did not rerun them. TypeSafe's Master Customer Agreement section 2.3(f) forbids customers from publishing benchmarks of the Services. Method records how this desk treats that rule.
Pricing and speed: read the fine print
Opened October 1, 2026, TypeSafe's models page lists jev-1.13.0 at $0.042 per million input tokens, with output free.
The same page lists 100K tokens per second and 40 requests per second, a 64k limit on the whole request, and 32k for the state plus the longest question. It says those limits can change without notice. Opened October 7, 2026, the same page lists 80 requests per second. The price, the token rate, and the two context caps were unchanged. The 70 to 500 ms band is TypeSafe's latency claim. The 32-ticket p50 of 913.5 ms includes the round trip from one machine. Other hosts publish other rates. Vercel's card listed the model as Free, and on September 26 it still printed a promotional end date of September 25, 2026. New TypeSafe signups paused on September 22 and reopened on September 27 without the $5 credit. On October 6, 2026, at 17:05 UTC, a new console account showed available credits of $0.00. On October 8, 2026 the same signed-in console still showed available credits of $0.00, no payment method, and no billing history. That thread is signups paused, then reopened. Host-by-host notes are on Jev pricing.
Where Jev fits in an agent loop
In the loops on this desk, Jev chooses and some other program acts.
LangChain's ModelRouterMiddleware asks Jev which chat model should take the turn. AutoModeMiddleware can block a tool call it scores as risky. Both are marked experimental. Vercel's form router keeps the destination when confidence is at least 0.95 and sends the rest to gpt-5.6-luna-fast. A DSPy patterns repo refuses a shell command below 0.2 and runs it at 0.6 or above. That cutoff is a constant in one file, not a TypeSafe default. The longer reading is patterns. The Python clients are on Jev in Python.
The bottom line
Jev answers a closed question. A language model still writes any sentence the product has to show.
The list price, as opened October 1, 2026, is $0.042 per million input tokens, and output is free. The speed multiples in the launch post are the company's, and they have not been reproduced here. A console key, the request path, and the signup check are on the Jev API page. Why the open checkpoints start from a public language model is on Open decision models. Laya's open weights are on Jev vs Laya. The public suite is on the Jev benchmark page. What the launch post means by can't hallucinate is on decision model hallucination.
FAQ
Is Jev the same thing as Jev News?
No. Jev is TypeSafe AI's model. Jev News, at jevainews.com, records decision models, and it is not affiliated with TypeSafe.
Does a Noul include a confidence field?
No. A Noul is a probability between 0 and 1. Choice and Score are the returns that include a separate confidence, computed from the distribution.
Are the 200x figures an independent benchmark?
No. The roughly 20 to 200 times speed line and the roughly 40 to 400 times cost line are TypeSafe's own workflow test. This desk has not reproduced those four workflows.
Can I download Jev's weights?
No. The weights stay with TypeSafe. Open checkpoints that return a choice, a score, or a yes/no are indexed on the repos page. Laya is the comparison written up beside this one.
Stories grouped by tag are on topics. Similar tests sit next to each other on compare. The doors this desk has cited are on access. Short answers from the same material are on the FAQ.
Latest stories
- TypeLLM scores Qwen3.8-27B at 228 of 231 public JevBench tasks when thinking is on (Oct 10, 2026)
- On WANDS, more Jev questions move the mean to 0.610 and leave the median at 0.561 (Oct 9, 2026)
- Microsoft-Decision-1 lists $0.042 per million input tokens and calls itself 4.5 times quicker than Quyet-1.0-Large (Oct 9, 2026)