Published
Updated
TypeLLM scores Qwen3.8-27B at 228 of 231 public JevBench tasks when thinking is on
TypeLLM's September 23, 2026 file scores Qwen3.8-27B at 228 of 231 public JevBench tasks with thinking, and 195 of 231 without it. The hard tier accounts for the gain, 76 of 111 to 109 of 111, while easy falls from 48 of 48 to 47 of 48. The homepage prices input at $0.05 per million tokens and thinking at $0.50. The hosted model name is typellm-latest, and this file is not that model. We did not call the API.
TypeLLM is a Python library, and a hosted API, that keeps an autoregressive model and constrains each field to a JSON Schema. The September 17 introduction says the project was prompted by TypeSafe’s Jev. Jev is a purpose-built decision model. TypeLLM’s pitch is the open models a caller already serves. The library is on GitHub. PyPI typellm 0.6.8 was uploaded at 08:41 UTC on October 9, 2026. The license expression is Apache-2.0. It requires Python 3.10 or newer. The project owner field reads rejojer. The pages we opened do not give that account a personal name.
The homepage table is the September 23 self-hosted run. The file scores RadixArk/Qwen3.8-27B-NVFP4-BF16-LMHead on 231 public JevBench tasks. No thinking: 195 correct. Thinking: 228 correct. The homepage prints 84.42% and 98.70%. Those are 195/231 and 228/231 at two decimal places. The method note says this is one run per setting, on the public subset, and not a full 534-task board score. Schema validity is 231 of 231 in both files. We did not run verify.py.
The 33-task gap is the hard tier. Hard goes from 76 of 111 to 109 of 111. Easy goes from 48 of 48 to 47 of 48. Original goes from 71 of 72 to 72 of 72. Those three movements net to 33, which is 228 minus 195. The method splits the 231 tasks into 72 Original, 48 Easy, and 111 Hard, and into 139 Choice, 74 Noul, and 18 Score questions. The upstream pin is f8ce71361165846101d02ebc83ad44e47ae44fc3.
Four families in the same files supply 25 of the 33. Temporal numeric goes from 8 of 15 to 15 of 15. Probability goes from 4 of 10 to 10 of 10. Multi-hop goes from 12 of 18 to 18 of 18. Long policy goes from 12 of 19 to 18 of 19. Extraction goes the other way, from 24 of 24 to 23 of 24. The file does not say whether that lost extraction item is the lost easy item.
The three thinking misses sit in the high bin. The thinking summary’s bin from 0.9 to 1.0 holds 225 tasks, and 222 of them are correct. The printed accuracy on that bin is 0.9866666666666667. The other bins in that run, two tasks from 0.5 to 0.6 and four from 0.8 to 0.9, are all correct. The three errors are among the answers whose top probability was at least 0.9. Mean confidence in that bin is 0.9976.
Brier falls from 0.240853 to 0.029395. Top-label ECE falls from 0.052494 to 0.016979. Score ordinal MAE falls from 0.272135 to 0.000002. Paraphrase agreement is 35 of 36 pairs without thinking and 36 of 36 with it. The method says accuracy uses the highest-probability label, and that the typed score is the probability-weighted level. The 18 score questions are mapped to an integer enum. There is no permutation averaging in this run.
Latency moves the other way. Median request time goes from 1.296 seconds to 6.677 seconds, about 5.2 times. The 95th percentile goes from 2.150 seconds to 61.083 seconds, about 28 times. The method says those clocks include the client and an SSH tunnel, and that the thinking run reused a warm cache. Thinking tokens in the scored responses total 212,255. Divided by 231, that is 918.9 tokens a task, the average the eval README prints. The run set thinking_budget to null. All 231 thinking answers closed on their own. The local seed of 42 does not seed the server-side reasoning. Temperature on that reasoning is 0.6, with top-p 0.95 and top-k 20. Classification itself is argmax.
The homepage also prints Jev 1.13.0 at 86.58%, Open-Jev 27B v1.1 at 85.28%, GPT-5.6 Luna at 89.18%, and GPT-6 Astra at 100.00%. At two decimal places those four percents are 200, 197, 206, and 231 of 231. The eval README says the external rows come from Open-Jev and that the settings differ. Open-Jev’s row in that file is the same Qwen3.8-27B base with a rank-8 LoRA and a decision head, in BF16. TypeLLM’s row is NVFP4 weights, a BF16 language-model head, and no extra training. We did not open the Open-Jev page. This site’s JevBench composite, and Drex’s 199 of 231, are other tables.
The hosted API names a different model. Docs say it serves typellm-latest, which reads text and images and can think. The caller can leave model out. GET /v1/models is documented as returning that one name. The September 23 checkpoint name does not appear on that page. We did not call https://api.typellm.ai/v1/generate.
The homepage prices input at $0.05 per million tokens and thinking at $0.50 per million. Typed answers are free. New users get $5. At those rates, and with no thinking tokens, $5 pays for 100 million input tokens. With no input tokens, $5 pays for 10 million thinking tokens. That is rate-card arithmetic, not an invoice. TypeSafe’s models page, in the October 7 reading on this desk, lists $0.042 per million input tokens. $0.05 is 1.19 times that figure. The two cards are separate. The 212,255 thinking tokens in the September 23 file, priced at $0.50 per million, come to $0.106. The summary marks the run self-hosted and unmetered, so $0.106 was not charged there.
Hosted limits on the API page, opened October 10: 64 questions a call, 128K tokens of input, a 128 MB body, eight images, and string answers capped at 128 tokens. A thinking budget defaults to 1,024 tokens and may not exceed 4,096. The September 23 run used no budget, so its 228 of 231 is not a measurement of that cap. Timeout defaults to 60 seconds and may not exceed 90. The account rate is 200 requests a minute, shared by its keys. The playground line says 20 a minute before an access code. The library README says a self-hosted caller can pass text_max_tokens for a longer string. The output-types page calls 128 tokens a fixed service limit. Those two sentences describe different places to run.
@TypeLLM posted on September 29 that the playground was open with $5 of credit, and that the API was rolling out to early-access users. On October 7 the same account posted that OpenAI Decisions code runs on TypeLLM after the caller changes the API key and the base URL. The API reference says POST /v1/decisions uses base https://api.typellm.ai/v1, and POST /v1/systemone uses base https://api.typellm.ai. Both are answered and billed as /v1/generate. We did not send either call.
The docs models page lists five self-host rows: Qwen3.8-27B, Qwen3.5 at 0.8B, 4B, and 9B, MiniCPM5-1B, Ling-mini-2.0 with thinking off only, and Ring-mini-2.0 with thinking always on. Image input is tested with Qwen3.8-27B. The GitHub README’s supported-models table, read the same day, names the Qwen rows and does not name MiniCPM, Ling, or Ring. The README’s comparison table says rubric scoring is a numeric enum and that TypeLLM has no dedicated Score API. Further up, the same README documents a levels field that returns the probability-weighted index of the levels. The output-types page has a Score row with the same shape. Both sentences were on the pages we opened.
A field can name depends_on. The child then sees the parent’s typed answer, and a cycle is rejected before inference. The value is still only guaranteed to land inside the child’s own domain. A wrong parent is an input to the child, and the schema does not correct it. when tests those typed answers in code and adds no model request. A failed test skips the field. The README’s confidence formula is (p_max - 1/n) / (1 - 1/n). On the printed three-way example, probabilities 0.04, 0.93, and 0.03, that formula returns 0.895. The README prints confidence 0.9.
The September 22 die post is a different measurement from the 231 tasks. One Jev call, options written one through six, puts 83% on one. Across all 720 orderings, Jev’s highest probability stays on the label one every time, and lands on the first option 120 times. 120 of 720 is 1/6, which is how often one sits in front if the winner is always that label. TypeLLM plus Qwen, thinking off, puts the highest probability on the first option in 718 of 720 orderings, and on the label one in 122 of 720. The two counts are each two away from the “always the label” pair of 720 and 120. The post does not print those two orderings, so the overlap is a count check. After averaging all 720 orders, Qwen’s one is 25.41% and Jev’s one is 86.01%, against a fair 16.67%. Eight sampled orders, the post says, cut Qwen’s KL error by 79% relative to its average single-order error. All 720 cut it by 97%. Jev moved only slightly. The script named in that post expects the same NVFP4 checkpoint. We did not roll the die.
Two letter-count traces are also different prints. The thinking docs show 68 thinking tokens for strawberry. The example page shows 157 input tokens, 123 thinking tokens, 0.84 seconds, and a cost of $0.000069. At $0.05 and $0.50 per million, 157 input tokens plus 123 thinking tokens come to $0.00006935, which rounds to the printed cost. The 68-token trace is a different line.
Other hosted calls of this shape are on decision models like Jev.
This site's reading
Editorial notes evaluating claims against primary sources, contextualizing findings alongside related implementations, and defining technical terms.
Verify
Homepage https://typellm.ai/, opened October 10, 2026. The benchmark table prints TypeLLM + Qwen3.8-27B at 84.42% with no thinking and 98.70% with thinking, on 231 public JevBench tasks. The same table prints Open-Jev 27B v1.1 at 85.28%, Jev 1.13.0 at 86.58%, GPT-5.6 Luna (none) at 89.18% with type-safe output marked No, and GPT-6 Astra (low) at 100.00% with type-safe output marked No. Pricing on that page: input $0.05 per million tokens, thinking $0.50 per million tokens, typed outputs free. New users get $5 to start.
GitHub README, same day, https://github.com/TypeLLM/TypeLLM. The September 23 update line says 195 of 231 without thinking and 228 of 231 with thinking. 195/231 is 84.42% at two decimal places. 228/231 is 98.70%. The eval README dates the runs 2026-09-23 and says external results are reported by Open-Jev, with model and inference settings that differ. It says this is not the full 534-task benchmark. Both runs use RadixArk/Qwen3.8-27B-NVFP4-BF16-LMHead. Open-Jev 27B v1.1, in that file's comparison, is Qwen3.8-27B with a rank-8 LoRA and a scalar decision head. TypeLLM's row says additional training is none.
summary.json for the two runs. No thinking: 195 correct of 231, Brier 0.240853, top-label ECE 0.052494, score ordinal MAE 0.272135, p50 1.296 seconds, p95 2.150 seconds. Thinking: 228 correct of 231, Brier 0.029395, ECE 0.016979, ordinal MAE 0.000002, p50 6.677 seconds, p95 61.083 seconds, 212,255 thinking tokens, 231 answer tokens. Schema validity is 231 of 231 in both files. price_per_1000_decisions_usd is null, and the cost basis string is self_hosted_unmetered. Tiers, no thinking: original 71 of 72, easy 48 of 48, hard 76 of 111. Thinking: original 72 of 72, easy 47 of 48, hard 109 of 111. Paraphrase pairs: 35 of 36 agree without thinking, 36 of 36 with thinking. The thinking ECE bin from 0.9 to 1.0 holds 225 tasks at accuracy 0.9866666666666667, which is 222 of 225. The other thinking bins in that file are correct. We did not rerun verify.py.
METHOD.md. GPU is an NVIDIA RTX PRO 6000 Blackwell Server Edition. Serving is SGLang 0.5.19. Tasks are 72 Original, 48 Easy, and 111 Hard, and 139 Choice, 74 Noul, and 18 Score. Upstream revision f8ce71361165846101d02ebc83ad44e47ae44fc3. Argmax, no permutation averaging. Thinking uses thinking_budget null, temperature 0.6, top-p 0.95, and top-k 20. The local seed 42 does not seed server-side reasoning. One run per setting. Timings include an SSH/IAP tunnel. The thinking run reused the cache.
API reference, https://typellm.ai/docs/api. Hosted model typellm-latest. POST https://api.typellm.ai/v1/generate. Input $0.05 per million tokens, thinking $0.50 per million, answers free. Limits printed there: 64 questions, 128K input tokens, string answers 128 tokens, thinking budget 1,024 by default and at most 4,096, timeout 60 seconds by default and at most 90, 200 requests per minute per account. The playground line says 20 requests per minute before an access code. Compatible endpoints: POST /v1/decisions with base URL https://api.typellm.ai/v1, and POST /v1/systemone with base URL https://api.typellm.ai. Each of those is billed as /v1/generate. We did not send a request.
Docs models page lists Qwen3.8-27B, Qwen3.5 0.8B/4B/9B, MiniCPM5-1B, Ling-mini-2.0 (off only), and Ring-mini-2.0 (always on). The GitHub README's supported-models table, opened the same day, names the Qwen rows and does not name MiniCPM, Ling, or Ring. The README comparison table says rubric scoring is a numeric enum and that there is no dedicated Score API. The same README documents a levels field. The output-types page has a Score row.
Fair-die post, September 22, 2026. One Jev call prints one 83%, two 1%, three 4%, four 3%, five 1%, six 8%. Across 720 orderings, Jev's highest probability is the label one in 720 of 720, and the first option in 120 of 720. TypeLLM + Qwen, thinking off: the label one in 122 of 720, the first option in 718 of 720. After a 720-order mean, Qwen's one is 25.41% and Jev's one is 86.01%. The post says eight permutations cut Qwen's KL error by 79%, and all 720 cut it by 97%, against its average single-order error. We did not rerun it.
PyPI typellm 0.6.8, uploaded 2026-10-09T08:41:47Z. License expression Apache-2.0. Requires Python >=3.10. Project owner field rejojer. September 29 post 2104934225480708133 says the playground is open with $5 credit. October 7 post 2107745238513312154 says OpenAI Decisions code runs after the key and the base URL change.
Compare
The two summary files are one public subset, one checkpoint, two settings. Hard moves from 76 of 111 to 109 of 111. Easy moves from 48 of 48 to 47 of 48. Original moves from 71 of 72 to 72 of 72. The net is 33 tasks, which is 228 minus 195. Four families supply 25 of those 33: temporal_numeric 8 of 15 to 15 of 15, probability 4 of 10 to 10 of 10, multi_hop 12 of 18 to 18 of 18, and long_policy 12 of 19 to 18 of 19. Extraction goes from 24 of 24 to 23 of 24. Those family counts are in the same two files. They are not a second benchmark.
The homepage's 86.58%, 85.28%, 89.18%, and 100.00% are the external rows. At two decimal places, 86.58% of 231 is 200, 85.28% is 197, and 89.18% is 206. The eval README attributes those rows to Open-Jev and says the settings differ. This desk did not open that Open-Jev page. JevBench's hosted composite on this site, and Drex's 199 of 231, are other tables.
Hosted thinking, in the API reference, defaults to a 1,024 token budget and refuses a budget above 4,096. The September 23 thinking run set no budget. Its 212,255 thinking tokens, priced at the homepage's $0.50 per million, come to $0.106. The summary file marks that run unmetered. The $0.106 figure is this desk's rate-card arithmetic, not a bill. TypeSafe's models page, in the October 7 reading on this desk, lists $0.042 per million input tokens. $0.05 is a separate card.
Terms
- 228 of 231
- Correct tasks in TypeLLM's September 23, 2026 thinking summary for RadixArk/Qwen3.8-27B-NVFP4-BF16-LMHead on 231 public JevBench tasks. The no-thinking file is 195 of 231. The homepage prints those rates as 98.70% and 84.42%. The hosted name typellm-latest is not this checkpoint.
- Hard tier, 109 of 111
- The thinking summary's hard tier. The no-thinking file has 76 of 111. Easy in those files is 48 of 48 without thinking and 47 of 48 with thinking. Original is 71 of 72 and then 72 of 72.
- $0.05 and $0.50 per million tokens
- Input and thinking prices on the TypeLLM homepage and in the API reference, opened October 10, 2026. Typed answers are free. New accounts get $5. A field that thinks on the hosted API gets a 1,024 token budget unless the caller sets another, up to 4,096.
- 718 of 720
- Orderings, in the September 22 fair-die post, where TypeLLM + Qwen gave the highest probability to the first option. Jev gave it to the label one in 720 of 720 orderings, and to the first option in 120 of 720. That post is not the JevBench file.
Sources
- TypeLLM homepage
- TypeLLM on GitHub
- JevBench public subset, September 23
- Method and configuration
- No-thinking summary
- Thinking summary
- Can Jev roll a die?, September 22
- Introducing TypeLLM, September 17
- API reference
- Supported models
- Output types
- Thinking
- Letter-count example
- typellm 0.6.8 on PyPI
- TypeLLM, September 29
- TypeLLM, October 7