Published
Updated
Arena scores typesafe/jev-router at +3.6% and $0.155 a task
Arena's October 7, 2026 thread scores typesafe/jev-router at +3.6% net and $0.155 a task. The same thread says similar performance to DeepSeek V4.1 Flash (Max) costs 38% more, with median latency 1.7 times higher, on more than 4,700 agent sessions. The model page FAQ says the router has no price of its own. We did not call it.
typesafe/jev-router is OpenRouter’s chat model id for Jev Router. The model page, opened October 8, 2026, gives the display name TypeSafe: Jev Router and a created time of 2026-09-25T19:12:40.495Z. The FAQ on that page says the release date is September 25, 2026. The call is a chat completion. The text comes back from whichever model served the request. OpenRouter’s Decisions API for Jev is a separate door, already on the gateways story.
The router guide says Jev reads the conversation and judges the task type, the difficulty, and how much a stronger model would help. The router then picks the cheapest pool model that still meets that bar. The response model field names the model that served the request. The guide says the same id works for Chat Completions, Responses, and Messages, with or without streaming.
A plugin whose id is jev-router can pass models, allowed_models, and excluded_models. Each list takes up to 1,024 patterns. An unknown key returns 400. An include list that matches nothing in the pool is ignored, and the stage reports list_fallback as models_ignored. Exclusions still apply. If they remove every pool model, the request fails with 404. The lists can lower the tier, reported as list_tier_cap. On the hardest requests they can also drop the advisor, reported as max_fallback with the value deep. The header X-OpenRouter-Metadata: enabled asks for that stage on the response. We did not send a request.
The FAQ says Jev Router has no fixed price of its own. Each request is billed at the price of the model that serves it. The same FAQ says a request can go to a model with a context window of up to 1,000,000 tokens, and that the window on that request is the serving model’s. Inputs are audio, files such as PDFs, images, text, and video. The output is text.
The HTML meta description on the same page says “This model is free to use” and “1,000,000 token context window.” On October 8, GET https://openrouter.ai/api/v1/models listed this id with prompt and completion prices of -1, a context length of 1,000,000, tokenizer Router, and text as the only output modality. The page payload for one endpoint listed prompt "0" and completion "0", is_free false, data_policy.training false, and has_chat_completions true. GET https://openrouter.ai/api/v1/models/typesafe/jev-router/endpoints returned an empty endpoints array.
On September 25 at 22:21 UTC, OpenRouter posted the id. The thread says that before each turn Jev reads the prompt and scores difficulty and precision, whether a bigger model or more effort would help, whether a cheaper model is enough, and whether the task changed. Jev reads the conversation text only. Attachments are not sent to Jev. The post says the call is under zero data retention, and that requests with zdr: true work with this router.
The router keeps a working model for the rest of the session and can change effort without switching. It switches when the expected gain is larger than the cost, including the cached chat it would lose. If the Jev call times out or the output is invalid, the request fails instead of falling back to another router. The post says each response includes routing metadata with the reason for the choice.
The same thread says that on four agent benchmarks Jev Router solved 237 of 423 tasks and Auto Router solved 130 of 423. The post calls that 82% more tasks. The next post says that on five agent benchmarks Jev Router had a faster median time to first token than every other router tested, and it prints no times. The model page we opened does not reprint 237 or 130. TypeSafe quoted the announcement at 22:29 UTC. Those September counts are also on the gateways story.
Theo replied in that thread that Jev categorizes, and that the prompt does not tell it the agent’s tools or the size of the codebase. On September 26 he posted that he spent $1,000 benchmarking Jev Router. He wrote that DeepSWE was roughly the same as GPT-6 Astra on low, slightly more expensive, and almost 5 times longer. The post includes a chart. We did not read a task table off that image.
On September 28 Christoph Görn logged sixteen of his own prompts through the same model id. The answered_by field names openai/gpt-6-luna five times, deepseek/deepseek-v4.1-flash three times, and moonshotai/kimi-k3 once. Seven of the sixteen calls errored. Nine successful calls averaged 2.99 seconds, from 1.26 to 4.41. Sizing cost about $0.0032. The outcomes he counted were tiny 3, everyday 1, large 1, self 4, and error 7. The local Kev-4B arm of that write-up is in the Kev story.
On October 7 at 22:29 UTC, Arena posted an Agent Arena run of Jev Router. The thread says the tests span more than 4,700 real-world agent sessions. It says the router does not improve on the current Pareto frontier. For similar performance to DeepSeek V4.1 Flash (Max), the solution costs 38% more, and its median model request latency is 1.7 times higher. The router mostly picks models the post calls Pareto-efficient. The one it picks most is DeepSeek V4.1 Flash. GPT-6.1 Sol and GPT-6 Luna are also frequent. Steerability is +10%, against Claude Opus 5.5 (High) at +10.48.
The next post says DeepSeek V4.1 Flash took 28.1% of model calls. It says the router preferred OpenAI models over Anthropic’s. GPT-6 Astra was second at 23.5%, then Sol at 22.8% and Luna at 11.1%. Those four shares sum to 85.5%. The thread does not name the remaining calls.
A later post says that if Jev Router were on the Agent Arena leaderboard it would post a +3.6% net improvement at a median cost of $0.155 per task. It would rank below DeepSeek V4.1 Flash. On that comparison the router is ahead on Steerability, +10.0% against -0.2%, and on Praise vs Complaint, +6.3% against +2.1%. The post says DeepSeek is ahead on Confirmed Success, +8.3% against +2.2%, on Bash Recovery, +7.8% against -0.7%, and on Tool Hallucination, +0.4% against +0.2%.
Median latency in the next post is 6.18 seconds against 3.64 seconds for DeepSeek V4.1 Flash (Max). P90 is 30.59 seconds against 14.26 seconds. 6.18 / 3.64 is 1.70 to two decimals, which matches the 1.7 times line. The thread does not print the dollar figure behind the 38%. The last post says 4 of the router’s 5 most-used models sit on the Pareto frontier, and that calling DeepSeek V4.1 Flash directly gets similar performance at lower cost and latency. We did not rerun the sessions. The share image is the chart on the first post.
This site's reading
Editorial notes evaluating claims against primary sources, contextualizing findings alongside related implementations, and defining technical terms.
Verify
Primary post for the October measurement is @arena, October 7, 2026, 22:29 UTC, status 2107961555363213482. It says the tests span more than 4,700 real-world agent sessions. Similar performance to DeepSeek V4.1 Flash (Max) costs 38% more, and median model request latency is 1.7 times higher. Steerability is +10% against Claude Opus 5.5 (High) at +10.48. The share image is the chart on that post. We did not read extra counts off the pictures.
Reply 2107961558542557538 prints the mix: DeepSeek V4.1 Flash 28.1% of model calls, GPT-6 Astra 23.5%, Sol 22.8%, Luna 11.1%. Those four shares sum to 85.5%. The thread does not name the rest. Reply 2107961560794866099 prints +3.6% net improvement and a median $0.155 per task, and the pairs against DeepSeek: Steerability +10.0% against -0.2%, Praise vs Complaint +6.3% against +2.1%, Confirmed Success +8.3% against +2.2%, Bash Recovery +7.8% against -0.7%, Tool Hallucination +0.4% against +0.2%. The post says DeepSeek is ahead on the last three. Reply 2107961562858492338 prints median latency 6.18 seconds against 3.64 seconds, and P90 30.59 seconds against 14.26 seconds. 6.18 / 3.64 is 1.70 to two decimals. Reply 2107961564821405897 says 4 of 5 most-used models sit on the Pareto frontier. We did not rerun the sessions.
The model page https://openrouter.ai/typesafe/jev-router, opened October 8, 2026, names the id typesafe/jev-router. The page payload has created_at 2026-09-25T19:12:40.495Z. The FAQ says the release date is September 25, 2026, that the router has no fixed price of its own, that context can be up to 1,000,000 tokens and is the serving model's window, and that inputs are audio, files such as PDFs, images, text, and video, with text output. The HTML meta description says "This model is free to use" and "1,000,000 token context window."
GET https://openrouter.ai/api/v1/models on October 8 listed prompt and completion as "-1", context_length 1000000, tokenizer Router, and output modality text. The page payload for one endpoint listed prompt "0" and completion "0", is_free false, data_policy.training false, and has_chat_completions true. GET https://openrouter.ai/api/v1/models/typesafe/jev-router/endpoints returned data.endpoints as an empty array. We did not call the router.
The guide https://openrouter.ai/docs/guides/routing/routers/jev-router is the plugin and metadata page. Include lists that match nothing are ignored, with list_fallback "models_ignored". Exclusions that empty the pool return 404. Unknown plugin keys return 400. Each list takes up to 1,024 patterns. We did not send the metadata header.
OpenRouter's September 25 thread starts at status 2103610898690855161, 22:21 UTC. Status 2103610914503315515 prints 237 against 130 of 423. Status 2103610929493778525 says the median time to first token was faster than every other router tested and prints no times. Status 2103610965132779966 says Jev reads conversation text only, attachments are not sent to Jev, and the call is under zero data retention. Status 2103610988432126195 says a timeout or invalid Jev output fails the request. The model page does not reprint 237 or 130. Those September posts are also on the gateways story.
Compare
Arena's October 7 pairs are one Agent Arena run. DeepSeek V4.1 Flash (Max) is the model the router picked most, at 28.1% of calls. The router's printed net is +3.6% at a median $0.155 a task. The thread does not print DeepSeek's net score or DeepSeek's dollar cost. The 38% line and the 6.18 second median are the same comparison named two ways. 6.18 / 3.64 matches the 1.7 times line. Opus 5.5 (High) at +10.48 sits beside that steerability pair. It is outside the DeepSeek columns.
OpenRouter's September 25 count, 237 of 423 against Auto Router's 130 of 423, is four agent benchmarks and a different denominator. Theo's September 26 DeepSWE note is a $1,000 run against GPT-6 Astra on low, with no task count in the text. Görn's September 28 log is 16 prompts on the same model id, 7 errors, and a different clock. None of those is the Arena table.
Terms
- typesafe/jev-router
- OpenRouter's chat model id for Jev Router. The guide says Jev scores the conversation and the router picks a model and a reasoning effort. The response is text from the model that served the request.
- list_fallback
- A field on the jev-router pipeline stage. The guide says the value models_ignored means the include list matched no pool model, so the router used the whole pool. excluded_models still applies.
- Agent Arena net
- Arena's October 7 figure for typesafe/jev-router, +3.6% net improvement at a median $0.155 per task, on more than 4,700 agent sessions. The post says that would rank below DeepSeek V4.1 Flash.
Sources
- Arena, Agent Arena evaluation
- Arena, model mix
- Arena, leaderboard line
- Arena, latency
- Arena, summary
- OpenRouter, introducing typesafe/jev-router
- OpenRouter, 237 of 423 against Auto Router's 130
- OpenRouter, median time to first token
- OpenRouter, what Jev scores
- OpenRouter, text only, attachments, ZDR
- OpenRouter, session model and effort
- OpenRouter, timeout fails the request
- TypeSafe, quoting the Jev Router post
- Theo, Jev does not see the codebase
- Theo, DeepSWE timing on Jev Router
- Christoph Görn, who answered jev-router
- Görn, router log
- OpenRouter, Jev Router model page
- OpenRouter, Jev Router guide
- OpenRouter models list
- OpenRouter, jev-router endpoints