Published
TechCrunch covers Jev, and quotes a Vercel engineer who swapped Luna for it
Tim Fernholz's TechCrunch story is the first large press write-up of Jev. It reports that TypeSafe briefly could not serve the API under demand, quotes Pranit Sharma at Vercel on a 5 to 18 times speedup versus gpt-5.6-luna, and points at Nikhil Mudholkar's email test against Gemini.
Tim Fernholz published a TechCrunch story on Jev on September 18. TypeSafe quoted the post.
Fernholz writes that TypeSafe briefly lost the ability to serve the API because demand was so high. He interviews Diogo Almeida on why a model that does not write text might be cheaper to run inside software, and on the synthetic-data bet behind RLCD. Almeida told him that making all of the training data was “one of the best bets I’ve ever made in my life,” better than the launch and better than RLHF.
Pranit Sharma, a software engineer at Vercel, posted on September 16 that the company had used gpt-5.6-luna to run a classifier that reviews commands for safety in fx auto mode. After swapping in Jev, he wrote, the classifier was about 5 to 18 times faster and more accurate. The post does not publish the accuracy percentages, the sample, or a cost figure.
Nikhil Mudholkar, CTO of Bryo AI, tested Jev against Gemini on business email routing. TechCrunch reports Gemini as slightly more accurate and 10 to 20 times more expensive, and quotes Mudholkar on confidence scores: “it is the only one that hands back a real probability which makes it ideal for automating workflows!!” The full counts from that thread are in a separate story.
Fernholz also quotes Armin Ronacher, CTO of Earendil, on what a probability is for. If the model returns 50%, Ronacher says, treat it as a coin toss. If it returns 95%, the application can act. He also names model routing: predicting which generator a prompt needs is useful, and using an LLM for that job is expensive.
Almeida, asked whether TypeSafe is a frontier lab, told Fernholz that the main product of frontier labs is fear or hype, and that he would like TypeSafe’s main product to be intelligence.
This site's reading
Editorial notes evaluating claims against primary sources, contextualizing findings alongside related implementations, and defining technical terms.
Verify
Fernholz's article is dated September 18, 2026. Sharma's September 16 post is the source of the Vercel figure: the fx auto mode safety classifier was "about 5-18x faster and more accurate" than gpt-5.6-luna. That post does not publish accuracy percentages or the sample size. Mudholkar's numbers are in a separate story on this desk; TechCrunch compresses them to "slightly more accurate" and "10 to 20 times more expensive." Almeida's quotes on synthetic data and on TypeSafe not being a "fear or hype" lab are from Fernholz's interviews. The claim that TypeSafe briefly lost the ability to serve the API is Fernholz's reporting. We did not rerun the Vercel classifier.
Compare
This desk already covered the September 15 launch. The new material is a named-company speed claim (Sharma at Vercel) and a named-company accuracy claim (Mudholkar at Bryo). Armin Ronacher, quoted in the article, describes confidence as a halt the application owns, and cheap routing as a reason to put Jev in front of a slower model. Those two uses match LangChain's refusal gate and Hono's semantic router more closely than they match TypeSafe's internal 193.6x workflow table.
Terms
- fx auto mode
- A Vercel safety classifier Sharma says previously ran on gpt-5.6-luna. The September 16 post is the source for the 5 to 18 times figure.
- Calibrated decision
- TypeSafe's phrase, used in the TechCrunch piece, for a typed answer with a probability rather than generated text.