Updated

Published

Jason Lemkin tried Jev on SaaStr Connect and stopped on a 32% false-admission rate

Jason Lemkin ran 600 judgments against Claude Sonnet. Jev cost $0.064 against about $5, roughly 90 times cheaper. A blind third-model referee scored Jev 70.5% and Sonnet 77.5% on 200 random pairs. Jev let through 47 of 148 negatives. He says that is too high for Connect.

Jason Lemkin posted on September 19 that he spent the day trying Jev for SaaStr Connect. He reports $0.064 for 600 judgments against about $5 on Claude Sonnet, using TypeSafe’s $0.042 per million input tokens and free output. On identical inputs that is about 90 times cheaper.

Agreement with Sonnet: 59.5% on 200 pairs he weighted toward borderline cases, 67% on 200 random pairs. A blind third-model referee on the random set scored Jev 70.5% and Sonnet 77.5%.

The misses were not symmetric. On 148 negatives, Jev admitted 47 (32%) and Sonnet 20 (13.5%). On 52 positives, Jev rejected 2 and Sonnet 15. Connect, as he describes it, cannot take a 32% rate of weak answers marked acceptable.

The post does not include the prompts, the pair list, or the referee’s name. There is no gist. We did not rerun the 600 judgments.

This site's reading

Editorial notes evaluating claims against primary sources, contextualizing findings alongside related implementations, and defining technical terms.

Verify

The September 19 post is the source. There is no repository. Lemkin reports $0.064 for 600 judgments against roughly $5 on Sonnet, input at $0.042 per million tokens and free output, about 90× cheaper on identical inputs. Agreement with Sonnet: 59.5% on 200 borderline-weighted pairs, 67% on 200 random pairs. Against a blind third-model referee on the random set: Jev 70.5% accurate, Sonnet 77.5%. Error split on the negatives/positives he posted: Jev 47 false admissions out of 148 negatives (32%) and 2 false rejections out of 52 positives; Sonnet 20 false admissions (13.5%) and 15 false rejections. He says the 32% false-admission rate is too high for SaaStr Connect. We did not see the prompts, the pairs, or the referee model. We did not rerun the set. TypeSafe's Master Customer Agreement section 2.3(f) forbids publishing benchmarks of the Services; this story reports the post as published.

Compare

LangChain's judge eval is the same job on frozen weather traces with a human oracle; Jev matched every binary label there and the sample is five traces. Bryo auto-routes the high-confidence band and leaves the rest to a person, which is the override pattern Lemkin did not keep. Kumar's reject filter is built to drop work, not to admit candidates. Poly's 90% precision gate is the same caution on Banking77. This post is the first Connect-style admission task we have with a named false-admission count.

Terms

False admission
A negative case the model treats as acceptable. Lemkin reports 47 of 148 for Jev (32%) and 20 for Sonnet (13.5%) on this SaaStr Connect sample.
Borderline-weighted pairs
Lemkin's 200-pair slice that over-represents close cases. He reports 59.5% agreement with Sonnet there, against 67% on 200 random pairs.

Sources

  1. Jason Lemkin, SaaStr Connect trial