Updated

Published

Every ran 777 judgments in under a second and still wanted a better accuracy check

Mike Taylor at Every put 37 documents through 21 AI-writing questions at once. Jev returned 777 answers in under 0.7 seconds for about a quarter of a cent. A later defect test with Dan Shipper missed one of seven planted problems that Claude Fable 5.1 caught.

Mike Taylor, head of evals at Every, published a launch-day write-up of Jev on September 15, 2026, and updated it on September 18.

He fed Jev the text of his 27 Every articles plus 10 AI-styled counterparts, then asked the same 21 questions of every document. The questions came from Every’s list of AI writing tells: repeating an idea without adding evidence, forcing a symmetrical both-sides argument, overexplaining a straightforward point.

Jev returned 777 answers in under 0.7 seconds. Taylor estimated the cost at a quarter of a cent. He wrote that the model correctly flagged pieces of his that leaned more on AI, and that he would want a more thorough accuracy check before putting it into production. He called that useful as an early warning. The alternative, he wrote, is not checking at all.

He ran 11 experiments in total, covering context lookup, work checks, and routing decisions. Across that suite Jev produced 1,709 judgments for an estimated total of less than a cent. GPT-6 Astra, running in Codex, wrote the queries from TypeSafe’s docs.

Every CEO Dan Shipper ran a tighter test: 12 synthetic passages, six clean and six with planted problems, scored on four writing checks. Jev’s median time was 0.35 seconds per passage. Claude Fable 5.1 at high effort took 8.83 seconds, about 25 times slower and about 580 times more expensive on Taylor’s estimate. Fable caught all seven intended defects. Jev caught six, and missed the same one in all three runs: “a shared appointment calendar that parents and staff teach together,” which Fable treated as an unexplained action.

Taylor’s advice was to pick a judgment you already make repeatedly, compare accuracy against the model you use today, and see whether acting on the answers makes the work better.

This site's reading

Editorial notes evaluating claims against primary sources, contextualizing findings alongside related implementations, and defining technical terms.

Verify

The 777 figure is 37 documents times 21 questions, as stated in Taylor's Every essay published September 15, 2026 and updated September 18. Latency ("under 0.7 seconds") and cost ("a quarter of a cent", then 1,709 judgments for under a cent across 11 experiments) are author estimates. We did not rerun the lab. The Fable comparison (median 0.35 seconds versus 8.83 seconds at high effort, about 25 times faster and about 580 times cheaper, six of seven planted defects) is from Every CEO Dan Shipper's 12-passage test in the same essay. The missed defect is quoted there: a calendar that "parents and staff teach together."

Compare

TypeSafe's launch evals compare Jev to large chat models on internal workflows and report the high-end 193.6x / 444.6x figures. Every's tests are smaller and independent: a writing-pattern sweep over Taylor's own archive, plus a 12-passage defect check against Fable 5.1. Ryan Vogel's 1,500-email run is a personal demonstration without a comparator or published error rate. Mudholkar's later 1,565-email table is the named-company accuracy baseline against Gemini; Every remains the write-up that times Jev and scores it against a named frontier model on planted defects. The sample is still small.

Terms

Judgment
One typed answer to one question about one document. 37 documents times 21 questions produced the 777 count.
Planted defect
A problem inserted into synthetic text so the tester knows which passages should be flagged.
Early warning
Taylor's description of using Jev as a first pass that is cheap enough to run even when a later human or larger model still reviews the result.

Sources

  1. Mini-Vibe Check (Every / Mike Taylor)
  2. Introducing System One Models & Jev