Updated

Published

Jev scored 384 morning headlines in 25 seconds; Opus 5 finished four

Elvis ran a morning Google News wire through Jev and Claude Opus 5 at the same time. Jev tagged 384 headlines for 15 brand desks in 24.9 seconds at $0.19. Opus got through four headlines and cost $0.77 before the run was stopped.

Elvis, who builds Newsjack, posted a clip of Jev reading 384 morning headlines in 24.9 seconds and sending them to 15 brand desks for $0.19. Claude Opus 5, on the same feed at the same time, finished 4 of the 384 and cost $0.77. He put the per-headline gap at about 390 times cheaper. The code lives in demos/news-desk-dealer.

The wire is a Google News pull (six topic feeds plus Techmeme). Set A stamps each headline with a desk, a story type, and six 0-4 scores that draw a radar: magnitude, velocity, novelty, window, heat, risk. Those weights become a newsworthiness mark out of 10, using Newsjack’s public newsworthiness-check skill. Set B scores each story for each company on standing (none through direct) and journalist shape, then tiers it. A story can land on several desks or none. The 15 companies are public names, one per industry, used as sample desks.

When Jev finishes, the Opus run is stopped so it does not keep spending tokens. The README estimates a full Opus pass over 398 headlines at about $25 if that abort is off. The checked-in UI uses a mock engine unless you set live keys; the September 18 numbers are from the live path. In that run Jev took about 190 ms for set A and about 320 ms for set B (90 questions, about 10k input tokens), around 1,200 judgments per second with eight workers. Opus 5 took 4 to 6 seconds for set A and longer for set B.

The post does not say how often Jev’s desk or newsworthiness marks match a human editor. The published comparison is how far each model got, and what it cost, before the slower column was cut off.

This site's reading

Editorial notes evaluating claims against primary sources, contextualizing findings alongside related implementations, and defining technical terms.

Verify

The 384 count, 24.9 seconds, $0.19, and the Opus 5 comparison (4 of 384 headlines, $0.77) come from Elvis's September 18 post. The ~390x cheaper per-headline figure is his: Jev's finished 384 against Opus's unfinished four. The demo README records a live run the same day: Jev set A about 190 ms per call, set B (90 questions, about 10k input tokens) about 320 ms, roughly 1,200 judgments per second with eight workers; Opus 5 about 4 to 6 seconds for set A. The UI ships with a mock engine on by default. A full Opus pass over 398 headlines is estimated at about $25 if it is not stopped. We did not rerun the live engines. The post does not publish agreement with human editors.

Compare

Ryan Vogel classified 1,500 personal emails and posted no comparator. Every's 777 writing judgments were timed against Claude Fable 5.1 on a known answer key. This demo races a named frontier model on the same morning wire and stops the slower column when Jev finishes. Throughput and spend are the published metrics. Accuracy against a newsroom is not.

Terms

Set A
Per-headline questions: a desk, a story type, and six 0-4 scores (magnitude, velocity, novelty, window, heat, risk) that feed a newsworthiness mark out of 10.
Set B
Per-company questions on standing and journalist shape, then a tier. A story can land on several of the 15 illustrative desks, or none.
Abort-on-finish
The Opus 5 column is stopped when Jev completes the wire, so the slower model is not billed for the remaining headlines.

Sources

  1. Elvis, Jev vs Opus 5 on 384 headlines
  2. News Desk Dealer README (Newsjack demo)
  3. Newsjack repository