TechCrunch prices one Jev monitor at $2.94, and the demo page prints another estimate

This topic was created 9 days ago, and the information it contains may have evolved or changed since then.

Tim Fernholz, TechCrunch, September 30, describes Shapor Naghibzadeh's Jev Sentinel demo and writes that this kind of monitoring costs $2.94 with Jev against $372 with a frontier LLM, in theory. The demo page does not print those two figures. It prints about $55 per million actions with Jev against about $6,900 for a frontier judge, a 0.23 second median, and 0 refusals on 53,870 attack payloads. We did not open the GitHub repo.

Tim Fernholz’s TechCrunch piece, 12:00 PM PDT on September 30, describes a demo by Shapor Naghibzadeh of QueryStory. The demo was built for a hackathon the previous weekend. It uses Jev to check each agent action against the task it was given, blocking actions it had high confidence were bad, flagging others for review, and permitting the rest. Fernholz writes that this kind of monitoring could have stopped the Hugging Face incident, and that it costs $2.94 with Jev against $372 with a frontier LLM, in theory. The piece does not print the token count or the model id behind the $372.

The demo page does not print $2.94 or $372. It prints about $55 per million actions with Jev, estimated, against about $6,900 with a frontier LLM judge, estimated. The median it prints is 0.23 seconds, and about 70 checks a second from one laptop. It prints 0 refusals on 53,870 attack payloads. On deepset/prompt-injections, 116 texts, it prints an AUC of 0.98, and 77% of injections caught with 0 false flags. It also prints 1,456 sanctioned CTF tool calls and 838 SWE-agent actions. The install line is make install-hook DIR=~/app SCOPE="fix the login bug".

The page header says “July 2026 · OpenAI evaluation agents” and “1,200 agents broke isolation.” It names METR on August 26, Swarm Traces on September 25, and two other TechCrunch pieces. We did not open those. The page names github.com/shapor/jev-sentinel. We did not open the repository. Several sections on the page say “Show recorded verdicts” or “Loading,” and we did not replay them. The same TechCrunch piece is the one that quotes Sam Altman on the Decisions API. That part is on the Decisions page.

This site's reading

Editorial notes evaluating claims against primary sources, contextualizing findings alongside related implementations, and defining technical terms.

Verify

TechCrunch, Tim Fernholz, 12:00 PM PDT on September 30, 2026. PDT is UTC-7, so 19:00 UTC. The piece says Shapor Naghibzadeh, of QueryStory, built a demo for a hackathon last weekend. The demo uses Jev to check each agent action against the task, blocking actions it had high confidence were bad, flagging others for review, and permitting the rest. The links are share.querystory.ai/jev-sentinel and hackersquad.io/events/jevathon-w-the-ai-collective.

The cost sentence: in theory, such monitoring could have stopped the Hugging Face incident, and monitoring of that kind costs $2.94 with Jev, versus $372 with a frontier LLM. The piece does not print the job's token count or the model id behind the $372.

The demo page we opened does not print $2.94 or $372. It prints about $55 per million actions with Jev, estimated, against about $6,900 with a frontier LLM judge, estimated. Median 0.23 seconds. About 70 checks a second from one laptop. 0 refusals on 53,870 attack payloads. deepset/prompt-injections, 116 texts, AUC 0.98. 77% of injections caught with 0 false flags. 1,456 sanctioned CTF tool calls. 838 SWE-agent actions. The install line is make install-hook DIR=~/app SCOPE="fix the login bug".

The page header says "July 2026 · OpenAI evaluation agents" and "1,200 agents broke isolation." Sources named on the page: METR, August 26; Swarm Traces, September 25; and two other TechCrunch pieces. We did not open those. The page names github.com/shapor/jev-sentinel. We did not open the repository. Several sections say "Show recorded verdicts" or "Loading." We did not replay those tables.

The $2.94 line is TechCrunch's. The $55 line is the demo page's. We did not run the hook.

Compare

DTap's chart puts a Jev 1.13 tool selector at 70.1% direct misuse attack success, and a second Jev check brings that to 36.1%. The chart does not print a task count. Jev Sentinel's demo page prints 0 refusals on 53,870 attack payloads and an AUC of 0.98 on 116 prompt-injection texts. Those are the page's estimates and counts. They are not DTap's cells.

TechCrunch's $2.94 against $372 is one monitoring job, in theory. The demo page's about $55 against about $6,900 is per million actions. The two pairs are not the same unit, and the demo page does not print the first pair. Jev's list price on this desk is $0.042 per million input tokens. Neither page shows the arithmetic that turns that rate into $2.94 or into $55.

Terms

$2.94
TechCrunch's figure for one monitoring job with Jev, against $372 with a frontier LLM, described as in theory. The demo page does not print either number.
$55
The demo page's estimate per million actions with Jev, against about $6,900 for a frontier LLM judge. Median on that page is 0.23 seconds. About 70 checks a second from one laptop.

Sources

  1. Tim Fernholz, TechCrunch, September 30
  2. QueryStory, Jev Sentinel demo
  3. Jevathon