You.com's risk sweeps cost $0.14 each, and the report means tie

This topic was created 8 days ago, and the information it contains may have evolved or changed since then.

Edward Irby's three escalated sweeps cost $0.43 in all, and a blind read of nine Jev reports and six Qwen reports ties at a mean of 4.25.

Edward Irby posted a long write-up on October 1, 2026. He works at You.com. The code is youdotcom-oss/risk-analysis-server, MIT, and it runs on Bun.

A profile is a topic, a list of places, and a few triggers. You.com search returns web hits and licensed Knowledge facts. Jev returns a Noul for “is this a threat?” and for ranking proposed queries, a Score on a three-step relevance rubric, and a Choice of low, medium, or critical.

Qwen3.8 27B, qwen/qwen3.8-27b on OpenRouter, proposes the queries and writes the Markdown briefing. Application code owns the threshold, the budgets, and the SQLite row. MCP serves the same tools over stdio and over HTTP, as a handle the client polls.

Caps that showed up on a bill

The deep dive stops at 5 proposal steps, 8 queries after Jev ranks them (RISK_MAX_QUERIES), 30 results scored, 15 passed to synthesis, 10 pages fetched, 12,000 characters a page, and 100,000 characters in total.

An earlier loop let the model run searches itself and spent 19 to 26 searches a sweep. Ranking first cut an escalated sweep to a flat 12. One synthesis path, before the character cap, grew to about 212,000 tokens.

Three sweeps, one invoice

All three escalated, with six Qwen calls each.

PNW data center buildout came back low, with 6 knowledge hits, 22 searches, 71 seconds, and 20,658 / 881 Jev tokens. US AI lab operations came back critical, 1 hit, 19 searches, 101 seconds, 18,489 / 797. Gulf AI infrastructure came back medium, 0 hits, 26 searches, 83 seconds, 19,980 / 902.

Qwen inference was $0.0619 for 18 calls, 66,741 in and 13,632 out, or $0.0206 a sweep. You.com was $0.37 for 67 searches and 29 page fetches, or $0.12 a sweep. Jev was $0.0025 for 59,127 in and 2,580 out, or $0.0008 a sweep, at the listed $0.042 per million input tokens with output free.

The measured total is $0.43, $0.14 a sweep. A follow-up under the query budget cut retrieval from 67 searches to 36, about $0.21, where the earlier retrieval line was $0.37. Contents is billed per URL: three tool calls arrived as 29 API calls. Irby writes that this section does not compare the briefings with a frontier model.

0.5 against 0.7

RISK_TRIAGE_THRESHOLD ships at 0.5. The same three profiles, minutes apart, on fresh keys:

At 0.5, Gulf was 0.56, escalated, critical, 12 searches, 104 seconds. PNW was 0.63, medium, 12 searches, 209 seconds. US AI lab was 0.83, medium, 12 searches, 100 seconds.

At 0.7, Gulf was 0.58 and stopped, recorded as low, 1 search, 3 seconds. PNW was 0.63 and stopped, low, 1 search, 3 seconds. US AI lab was 0.85 and still escalated, medium, 12 searches, 86 seconds.

The 0.7 arm used 14 searches against 36, about 40 percent of the spend. Gulf had been the critical briefing on the 0.5 arm, so the higher cutoff dropped it. Irby kept 0.5. He prices a wasted deep dive at about $0.14, and a miss as silence.

Blind scores

RISK_JUDGE selects jev or qwen for the four decisions. Same questions, same batches, same rubric, strict JSON. Parse failures: zero.

Jev completed 9 sweeps and escalated all 9. Triage nouls stayed within 0.05 of each other. It escalated PNW at 22:06 and 22:29, at 0.63 to 0.65.

Qwen attempted 11 and escalated 6. Gulf read 0.35, then 0.68, then 0.50. PNW read 0.35 to 0.40 and was skipped at 22:22, 22:30, and 22:42.

One reader scored the deep-dive reports blind, dates stripped, rubric frozen first, one news day. Factual grounding was 4.44 for Jev (n=9) and 4.83 for Qwen (n=6). Signal capture was 4.00 and 4.17. Actionability was 4.22 and 3.67. False-signal avoidance was 4.33 and 4.33.

The means are 4.25 and 4.25. On a total of 20, the floor and ceiling are 14/20 and 11/19. Perfect reports against weak reports are 3 and 0, against 0 and 1. Qwen’s single PNW deep dive scored 19 of 20. The reader called each skip defensible from the stub the gate had seen. A skipped sweep looks like a quiet day.

Jev’s judgments on a deep dive are about $0.0005, around 13,000 input tokens. The Qwen judge averages 9,000 input tokens and 22,000 to 41,000 output tokens. At about $4.43 per million output tokens, that is $0.10 to $0.18 a deep dive, about 250 times the Jev judgment line, and more than the $0.14 all-in figure for a Jev sweep.

A skip still bills 8 to 27 times a full Jev-judged sweep. Qwen deep dives took 7 to 12 minutes. Jev’s took about 2.

TypeSafe’s quote says the language-model judge missed 5 of 11 investigations and Jev missed none, and it puts the price at $0.02 a sweep.

The article’s escalation counts are 6 of 11 attempts and 9 of 9. The $0.0206 line is Qwen inference alone. The invoice for the sweep is $0.14.

This site's reading

Editorial notes evaluating claims against primary sources, contextualizing findings alongside related implementations, and defining technical terms.

Verify

Edward Irby, @edwardirby, You.com, October 1, 2026, 16:25 UTC, status 2105695667486355848. The article is edwardirby/article/2105695667486355848. We opened it in the browser. The repository is youdotcom-oss/risk-analysis-server. GitHub lists the license as MIT.

The invoice for three escalated sweeps is Qwen $0.0619, You.com $0.37, Jev $0.0025, total $0.43, or $0.14 per sweep. Jev is priced in the article at $0.042 per million input tokens, output free. The blind means are 4.44 against 4.83, 4.00 against 4.17, 4.22 against 3.67, and 4.33 against 4.33, averaging 4.25 and 4.25.

TypeSafe's quote, status 2105733977147617750, says the language model missed 5 of 11 and Jev missed none, at $0.02 per sweep. The article's counts are 6 of 11 escalations and 9 of 9.

The $0.0206 figure is the Qwen inference line. We did not run a sweep.

Compare

The $0.14 sweep is retrieval plus two models. HydroJEV's median compute, 1.215 seconds on 2,722 calls, is a different pipeline and a simulated alarm set. This article's wall clocks are 71 to 209 seconds on the invoice run.

Qwen's Gulf triage moved 0.35, 0.68, 0.50 on the same profile and the same day. The blind means are both 4.25. A Noul near 0.5 is also the band the jaggedness notes treat as least certain.

Terms

RISK_TRIAGE_THRESHOLD
The Noul cutoff that decides whether a sweep escalates. The article ships 0.5. At 0.7, Gulf scored 0.58 and the run stopped, after the 0.5 arm had marked that profile critical.
RISK_JUDGE
The switch, jev or qwen, behind triage, ranking, scoring, and severity. The Qwen arm answers the same questions in strict JSON. The article reports zero parse failures.

Sources

  1. Edward Irby, October 1
  2. The X article
  3. youdotcom-oss/risk-analysis-server
  4. TypeSafe AI, October 1