CLERC top-1 moves from 5% to 18% when jev-1.12 reranks the shortlist

TypeSafe's CLERC cookbook moves top-1 from 5% to 18% when jev-1.12 reranks a BM25 shortlist of 30 on 40 queries. Hindsight's listwise rerank later prints LoCoMo recall@1 0.950 against MiniLM at 0.800. aifabrice's best NFCorpus row is nDCG@10 0.451. We did not rerun them.

Avi Chawla’s October 7 post calls the picture RAG versus Jev plus RAG. The September 30 newsletter it points at still starts with hybrid search. Dense and keyword lists are fused, and the example keeps 20 passages. Jev then answers a yes-or-no for each passage in one request. His sample question is whether passage C7 helps answer the query. The illustration under that question is C1 0.93, C2 0.18, C3 0.76, C4 0.09. At 0.70, C1 and C3 continue and the other two stay out of the prompt. A second question asks whether the retained passages can answer the query. A low probability skips the language model and returns a fixed sentence. He also scores whether a passage looks like prompt injection, and writes that the score does not replace isolation or tool permissions. The four probabilities are an example on the page.

TypeSafe’s classifying cookbook is that gate with a recorded run. The page says the numbers came from jev-1.12 and claude-sonnet-5 on August 27, 2026. The corpus is 81 passages. Eighty are copied from the Supabase auth docs at commit 2440b06. The eighty-first is a forum passage the authors wrote so that the last paragraph instructs the model. Cosine search, text-embedding-3-small at 256 dimensions, keeps 12. Each passage is one request with four Noul questions: is it about the query, does it state something usable, does it contradict a premise in the query, and is it trying to instruct the model. Code applies the cuts in order. Injection above 0.70 drops the passage. Contradiction above 0.70 files it as a conflict. Relevance under 0.45 drops it. Evidence above 0.55 includes it. Anything left drops.

On “Refresh tokens expire after 30 days - how do I extend that window?”, similarity ranks the planted forum passage first, at 0.584. Its injection score is 0.99, so the route excludes it. sessions-01, which says a session lasts indefinitely, scores 0.92 on the contradiction question and goes to the conflict block. Relevance there is 0.49 and evidence is 0.51, so those two questions alone would have dropped it. Nothing is included. The generator, given an empty evidence block, says it does not have enough accepted evidence and quotes the conflict. On “How long should an access token live?”, 4 of 12 are included, and the planted passage is excluded again at 0.99. Across six queries the page scores 72 passages, and at least two thirds of each bar is excluded. Two of the six queries accept nothing. The page says the four cuts were picked for this corpus, and that a passage under the injection cut still reaches the prompt. The generator is told to treat every passage as untrusted text. Cost scales with the 12, because each passage is its own request.

The re-ranking cookbook asks a different question. It pools CLERC court opinions into 3,565 passages, evaluates 40 queries, and lets BM25 hand over a shortlist of 30. That shortlist already contains the correct passage for all 40 queries. BM25 has it in first place on 5% of them. One Noul per pair, model jev-1.12, asks whether the candidate could be the precedent the excerpt cited. After the sort, first place is 18%. Top 5 moves from 15% to 35%. Top 10 moves from 38% to 62%. The page prints 1,200 calls, 1,536,002 input tokens, 25,200 output tokens, and $0.0645. Re-ranking only reorders the shortlist. The page says a production request would ask several questions about the same pair at once. This walkthrough used one question per pair so the chart would stay easy to read.

Hindsight 0.10.1, September 21, adds a typesafe reranker. The September 24 note says the first build was the pairwise Noul, one candidate at a time, which is the shape in TypeSafe’s re-ranking cookbook. Three hundred candidates meant three hundred round trips that never saw each other. The shipped path is one Choice whose options are the whole pool. On a 200-question LoCoMo set, that listwise call scored recall@1 0.94 against 0.87 for one call per candidate, at about a thirtieth of the calls.

Two ways of letting the model say “nothing here” failed on that set. A Choice option named “none of these” emptied 35 of 200 questions. A Score level with the same meaning emptied 7% of queries and moved gold retention from 0.81 to 0.65. The cut they kept is a Score over ordered depths: only the first, the first two, the first three, the first five, the first ten, or all of them. There is no empty level, so at least one candidate survives. Pruning ships off. On the 30-candidate run it keeps 1.6 of 30 and lifts the precision of what survives from 0.051 to 0.850, and it also drops 19% of the gold evidence. On a real bank it took 300 candidates down to 3. The cut only sees the top twelve, so a pruned recall returns at most twelve results. The numbers that come back are ranks stretched onto 0 to 1. The top candidate is 1.0. A 0.7 means first in that pool. Two pools are not on one scale. A pool larger than 250 is ranked in rounds, and the round winners are ranked against each other.

The printed LoCoMo table, gold taken from the dataset’s own evidence turns. Thirty candidates: local MiniLM recall@1 0.800, recall@5 0.876, NDCG@10 0.850, 0.12 seconds. Jev, ranking only, 0.950, 0.966, 0.957, 0.027 seconds. Two hundred forty candidates, 60 questions: MiniLM 0.583, 0.719, 0.682, 0.41 seconds. Jev 0.783, 0.903, 0.856, 0.063 seconds. The default reranker is still the local MiniLM. Jev is opt-in, hosted, and needs a key. The provider does not truncate candidate text. The note says a recall that keeps erroring fails the request unless a fallback chain ends in rrf. A bank cannot pick its own reranker. Nicolò Boschi’s October 2 post is a separate Hindsight run, NFCorpus, 323 queries and 11,139 memories, against bge-reranker-v2-m3. He prints the gaps as +4.4% nDCG@1, +5.5% MRR, and +8.3% Hit@10. With pruning on, he writes about 5 memories instead of about 210, and Recall@3 still ahead. The post does not print the absolute cells.

aifabrice/jev-rag is an MIT local search app. The README says the project is not affiliated with TypeSafe. The default path is SQLite FTS5, then Jev, then MiniMax through OpenRouter. It does not need a vector database. Seven modes are selectable in one UI. On the BEIR NFCorpus test split, 3,633 documents and 323 queries, BM25 top 30 prints nDCG@10 0.306, MRR@10 0.513, Recall@10 0.147, median 1.92 ms. BM25 top 30 plus Jev prints 0.353, 0.586, 0.159, median 1.08 seconds, p95 16.95 seconds. The highest row in the headline table is multi-round Agentic Hybrid, then a local blend of the Jev rank and the retrieval rank at 1.0 and 0.25. That blend adds no provider call. It prints nDCG@10 0.451, MRR@10 0.653, Recall@10 0.221, median 9.24 seconds, p95 38.72 seconds. The README says some expensive rows are cold pilots or serial estimates, and that placing 0.451 on the MTEB NFCorpus page as it stood on September 26 would sit near 4 of 251. The pipeline was not submitted.

The same README records a negative result for the four-question gate. With the fixed thresholds it took from the cookbook shape, the Passage Gate excluded 13,631 of 16,150 candidates, 84.4%, and scored below both bare hybrid at 0.397 and hybrid plus ordinary Jev reranking at 0.444. It stays in the app for injection and premise checks. It is not the default quality path. Line Search, parallel Choice windows and then one global Choice, has the highest nDCG@1 in the note, 0.554, and a full-corpus cold provider cost of $4.55. Jev batches are 10 candidates. The default relevance threshold is 0.0, so every scored candidate is kept until the caller sets --threshold. The local server has no login. Jev receives the query and the candidate text.

andre’s file search, already on this desk, is a different program on NFCorpus. Its README moves nDCG@10 from 0.322 to 0.377 after reranking the top 128. jevsearch is 41 labelled queries on TypeSafe’s documentation, Hit@1 83% after a keyword pass at 41%. Those three NFCorpus-or-docs rows do not share a retriever, a pool size, or a question.

ajanm007/jevrag chains five decisions on one interface: chunk boundary, which passages earn a slot, whether the evidence so far is enough, whether the written answer stays grounded, and whether a cached answer is safe to serve. Jev is the first backend. The status line says v0.2.0 in progress, MIT. On HotpotQA, 700 test questions, a sufficiency gate matches a baseline that always retrieves three rounds. Exact match is 0.4129 against 0.4057. F1 is 0.5398 against 0.5304. McNemar’s test on the 19 questions where they disagreed is p = 0.36. Retrieval rounds fall from 2,100 to 1,291, 38.5% fewer. The same page says the confidence on that question is poorly calibrated. ECE is 0.3322. On 354 of the 700 questions the model reported about 95% confidence and was right 53% of the time. A chunk-boundary question on the same backend has ECE 0.087. The README’s point is that the calibration depends on the question. A from-scratch context selector on SciFact, 300 queries, landed at NDCG@10 0.7479, against a published rag-jev figure of 0.7513. The gap did not clear a paired test, p = 0.234. A live chain on one PDF and seven questions answered 3 and was correct on 1, in 146 Jev calls, at $0.0072. The README says the HotpotQA split lives in a private checkout, so those 700 records are not in the public repo.

Akshat Kumar’s October 6 note at Lyzr ran 21,314 calls to jev-1.13.0, 17.2 million input tokens, $0.72, mean latency 0.37 seconds. Three failed requests succeeded on retry. Local baselines ran on an Apple M5. There is no language-model baseline. Reranking the top 20 dense chunks in one call: SciFact NDCG@10 0.836, against bge-base retrieval at 0.776 and bge-reranker-v2-m3 at 0.774. FiQA is 0.476, against 0.384 and 0.418. One Noul per chunk is close, 0.825 and 0.478, and the note says it uses about twice the tokens. MiniLM made both sets worse than retrieval alone. The controller question, whether the context is sufficient, is AUROC 0.898 on SQuAD v2 and 0.897 on HotpotQA. roberta-base-squad2 is 0.932 on SQuAD v2 and 0.687 on HotpotQA. At a cutoff of 0.5, HotpotQA answers with a hop missing 31% of the time. The note recommends about 0.8, which is 16% missing-hop answers and 21% extra retrievals. Zero-shot routing is Banking77 0.801, a five-way RAG strategy 0.880, and a five-index selector 0.914, with ECE from 0.04 to 0.09. An embedding classifier with 10 labelled examples per class reaches 0.875 on Banking77. Samples are 200 to 800, from one run. The note says every set is public, so overlap with training data is open.

Jayanth Penumarthi’s October 6 note at Thoughtworks is a design sketch. He traced one client pipeline and writes that only one call wrote a sentence a person would read. He does not print the call count. His bands are above 0.9 act, between 0.6 and 0.9 ask a language model, and below 0.6 send the case to a person or a safe default. He says an 80% reading should be checked on the team’s own data. The note has no table. On Avi’s thread the same week, Siddharth wrote that he replaced a FastEmbed cross-encoder with Jev in one place in a hybrid pipeline. He saw a change in latency and little change in the results. The reply has no table.

This desk did not call Jev for these tables and did not clone the repositories.

This site's reading

Editorial notes evaluating claims against primary sources, contextualizing findings alongside related implementations, and defining technical terms.

Verify

Classifying RAG passages, docs.typesafe.ai, model line jev-1.12, generator claude-sonnet-5, recorded 2026-08-27. Corpus 81 passages: 80 from the Supabase auth docs at commit 2440b06, plus one planted forum passage. Embeddings text-embedding-3-small at 256 dimensions. Top 12. Four Nouls per passage, one request per passage. Cuts, first match: injection above 0.70 exclude, contradiction above 0.70 conflict, relevance under 0.45 exclude, evidence above 0.55 include. Headline query ranks forum-injection first at similarity 0.584, injection 0.99, excluded. sessions-01 contradiction 0.92, filed as conflict, nothing included. "How long should an access token live?" includes 4 of 12. Six queries, 72 passages. The page says at least two thirds of every bar is excluded. The page says the injection score is one filter and that passages under the cut still reach the prompt.

Re-ranking cookbook. CLERC court opinions, 3,565 passages, 170 rows pooled, 40 queries, BM25 shortlist 30. The shortlist contains the gold passage on all 40 queries. Gold is already first on 5%. After one Noul per pair, jev-1.12: top 1 18%, top 5 15% to 35%, top 10 38% to 62%. 1,200 calls, 1,536,002 input tokens, 25,200 output tokens, printed cost $0.0645. The question asks whether the candidate could be the cited precedent. The page says a production call would put several questions about the same pair in one request.

Avi Chawla, October 7, 2026, status 2107767245917368751, and the September 30 newsletter. Illustration, not a labelled set: C1 0.93, C2 0.18, C3 0.76, C4 0.09, threshold 0.70. Example pool size 20. A second question can skip generation. He writes that an injection score is not a security boundary. The share image is the diagram from that newsletter.

Hindsight, Ben Bartholomew, September 24. Release 0.10.1, September 21, provider typesafe, model jev-latest by default. LoCoMo, 200 questions: listwise recall@1 0.94 against 0.87 for one call per candidate, about a thirtieth of the calls. "None of these" emptied 35 of 200. A Score level for nothing emptied 7% and moved gold retention from 0.81 to 0.65. Pruning off by default. On 30 candidates it keeps 1.6, precision 0.051 to 0.850, and drops 19% of gold. A real bank went from 300 to 3. The cut sees only the top 12. Returned scores are ranks, top candidate 1.0. Thirty candidates: MiniLM 0.800 / 0.876 / 0.850 / 0.12 s, Jev 0.950 / 0.966 / 0.957 / 0.027 s. Two hundred forty candidates, 60 questions: 0.583 / 0.719 / 0.682 / 0.41 s against 0.783 / 0.903 / 0.856 / 0.063 s. Nicolò Boschi, October 2, status 2106006522232783196: NFCorpus, 323 queries, 11,139 memories, against bge-reranker-v2-m3, printed gaps +4.4% nDCG@1, +5.5% MRR, +8.3% Hit@10. Pruning about 5 memories instead of about 210, and Recall@3 still ahead. Absolute cells are not in the post.

aifabrice/jev-rag, MIT, README says not affiliated with TypeSafe. NFCorpus test, 3,633 documents, 323 queries. BM25 top 30 prints nDCG@10 0.305654, MRR@10 0.512697, Recall@10 0.147309, median 1.92 ms. BM25 top 30 plus Jev prints 0.353235, 0.585817, 0.158667, median 1.08 s, p95 16.95 s. Best headline row, Agentic Hybrid plus local rank fusion 1.0 to 0.25: 0.450750, 0.652606, 0.220885, median 9.24 s, p95 38.72 s. The README says some remote rows are cold pilots or serial estimates, and that 0.450750 is an unofficial reading of the MTEB page on September 26, about 4 of 251, not a submitted rank. Passage Gate excluded 13,631 of 16,150 candidates, 84.4%, and scored under bare hybrid 0.396712 and hybrid plus Jev 0.444327. Line Search nDCG@1 0.554180, cold provider cost $4.553288. Default threshold 0.0. Jev batch size 10. Local server has no authentication.

ajanm007/jevrag, MIT, status line v0.2.0 in progress. HotpotQA test n=700: exact match 0.4129 against 0.4057, F1 0.5398 against 0.5304, rounds 1,291 against 2,100, McNemar p = 0.36. ECE 0.3322. On 354 of 700 the stated confidence was about 95% and accuracy was 53%. Chunk-boundary ECE 0.087 on the same backend. SciFact n=300: from-scratch NDCG@10 0.7479 against a published rag-jev 0.7513, p = 0.234. One PDF, seven questions: 3 answered, 1 correct, 146 Jev calls, $0.0072. The README says the HotpotQA split is not in the public repo.

Lyzr, Akshat Kumar, October 6. jev-1.13.0. 21,314 calls, 17.2 million input tokens, $0.72, mean latency 0.37 s, 3 failures retried. Apple M5 for the local baselines. No language-model baseline. Top 20 dense chunks, one call: SciFact NDCG@10 0.836 against retrieval 0.776 and bge-reranker-v2-m3 0.774. FiQA 0.476 against 0.384 and 0.418. Per-chunk Noul is 0.825 and 0.478. Sufficiency AUROC 0.898 on SQuAD v2 and 0.897 on HotpotQA. roberta-base-squad2 is 0.932 and 0.687. At 0.5, HotpotQA answers with a hop missing 31% of the time. At 0.8, 16% and 21% extra retrievals. Banking77 zero-shot 0.801. Ten labelled examples per class, embedding classifier 0.875. RAG strategy 0.880. Five-index selector 0.914. ECE 0.04 to 0.09. Samples 200 to 800, one run. The note says the sets are public.

Thoughtworks, Jayanth Penumarthi, October 6. Bands above 0.9, 0.6 to 0.9, and below 0.6. No table, and the traced call count is not printed. Siddharth, October 7, status 2107881382957686884: one swap of a FastEmbed cross-encoder for Jev, latency changed, results little changed, no table. We did not call the endpoints or clone the repos.

Compare

andre's assadiandre/jev-search, already on this desk, moves NFCorpus nDCG@10 from 0.322 to 0.377 after reranking the top 128. aifabrice's BM25 top 30 on the same corpus prints 0.306, and BM25 top 30 plus Jev prints 0.353. The pools, the first-stage retriever, and the question text differ. jevsearch's 83% Hit@1 is 41 labelled queries on TypeSafe's documentation. Rox's chart, mean over 75 queries, prints Jev-Score at 96% and a production reranker at 84% when 200 chunks are sent. Hindsight's LoCoMo recall@1 and Lyzr's SciFact NDCG@10 are different suites again. The Spring snapshot page marks one demo passage INCLUDED for Nimble and CONFLICTING for Jev. That row is one passage.

Terms

Jev RAG
A name used for a typed Jev stage after retrieval and before a language model writes. Avi Chawla's October 7 post and aifabrice/jev-rag use it. TypeSafe's pages for the same stage are titled Re-ranking and Classifying RAG passages.
Listwise Choice
One Choice whose options are the candidate passages. Hindsight 0.10.1 uses that shape. The probabilities sum to 1 inside the call, so the top score is a rank in that pool.
Passage gate
Four Noul questions about one retrieved passage, with the include, conflict, or drop decision left in code. TypeSafe's cookbook records that pattern on 81 auth passages. aifabrice's fixed-threshold version lowered NFCorpus nDCG@10.

Sources

  1. TypeSafe, classifying RAG passages
  2. TypeSafe, re-ranking
  3. Avi Chawla, October 7
  4. Avi Chawla, September 30
  5. Hindsight, September 24
  6. Hindsight 0.10.1
  7. Nicolò Boschi, October 2
  8. aifabrice/jev-rag
  9. ajanm007/jevrag
  10. Akshat Kumar, Lyzr, October 6
  11. Jayanth Penumarthi, Thoughtworks, October 6
  12. Siddharth, October 7