Published
Updated
On WANDS, more Jev questions move the mean to 0.610 and leave the median at 0.561
Doug Turnbull's WANDS table prints mean NDCG@10 of 0.610 for a bag of Jev questions and 0.574 for one question. Both medians print 0.561, and the CSV stores one shared float. The score adds ten times the summed Nouls to BM25. This desk checked the CSV and did not rerun the searches.
Nathan LeClaire posted on October 9, 2026, at 06:59 UTC, that Jev is a reranker, “and not just for text.” The post is that sentence. It has no chart, no corpus, and no question text. His bio names TypeSafe.
Yechan Do asked, three minutes later, whether multimodal input was coming. Opened the same day, the models page says Jev 1.13 takes text only: a string, a JSON object, or an array of text values. Images, audio, and video are not accepted. jev-latest and jev-preview both point at jev-1.13.0. On October 4 LeClaire had written that multimodal support would be good, and that he was not promising it.
The reply with a measurement is Ravi’s. He linked Doug Turnbull’s October 8 note, Bag of Decisions Reranker, and wrote that a better question than “is it relevant?” worked better. Turnbull has GPT-5 write yes/no questions for a single shopping query, sends them to Jev as Nouls, and adds the probabilities to a BM25 score.
The blog’s middle column is named jev_reranker. The CSV and the YAML call that row bag_of_decisions_direct. It is one configured question, “Does this candidate product satisfy the search query?”, with written criteria for true and false. The state includes the query, the title, and the description. The third column, bag_of_decisions on the blog and bag_of_decisions_example in the repo, is the GPT-5 list. Its state is the title and the description. The query sits in the generated questions. strategy.py refuses criteria on that LLM generator. Both configs name the model jev/jev-latest.
| Dataset | Blog name | Repo strategy | Mean | Median |
|---|---|---|---|---|
| WANDS | bm25 | bm25 | 0.540755 | 0.474620 |
| WANDS | jev_reranker | bag_of_decisions_direct | 0.574262 | 0.560945 |
| WANDS | bag_of_decisions | bag_of_decisions_example | 0.609814 | 0.560945 |
| ESCI | bm25 | bm25 | 0.289473 | 0.170709 |
| ESCI | jev_reranker | bag_of_decisions_direct | 0.313621 | 0.204850 |
| ESCI | bag_of_decisions | bag_of_decisions_example | 0.360079 | 0.341417 |
The two WANDS medians are one stored float, 0.5609447680982702. The means differ. The blog prints both medians as 0.561 and the bag mean as 0.610. Rounding the CSV mean 0.6098137673305334 to three decimals is 0.610. Rounding the BM25 mean 0.5407553553417934 to four decimals is 0.5408. The blog prints 0.5407.
The note at research/decisions/ecom.md says an LLM generated it. Its ESCI bag median is the string 0.34141707368146015. The CSV median is 0.34141715214740553. The strings share 0.341417 and then diverge. The note’s median and its mean, 0.36007937368146015, both end in the digits 7368146015. The blog’s 0.341 still matches the CSV at three decimals.
The published commands and scripts/run_ecom_decisions_variants.sh use 480 WANDS queries, 1,000 ESCI queries, and seed 42. The CSV stores the commit 8734fde3b23ce17eafdf60b2f55dfa6bced47e43 and does not store the query count. The note’s commands pass 8 workers. The script’s own default is 16. tool_calls_mean is 1.0 on the BM25 rows too, so that column is not a count of Jev calls. The CSV has no price.
The score is the retrieval score plus 10 times the sum of the Noul probabilities. decision_reranker.py adds a probability only when it is finite and between 0 and 1, so a missing answer drops out of the sum. All questions for one product go in one request. The strategy rescores the top 100 hits and writes those scores back. Every other product keeps its retrieval score. The metric then reads the top 10.
cheat_at_search keeps ranks through 10. The discount is 1/log2(2^rank), which equals 1/rank. Gain is (2^grade - 1) / rank. The ideal puts the dataset’s maximum grade in every slot from 1 to 10, and that ideal is the same for every query. A query with few relevant products cannot reach 1.
The two Jev rows share a first stage: title weight 9.4, description weight 4, k1 1.2, b 0.75. The printed BM25 row does not. configs/ecom_base/bm25.yml sets the title boost to 9.3 and the description boost to 4.1. The note calls 9.4 and 4 the baseline. On WANDS, the matched comparison is direct against the bag: mean 0.574262 to 0.609814, median unchanged. The gap against the printed BM25 row also changes the field weights.
The repo does not publish the per-query scores, so this desk could not see which queries moved. There is also no column that ranks by the summed probabilities alone. The blog says that sum could be used on its own. The six stored rows all add it to retrieval.
The other replies in the thread do not name a corpus. Oscar Romero wrote that their production recall was 5 to 10 points under a qwen3-reranker, and that Jev helps drop off-topic hits without much fine-tuning. Hev’s September table, a different test, prints mean nDCG@10 of 0.501 for a Jev batch reranker and 0.478 for Qwen3-Reranker-0.6B, across SciFact, NFCorpus, and FiQA. adaxon wrote that it was not able to beat a tuned light gbm, with no table. 3THER_ wrote that a guardrail can ride in the same request. TypeSafe’s classifying cookbook already asks an injection Noul beside relevance. That result stays on Jev RAG.
Qdrant posted the same day. On 100 NFCorpus queries, a Jev Score lift over BGE-small is +0.0665 at depth 30 and +0.0741 at depth 50. Their Iterative method uses 10 requests and reaches +0.0746 and +0.0789. On 240 WANDS queries their Jev rerank prints nDCG@10 0.769 against hybrid search at 0.683, and the number of product classes in the top 10 falls from 3.17 to 2.41. That 0.769 is a different pipeline from Turnbull’s 0.610.
This desk read the CSV, the configs, the scorer, and the NDCG function. It did not rerun the searches.
This site's reading
Editorial notes evaluating claims against primary sources, contextualizing findings alongside related implementations, and defining technical terms.
Verify
Primary post is @dotpem, Nathan LeClaire, October 9, 2026, 06:59 UTC, status 2108452170639220915. The text is one sentence: Jev is a reranker, and not just for text. No image, no chart, no corpus, no question text. His bio names TypeSafe. On October 4, status 2106840933279191068, he wrote that multimodal support would be good and that he was not promising it.
Yechan Do, 07:02 UTC, status 2108453088814313665, asked if multimodal was coming. Opened October 9, 2026, https://docs.typesafe.ai/models says Jev 1.13 input is text only: a string, a JSON object, or an array of text values. No image, audio, or video. jev-latest and jev-preview both point at jev-1.13.0. Price on that page is $0.042 per million input tokens, output free.
Oscar Romero, 08:37 UTC, status 2108476814460916046, wrote that their tests were 5 to 10 points lower on recall than a qwen3-reranker in production, and that Jev helps drop off-topic results without much fine-tuning. No dataset. adaxon, status 2108476812309492097, wrote that it did not beat a tuned light gbm. No table. 3THER_, 09:48 UTC, status 2108494737300738317, wrote that guardrailing can ride in the same request. Ravi, 08:55 UTC, status 2108481431311925400, linked Doug Turnbull's October 8 post.
The blog table prints WANDS means 0.5407, 0.574, 0.610 and medians 0.475, 0.561, 0.561. ESCI means 0.289, 0.314, 0.360 and medians 0.171, 0.205, 0.341. Column names are bm25, jev_reranker, bag_of_decisions.
The CSV at commit 8734fde3b23ce17eafdf60b2f55dfa6bced47e43 stores WANDS bm25 0.5407553553417934 / 0.47461951858375095, direct 0.5742621821794697 / 0.5609447680982702, example 0.6098137673305334 / 0.5609447680982702. ESCI bm25 0.28947299049683545 / 0.17070857607370277, direct 0.3136213056689957 / 0.2048502912884433, example 0.36007937368146015 / 0.34141715214740553. The two WANDS Jev medians are the same float. research/decisions/ecom.md, which says it was generated by an LLM, prints the ESCI example median as 0.34141707368146015. The CSV median is 0.34141715214740553. The note's median and its mean, 0.36007937368146015, both end in the digits 7368146015.
Rounding 0.5407553553417934 to four decimals is 0.5408. The blog prints 0.5407. The other blog cells match the CSV at the precision the blog shows. The CSV does not store the query count. The note and scripts/run_ecom_decisions_variants.sh default to 480 WANDS queries, 1,000 ESCI queries, and seed 42. The note's commands pass 8 workers. The script defaults to 16. tool_calls_mean is 1.0 on every row, including BM25.
configs/ecom_base/bm25.yml sets title_boost 9.3, description_boost 4.1, k1 1.2, b 0.75. Both decision configs set fields title^9.4 and description^4, and leave k1 and b unset. strategy.py then uses 1.2 and 0.75. decision_model is jev/jev-latest, decision_weight 10, k 100. The direct question includes the query in the state and sets true and false criteria. The bag state is title and description. strategy.py raises if the LLM generator is given criteria. It rescores the top 100, writes those scores back, and leaves every other document at its retrieval score. decision_reranker.py adds weight times the sum of finite Noul values in [0, 1]. A missing answer is left out of the sum. One request carries every question for one candidate.
cheat_at_search.eval.grade_results keeps rank <= 10. The discount is 1/log2(2^rank), which is 1/rank. Gain is (2^grade - 1) / rank. The ideal puts the dataset's maximum grade in all ten slots.
Qdrant, October 8, Andrei Cristea and Chadha Sridi. 100 NFCorpus queries, lift over BGE-small: Jev Score +0.0665 at top 30 and +0.0741 at top 50. Iterative, 10 requests, +0.0746 and +0.0789. On 240 WANDS queries their Jev rerank prints nDCG@10 0.769 against hybrid 0.683, and product classes in the top 10 fall from 3.17 to 2.41. Hev's September 17 table, updated September 24, prints a three-corpus mean nDCG@10 of 0.501 for Jev batch and 0.478 for Qwen3-Reranker-0.6B, on SciFact, NFCorpus, and FiQA, BM25 top 30. This desk did not rerun any of those searches and did not call Jev for this page.
Compare
The six CSV rows are one runner at one commit. Direct against the bag is the matched pair: same title^9.4 and description^4 retrieval, same top 100, same weight of 10. The printed BM25 row uses title 9.3 and description 4.1. Qdrant's WANDS 0.769 is hybrid search on 240 queries. Hev's 0.501 is a mean of SciFact, NFCorpus, and FiQA. CLERC top-1 5% to 18%, LoCoMo recall@1 0.950, and Rox's 96% at 200 chunks are other suites. Oscar's 5 to 10 point recall gap names no corpus.
Terms
- Bag of decisions
- GPT-5 writes yes/no questions for one query. Jev answers each as a Noul. The reranker adds 10 times the sum of those probabilities to the candidate's retrieval score.
- Direct question
- One configured Noul, with true and false criteria, asked of each of the top 100 products. The blog labels this column jev_reranker. The repo calls it bag_of_decisions_direct.
- NDCG@10 here
- cheat-at-search keeps ranks through 10. Gain is (2^grade - 1) / rank. The ideal is a perfect top 10 at the dataset's maximum grade, the same ideal for every query.
- decision_weight
- The multiplier on the summed Noul probabilities. Both published configs set it to 10.
Sources
- Nathan LeClaire, October 9
- Yechan Do, October 9
- Oscar Romero, October 9
- adaxon, October 9
- Ravi, October 9
- 3THER_, October 9
- Nathan LeClaire, October 4
- Doug Turnbull, October 8
- Experiment note
- Results CSV
- Comparison script
- Direct question config
- Bag config
- BM25 config
- Strategy
- Reranker
- cheat-at-search NDCG
- TypeSafe models
- Qdrant, October 8
- Hev, Jev as a reranker