Published
Updated
Decision Index 0.2 prints Jev at 51.67, and a later post names another first place
Opened October 8, 2026, the Decision Index space still headers v0.3, 115 models, updated 2026-10-07, and Perplexity Decider v1.1 is still first at 62.8. Opened October 7, 2026, the Decision Index space was edition 0.3, 112 models, and its full-score table ranked Perplexity Decider v1.1 first at 62.8. Jev 1.13.0 is a reference at 60.1. Rank 2 is Fastino GLiDE no-thinking at 60.2, and rank 3 is Torchcast Decision 27B at 59.9. apolinario/decision-index reproduces Decision Index 0.2, dated September 24: 40 benchmarks, chance-corrected, five equal areas. A test in the kit locks Jev at 51.67, AutoJev-27B at 50.94, Hopper at 30.01, and Verdict at 1.82. Posts the next day say Drex leads at 51.73. That 51.73 is not in the test we read. The board file generated September 28 is edition 0.2.1 and lists Jev 1.13.0 at 57.91 among 71 names, with no Drex. Nace's page prints Drex 1.5 at 58.28. Maincode's September 30 card prints a self-run of the 0.2.1 kit at 59.26 for Matilda, against this board's Jev cell of 57.91. The card says that run is being submitted, so 59.26 is not a row in the September 28 file. Nace's September 30 post also calls Drex 1.5 first on JevBench. The chart on the next post prints 201 of 231 for Jev and 199 of 231 for Drex 1.5. Liquid's September 29 chart is a separate reproduction at 58.9 against 57.9. Fastino's October 1 blog prints GLiDE at 64.81 on a self-run of the 0.2.1 scorer, against this board's Jev cell of 57.91. AutoTrust's October 3 card prints GEV-26B-Decide at 62.48 on another self-run. Neither figure is in the September 28 file. StartLux's October 3 README prints 63.88 for StartLux-Decision-27B on a self-run of the same scorer, higher than that Jev cell on 31 of 38 benchmarks, and says the run is not on the board. vLLM Semantic Router's Decision-2.0-Vega-27B card, opened October 5, prints 56.5 on its own 0.2.1 reproduction. The card does not print a Jev row, and 56.5 is not in the September 28 file.
apolinario/decision-index is a kit for rebuilding Decision Index, a benchmark for models that take a state and typed questions and return a probability for every option. The README on September 30, 2026 says the live board is Decision Index 0.2.1, dated September 27, and that the project is not affiliated with TypeSafe. Edition 0.2, dated September 24, is the earlier board. Code in the repository is MIT. The suite text comes from other datasets under their own terms, and the README says not to train on it.
Edition 0.2.1 rescores the same suite files. It counts 38 benchmarks. Arts and Human Taste is fixed at 10%, and the other four areas share the rest by the square root of how many benchmarks they hold. Gold benchmarks weigh 1.2 inside an area. RouterBench and SGD stay on the board and leave the index. The kit’s count for that edition is 120,340 requests, 119,898 that score after 442 exclusions, plus the same 30,419 for the seven benchmarks added in 0.2.
The index is the mean of five areas, times 100. Each area is the mean of its benchmarks. A benchmark is chance-corrected: 0 is random guessing and 100 is perfect, and a request that was not answered counts as wrong. MMLU left the index. Seven benchmarks were added, including MMLU-Pro, BBH, When2Call, RAGTruth, HoVer, New Yorker caption matching, and PhishNChips. The README’s count of the frozen rows is 121,057 requests after an ACOS subset, 120,615 that score after 442 exclusions, plus 30,419 requests for the new seven. Scores within 0.25 points share a rank.
tests/test_index02.py checks that math against published per-benchmark results. The four numbers named in the README are Jev 51.67, AutoJev-27B 50.94, Hopper 30.01, and Verdict 1.82. The README on September 30 says a rescore covers 64 entrants on the 0.2 board and 67 on the 0.2.1 board.
On September 30, 2026 we ran tests/test_index02.py and tests/test_index021.py from that day’s main. 13 tests passed in the first file and 19 in the second. The 0.2 locks are unchanged. The 0.2.1 test locks Jev at 57.89 in its fixture, with Rune 26B-A4B v3 at 57.44 and AutoJev-27B at 56.40, among 13 names. Drex is not one of those names.
On September 25, Florian S wrote that the Decision Index and JevBench measure different things. He called the index accuracy on 40 public benchmarks, and JevBench his own items plus a sealed set that also prices speed and cost. He wrote that Jev is first on raw intelligence on both, with JevBench Intelligence at 53.1 against 49.4, and that decider-4b v2 leads JevBench because it is faster and cheaper. He did not print 51.67.
The same afternoon, two posts said a model called Drex had taken first place: 51.73 against Jev 1.13.0 at 51.67, ahead on 23 of the 40 benchmarks, 70 tokens a decision against 367, and about 25 decisions a second on one H100. Movez also described a flap-or-glide game he marked as simulated, 20 seeds, 19 wins and one draw, about 136 ms against about 204 ms. The 51.67 matches the kit’s test. The 51.73 does not appear in that test, and we did not find Drex named in the README section we read.
The board we opened the same day does not name Drex either. data/index.json, generated 2026-09-28T00:39:36+00:00, is Decision Index 0.2.1. It lists 70 models plus Jev 1.13.0 at 57.91, then Surogate Rune 26B-A4B v3 at 57.44, Decider chat Gemma-4-31B at 57.33, and AutoJev-27B at 56.40. The 0.2 file from September 27 still has Jev at 51.67, and it has Decider chat Gemma-4-31B at 51.93, 0.26 above Jev. Feeding every per-benchmark skill in both files through the kit’s index02.aggregate reproduced the printed index within 0.01 on all 65 and 71 names. The largest gap was 0.0051.
Nace’s launch post prints the 51.73, and says Drex wins 22 of the 39 benchmarks scored for it. Its table lists Rune at 47.23 and Decider chat at 46.08, which match the September 27 file, and it does not list Decider chat Gemma-4-31B. The product page prints Drex 1.5 at 58.28 on Decision Index 0.2.1, first of 71, with Jev at 57.91 and a lead on 21 of 38. It says that row was scored on September 28 with the official 0.2.1 kit, and that the other cells are the public board’s scores from that day. 58.28 is not in the board file.
We sent one request through the kit’s http engine to https://api.typesafe.ai/v1/systemone, model jev-1.13.0. Warm-up was off, so this was the only request. The state was “The color is red.” The choice asked which color is named, red or blue. The noul asked whether the state names a color. The runner recorded status ok and named jev-1.13.0. The choice was red, with probabilities 1.0 and 0.0 and confidence 1.0. The noul was 0.94. Usage was 324 input tokens and 48 output tokens. Wall time was 646 ms. The runner finished at 2026-09-30T02:53:12Z.
We did not run the frozen suite. The repository does not ship the rows. Edition 0.2.1 scores 120,340 requests.
Maincode’s card on September 30 prints a self-run of this kit: Matilda 59.26 against the board’s Jev cell of 57.91. The card says the model is being submitted, so that 59.26 is not a row in the September 28 file. The area split, including Knowledge and Reasoning at 43.52 against Jev at 51.40, is on the Matilda page.
Nace’s September 30 post calls Drex 1.5 first on JevBench as well as on this index. The chart on the next post prints Jev 1.13.0 at 201 of 231 and Drex 1.5 at 199 of 231, with a hard-tier tie at 82 of 111. The product page prints a 128K context window and a size under 10 billion parameters. Those rows are on the Drex 1.5 page. Liquid’s chart the day before is a different reproduction, 58.9 against 57.9. We did not re-download the September 28 file, and we did not search it for the string d1. On October 1 a search of the JevBench page for “Drex 1” returned no match.
On October 1 Fastino’s blog printed a self-run of this same 0.2.1 scorer: GLiDE at 64.81 against the board’s Jev cell of 57.91. The blog says submissions for that edition are paused, so 64.81 is not a board row. On October 3 AutoTrust’s card printed another self-run, GEV-26B-Decide at an adaptive 62.48, with raw 70.66 and breadth 62.00. The card says the adaptive pass on Knowledge and Reasoning misses the kit’s latency cap. Both numbers are written up on their own pages, GLiDE and GEV-26B-Decide. We did not search the September 28 file for either name. The public kit still treats a gap of 0.25 as a tie, and both of these gaps are larger than that.
On October 3 StartLux’s README printed a further self-run: StartLux-Decision-27B at 63.88, the 35B-A3B at 61.55, and the 9B at 58.63, against the same 57.91. The 27B is ahead of that Jev cell on 31 of 38 benchmarks. The README says the runs are not on the board, and that 14 of the 38 benchmarks contributed their public train splits to training. Knowledge and Reasoning on the 27B is 44.3 against Jev at 51.4. The write-up is StartLux-Decision. We did not search the September 28 file for that name.
On October 5 the Decision-2.0-Vega-27B card printed another reproduction of the 0.2.1 kit, 56.5. The card does not print a Jev row. 56.5 is not in the September 28 file. The rest of that family is on the Decision 2.0 page. We did not search the September 28 file for Vega.
Opened October 8, 2026, the same space still says Decision Index 0.3. The header says 115 models, Jev jev-1.13.0, reproductions on one NVIDIA RTX PRO 6000, updated 2026-10-07. The lede says 114 open reproductions. The head-to-head panel still lists Perplexity Decider v1.1 at 62.8, Fastino GLiDE no-thinking at 60.2, and Torchcast Decision 27B at 59.9, with Jev as a reference at 60.1. The full-results table prints =2 on the GLiDE row, the Jev row, and the Torchcast row. We did not rerun the index.
Opened October 7, 2026, the same space headers edition v0.3, 112 models, Jev jev-1.13.0, reproductions on one NVIDIA RTX PRO 6000, updated 2026-10-06. The full-score summary, labelled all five areas, lists Jev at 60.1 as a reference. Rank 1 is Perplexity Decider v1.1 (27B) at 62.8. Rank 2 is Fastino GLiDE no-thinking (28B) at 60.2, and the button says open weights are coming soon. Rank 3 is Torchcast Decision 27B at 59.9. The v1.1 row then prints 45.8, 62.3, 64.7, 80.6, and 51.4, in the area order Knowledge, Language, Retrieval, Tools, Arts. It also prints an empty cell, 28B, 100.0%, and 104 ms. We did not keep the column titles for those three cells. Language lists Jev at 58.3 and v1.1 at 62.3. Retrieval lists Jev at 61.4 and v1.1 at 64.7. We did not copy a Jev reference for Knowledge, Tools, or Arts. The formula on the page is 0.20 times public, plus 0.50 times same-skill private, plus 0.30 times new domains. The public part is 100 times the weighted mean of the area scores, with GSM8K rebuilt and ForecastBench retired.
The v1.1 card prints its own table at 61.56 against Jev at 57.9. That table is a different reading from the space’s 62.8. Drex DLM prints 52.31 on a card labelled Decision Index 0.2, against the kit’s Jev cell of 51.67. A search of the 0.3 space for Drex returned no row. A search for “Decider v1 (27B)” returned no row. We did not rerun edition 0.3.
This site's reading
Editorial notes evaluating claims against primary sources, contextualizing findings alongside related implementations, and defining technical terms.
Verify
The README on September 30, 2026 says the live board is Decision Index 0.2.1, dated September 27, and that the project is not affiliated with TypeSafe. Edition 0.2, dated September 24, averages 40 benchmarks in five equal-weight areas. Scores are chance-corrected, 0 for random guessing and 100 for a perfect score, and unanswered requests count as wrong.
Scores within 0.25 index points rank as a tie. tests/test_index02.py checks the index math against published per-benchmark results for Jev at 51.67, AutoJev-27B at 50.94, Hopper at 30.01, and Verdict at 1.82. The README says a rescore covers 64 entrants on the 0.2 board and 67 on the 0.2.1 board.
Edition 0.2 adds seven benchmarks and drops MMLU from the index. The 0.2 row file is described as 121,057 requests after an ACOS subset cut, 120,615 scoreable after 442 exclusions, plus 30,419 requests for the new benchmarks.
The September 25 posts say Drex is 51.73 against Jev at 51.67, first on the index, and ahead on 23 of 40 benchmarks. Movez labels his own flap-or-glide bench simulated, 20 seeds, 19 wins and 1 draw, about 136 ms against about 204 ms. Those Drex figures are not in the parity test we read.
On September 30, 2026 we ran tests/test_index02.py and tests/test_index021.py from the kit's main. 13 tests passed in the first file and 19 in the second. The 0.2 locks are unchanged. The 0.2.1 test locks Jev at 57.89 in its fixture, with Rune 26B-A4B v3 at 57.44 and AutoJev-27B at 56.40, among 13 names. Drex is not one of those names.
The same day we opened the board at huggingface.co/spaces/multimodalart/jev-decision-index. data/index.json was generated 2026-09-28T00:39:36+00:00 and is Decision Index 0.2.1. It lists 70 models plus Jev 1.13.0 at 57.91. Next are Surogate Rune 26B-A4B v3 at 57.44, Decider chat Gemma-4-31B at 57.33, and AutoJev-27B at 56.40. The file does not contain the string Drex.
data/index-v2.json, generated 2026-09-27T01:35:16+00:00, is the 0.2 board. It lists 64 models plus Jev at 51.67. Decider chat Gemma-4-31B is 51.93. AutoJev-27B is 50.94. Drex does not appear in that file either. 51.93 is 0.26 above 51.67, outside the 0.25 tie band.
We passed every per-benchmark skill in both files through decision_index.scoring.index02.aggregate. All 65 names on the 0.2 file and all 71 names on the 0.2.1 file came back within 0.01 of the printed index. The largest gap was 0.0051.
www.nace.ai/blog/introducing-drex prints Drex at 51.73 on edition 0.2, ahead of Jev at 51.67 and AutoJev-27B at 50.94, and says Drex wins 22 of the 39 benchmarks scored for it. The table there lists Surogate Rune at 47.23 and Decider chat at 46.08. Those two cells match the September 27 file. Decider chat Gemma-4-31B at 51.93 is not in that table.
www.nace.ai/drex prints Drex 1.5 at 58.28 on Decision Index 0.2.1, first of 71, and Jev 1.13.0 at 57.91, ahead on 21 of 38. The page says Drex 1.5 was scored on 2026-09-28 with the official 0.2.1 kit, and that the other cells are the public board's scores read that day. 58.28 is not in the board file.
We sent one request through the kit's http engine to https://api.typesafe.ai/v1/systemone, model jev-1.13.0. Warm-up was off, so this was the only request. The state was "The color is red." The choice asked which color is named, with options red and blue. The noul asked whether the state names a color. The runner recorded status ok and named jev-1.13.0. The choice was red, probabilities 1.0 and 0.0, confidence 1.0. The noul was 0.94. Usage was 324 input tokens and 48 output tokens. Wall time was 646 ms. The runner finished at 2026-09-30T02:53:12Z.
We did not run the frozen suite. The repository does not ship the rows. Edition 0.2.1 scores 120,340 requests.
Maincode's card, posted September 30, prints a self-run of the 0.2.1 kit. Matilda is 59.26, raw 68.89, breadth 58.06. The card's Jev (hosted) cells are 57.91, 68.09, and 57.08. The card says the model is being submitted. 59.26 is not a name in the September 28 board file. Area scores for that run are on the Matilda page.
Nace's September 30 post calls Drex 1.5 first on JevBench as well as on this index. The chart on the next post prints 201/231 for Jev 1.13.0 and 199/231 for Drex 1.5, and a hard-tier tie at 82/111. The product page prints a context window of 128K and a size under 10B. Liquid's September 29 chart is a different reproduction, 58.9 against 57.9. We did not re-download the September 28 board file on October 1, and we did not search that file for the string d1. A search of the October 1 JevBench page for "Drex 1" returned no match.
Fastino's blog, opened for the GLiDE page, prints a self-run of the 0.2.1 scorer at 64.81 for GLiDE against the published Jev cell of 57.91. The blog says 0.2.1 submissions are paused and that this score is not on the public leaderboard. AutoTrust's GEV-26B-Decide card prints another self-run, adaptive balanced skill 62.48, raw 70.66, breadth 62.00. The card says adaptive mode misses the kit's latency cap on Knowledge and Reasoning. We did not re-download the September 28 file for either name.
StartLux's README, opened October 4, prints 63.88 for StartLux-Decision-27B, 61.55 for the 35B-A3B, and 58.63 for the 9B, scored with the board's kit. The Jev cell it cites is 57.91. The README says the runs are not on the board. The October 3 post says the same, and that the 27B is ahead on 31 of 38 benchmarks. We did not re-download the September 28 file for the name StartLux.
On October 7, 2026 the Decision Index space headers v0.3, 112 models, Jev jev-1.13.0, reproductions on one NVIDIA RTX PRO 6000, updated 2026-10-06. The full-score table ranks Perplexity Decider v1.1 at 62.8. Jev 1.13.0 is a reference at 60.1. Rank 2 is Fastino GLiDE no-thinking at 60.2. Rank 3 is Torchcast Decision 27B at 59.9. A search for Drex returned no row. We did not rerun the board.
Compare
JevBench v1.4.2 is 534 public decisions plus 308 sealed, scored as a harmonic mean of intelligence, calibration, speed, and cost. On that board Jev's Intelligence is 53.1 and decider-4b v2 leads the composite.
Florian's September 25 reply says the Decision Index is breadth across 40 public benchmarks and that Jev is first on raw intelligence on both. The 51.67 here is a chance-corrected average. It is not the 63.29 harmonic mean, and it is not the 75.4 geometric mean from v1.2.8.
The September 28 board file is edition 0.2.1. Jev 1.13.0 is 57.91 there, a chance-corrected average of 38 benchmarks.
Maincode's September 30 card prints 59.26 for Matilda on a self-run of this kit, against that 57.91. The card says the run is being submitted. Drex's 58.28, on Nace's page, is a different claim from both.
Nace's September 30 chart prints 199 of 231 for Drex 1.5 against 201 of 231 for Jev on 231 public items, and a hard-tier tie at 82 of 111. Liquid's chart prints 58.9 against 57.9 on its own reproduction. The October 1 JevBench page text has no "Drex 1".
GLiDE at 64.81 and GEV-26B-Decide at 62.48 are October self-runs of the 0.2.1 scorer. Matilda's 59.26, Drex 1.5's 58.28, and Liquid's 58.9 are earlier ones. The September 28 file still lists Jev at 57.91. The kit's tie band is 0.25. None of those self-runs is a row we found in that file.
StartLux-Decision-27B at 63.88 is another October self-run, written up on its own page. The README says 14 of the 38 benchmarks had their public train splits in the training data. Knowledge and Reasoning on that run is 44.3 against Jev at 51.4.
Edition 0.3, opened October 7, 2026, is a different board. The full-score table ranks Perplexity Decider v1.1 at 62.8 and lists Jev 1.13.0 as a reference at 60.1. The v1.1 card's own table prints 61.56 against Jev at 57.9. Drex DLM's card prints 52.31 on edition 0.2 against the kit's Jev cell of 51.67. A search of the 0.3 space for Drex returned no row.
Terms
- Decision Index
- A chance-corrected average of 40 public benchmarks in five equal areas. Zero is random guessing and 100 is perfect. Unanswered requests count as wrong. Edition 0.2 is dated September 24, 2026. The kit's test locks Jev at 51.67.
- 0.25
- The tie band in index02.ranks. Two scores within 0.25 index points share a rank. AutoJev-27B at 50.94 is 0.73 below Jev at 51.67, outside that band. On the September 27 file, Decider chat Gemma-4-31B at 51.93 is 0.26 above Jev, also outside that band.
- 0.2.1
- The live board edition in the September 28 file. 38 benchmarks. Arts and Human Taste weighs 10%. The other four areas share 90% by the square root of their benchmark count. Gold benchmarks weigh 1.2 inside an area. The file lists Jev 1.13.0 at 57.91. The kit's test locks that cell at 57.89.
- 0.3
- Opened October 8, 2026, the space header says v0.3, 115 models, Jev jev-1.13.0, updated 2026-10-07. Perplexity Decider v1.1 is still first at 62.8. Jev is a reference at 60.1. Opened October 7, 2026, the header said 112 models, updated 2026-10-06.
Sources
- Hesamation, Drex and the index
- Movez, Drex against Jev
- Florian S, the two boards
- apolinario/decision-index
- Decision Index board
- Nace, Introducing Drex
- Nace, Drex
- Maincode, Matilda self-score
- Nace AI, Drex 1.5
- Liquid AI, d1 chart
- Fastino, Introducing GLiDE
- AutoTrust, GEV-26B-Decide
- StartLux, October 3
- StartLux-Decision README
- Decision-2.0-Vega-27B
- Perplexity, v1.1, October 6
- pplx-decider-v1.1-27b