StartLux-Decision-27B prints 63.88 on its own Decision Index 0.2.1 run

StartLux posted StartLux-Decision on October 3, 2026. The README scores the 27B at 63.88 on Decision Index 0.2.1, above the September 28 board's Jev cell of 57.91 on 31 of 38 benchmarks, and says the run is not on the board. Code is Apache-2.0. Weights are CC BY-NC 4.0.

StartLux posted StartLux-Decision on October 3, 2026, at 13:41 UTC. The README scores StartLux-Decision-27B at 63.88 on Decision Index 0.2.1, StartLux-Decision-35B-A3B at 61.55, and StartLux-Decision-9B at 58.63. The highest entry it cites from the public board of September 28, 2026 is Jev 1.13 at 57.91. The 27B is higher than that Jev cell on 31 of the 38 benchmarks. The README says the runs were scored with the board’s own kit, and that they are not on the board.

The family is five dense sizes, 0.8B, 2B, 4B, 9B, and 27B, plus a 35B mixture of experts with about 3B parameters active per token. A request is a state plus questions: pick one option, yes or no, or a rating. The reply is a probability for each option, read from the option letters after one forward pass. The HTTP shape is TypeSafe’s POST /v1/systemone. Code in the repository is Apache-2.0. Weights on Hugging Face are CC BY-NC 4.0. Commercial use of the weights needs a separate license from StartLux Labs. The collection is startlux-models.

A note at the top of the README, dated 2026-10-03, says every size now reads images in the evidence and prompts up to 262,144 tokens. The MLX and GGUF backends read text only. The 35B fits one 80 GB GPU at about 71 GiB peak, and it is not in the GGUF table. Other calls of this shape are on decision models like Jev.

Training data includes the public train splits of 14 of the 38 benchmarks. The README marks those rows with a dagger and says the test items were filtered out. The index cells are chance-corrected, so 0 is random guessing and 100 is perfect. Starred benchmarks weigh 1.2. The other systems’ cells are taken from the public board.

Area27B35B-A3B9B4BJev 1.13
Decision Index 0.2.163.8861.5558.6352.7557.91
Knowledge & Reasoning44.342.438.032.151.4
Language Understanding74.573.671.363.962.0
Retrieval & Classification66.865.864.156.755.4
Tools & Automation82.276.273.272.375.1
Arts & Human Taste47.944.741.833.737.7

On Knowledge and Reasoning the 27B is behind Jev, 44.3 to 51.4. GPQA Diamond is 34.0 against Jev at 71.4. MMLU-Pro is 66.0 against 80.5. BBH is 71.3 against 89.7. HLE is 0.0 against 4.7. API-Bank is 83.8 against 88.0. Those five Jev leads are on rows the README does not mark with a dagger. SATA-Bench is 9.3 against Jev at 25.4, and against Decider chat 31B at 28.4.

The size table counts correct answers on 231 public items in the Intern-Decision bundle. StartLux-Decision-27B is 208. The 35B is 210. The 4B is 204. Jev 1.13 is 199. Intern-Decision-4B is 201. The average accuracy over the bundle’s seven suites is 91.82 for the 27B, 92.29 for the 35B, 90.02 for Intern-Decision-4B, and 88.74 for Jev. The README says the public count is not the official JevBench score. That score adds a sealed tier, speed, and cost, and only the maintainers measure it.

A later GGUF table, run through this repository’s package, prints 209 of 231 for the 27B original weights. The README says those original rows differ from the table above by one or two items of bf16 rounding.

Latency in the size table is one H200, three questions in one forward pass, fast kernels and CUDA graphs on. The 4B is 26.0 ms. The 0.8B is 12.2 ms. The 2B is 15.5 ms. The 9B is 35.7 ms. The 35B is 52.5 ms. The 27B is 102.3 ms. The same request through the TypeSafe API, mean of 100, is 64.0 ms of server time, with the network left out. The post’s 26 ms line is the 4B row. The 27B row in the same table is 102.3 ms. Without CUDA graphs the 4B takes 90.3 ms for that request. Intern-Decision’s own 4B measurement, on an RTX 4090, is 44.2 ms, and its end-to-end Jev figure is 109.7 ms.

The games table is one match they ran, plus numbers they attribute to other harness authors. Chess is 256 games at the rich level, no search, Jev called through its API in the same harness. StartLux-Decision-27B scores 149.5 points to Jev’s 106.5, which the README writes as 58.4%, +59 Elo, 95% interval +37 to +82. The record is 64 wins, 171 draws, and 21 losses. With the harness’s tactical hints, the README says the two are even, 48.6% and 46.1% over 256 games each. On 75 lines of clean NPC address detection the 27B has 1 wrong answer and Jev has 6. On the misheard-names variant the counts are 5 and 12. Dino Run, also run for Jev through the API, is 20 of 20 for both at the 300-obstacle cap.

The six clips on the README are described as real runs of the 27B. The computer-use clip is a mock shop in their harness. Mario, StarCraft, Doom, and JevBall are the same pattern: the model picks from a list, and code carries out the action. Those clips do not add a second score beyond the games table. We did not run the models. The Decision Index page keeps the September 28 board’s Jev cell at 57.91.

This site's reading

Editorial notes evaluating claims against primary sources, contextualizing findings alongside related implementations, and defining technical terms.

Verify

StartLux, @StartLuxAI, October 3, 2026, 13:41 UTC, status 2106379033722360053. The post prints 63.88, 31 of 38 benchmarks, an Intern-Decision average of 91.82% against Jev at 88.74% and Intern-Decision-4B at 90.02%, and 26 ms for a three-question request on the 4B. It says the Jev comparison uses the September 28, 2026 public leaderboard, and that the results are self-evaluated. The photo on that post is the share image. Its chart axis starts at 50. The caption says Jev's size is not published, and that StartLux-Decision is scored with the board's kit and is not listed on it.

The README at StartLuxLabs/StartLux-Decision was opened October 4. The same 63.88, 61.55, and 58.63 are in the Decision Index table, next to the board's Jev 57.91, Rune 57.44, Decider chat 31B 57.33, and AutoJev-27B 56.40. A note dated 2026-10-03 says every size now reads images and prompts up to 262,144 tokens. We did not run the models, and we did not re-download the September 28 board file.

Compare

GLiDE's self-run is 64.81. GEV-26B-Decide's adaptive figure is 62.48. Matilda's is 59.26. Those three, and this 63.88, cite the September 28 board's Jev cell of 57.91 and are not rows in that file. The kit's tie band is 0.25. The gap here is 5.97.

The size table's 208 of 231 is the Intern-Decision bundle's public items. The README says that count is not the official JevBench score, which adds a sealed tier, speed, and cost. Their Jev cell on that bundle is 199 of 231. The GGUF section's original-weight row for the 27B is 209 of 231, and the README says bf16 rounding moves one or two items off the table above.

Terms

StartLux-Decision
StartLux Labs' typed decision family. Five dense sizes from 0.8B to 27B, plus a 35B mixture with about 3B active. The server speaks POST /v1/systemone.
63.88
StartLux-Decision-27B on the authors' Decision Index 0.2.1 run, opened October 4, 2026. The README says the run is not on the September 28 board, where Jev 1.13 is 57.91.
CC BY-NC 4.0
The license on the StartLux-Decision weights. The code in the repository is Apache-2.0. Commercial use of the weights needs a separate license from StartLux Labs.

Sources

  1. StartLux, October 3
  2. StartLux-Decision README
  3. StartLux-Decision collection