Decision 2.0's Vega card prints 56.5 on its own Decision Index run

vLLM Semantic Router posted Decision 2.0 on October 3, 2026. Six Apache-2.0 checkpoints run from Kai at 0.6B to Vega at 27B. Vega's card prints 56.5 on a reproduction of Decision Index 0.2.1 and does not print a Jev row. Its JevArena cell is 74.0, which the card calls level with AutoJev-27B at 72.1. Vela 2.0, posted October 7, is a later fine-tune and has its own page.

Xunzhuo Liu posted Decision 2.0 on October 3, 2026, at 01:25 UTC. The vLLM account had posted the same release two minutes earlier. The models belong to vLLM Semantic Router. A call takes text or JSON, plus questions whose answers were listed in advance, and returns a probability for every option. The cards say the questions in one request are scored in one forward pass, and that the model does not write a reply. The license line on each card is Apache-2.0. Other calls of this shape are on decision models like Jev.

The six cards were opened October 5. Each prints JevArena, a human-labelled transfer score, and a Jev Decision Index cell. The index footnote says the Decision 2.0 cells are an independent reproduction with the official 0.2.1 kit on the released weights. Other models in those tables are a public board snapshot from September 28, 2026. The footnote says training data was audited at row level against all Index test items. No table on these cards includes Jev. The September 28 board’s Jev cell, 57.91, stays on the Decision Index page.

ModelParametersContextJevArenaTransferIndexMedian
Kai0.60B8,19248.645.916.34.9 ms
Eos0.75B16,38453.950.320.16.0 ms
Sol1.88B16,38452.151.329.57.2 ms
Nox4.21B16,38463.652.343.812.9 ms
Lux7.94B16,38468.156.246.318.4 ms
Vega29.37B32,76874.058.756.571.4 ms

The names say 0.6B, 0.8B, 2B, 4B, 9B, and 27B. The parameter column is the figure each card prints. Vega is an adapter on Qwen3.8-27B, and its install line adds peft. Kai is a fine-tune of Qwen3-0.6B-Base. Eos, Sol, and Lux are fine-tunes of the Decision 1.0 checkpoint with the same name. Nox is a fine-tune of Qwen3.5-4B-Base. The median is one question on a single GPU. The cards do not name that GPU.

JevArena’s footnote says every model answers the same frozen prompts, and that a missing or invalid answer counts as an error. It does not say how many prompts that is. Transfer is the median macro-F1 over 15 human-labelled tasks, times 100. That footnote is on all six cards.

Vega’s highlight calls 74.0 statistically level with AutoJev-27B at 72.1, and it does not print the interval. On transfer, Vega, AutoJev-27B, and Eikos-27B each print 58.7. Jebadiah-27B prints 57.8. Nox’s highlight calls 63.6 statistically level with Decider 4B at 61.9. On transfer that neighbor is ahead, 55.5 to 52.3. Sol’s JevArena cell is 52.1, below Eos at 53.9. The cards still call each of those two first inside its own size class.

Against the Decision 1.0 checkpoint of the same name, the index cells move from 6.5 to 16.3 for Kai, 18.4 to 20.1 for Eos, 25.3 to 29.5 for Sol, 34.4 to 43.8 for Nox, and 43.5 to 46.3 for Lux. The JevArena cells on that pairing move from 35.9 to 48.6, 42.5 to 53.9, 45.8 to 52.1, 56.5 to 63.6, and 65.8 to 68.1. Vega has no Decision 1.0 row. Its highlight says the index cell is 11.2 points above Lux. The printed cells, 56.5 and 46.3, differ by 10.2.

Loading on the cards is AutoModel.from_pretrained with trust_remote_code=True, and the install line asks for transformers>=5.17. The method is system_one, with choice, noul, and score questions in one dict. Kai’s decision_config.json prints max_options 255. That file does not print a question cap. The vLLM post says one request can hold 64 questions. The cards do not print 64.

Werdykt v1 already lists Decision-2.0-Nox-4B at 71.64%, p50 863.8 ms, on an H100 NVL. That export is a different test from the 63.6 and 43.8 on Nox’s card. We did not run these weights.

On October 7, 2026 the same account posted Vela 2.0. The 0.8B, 4B, and 9B cards say they are derived from Eos, Nox, and Lux. The 9B card prints 41.63 on Decision Index 0.2.1 with Noul calibration on. The write-up is Vela 2.0.

This site's reading

Editorial notes evaluating claims against primary sources, contextualizing findings alongside related implementations, and defining technical terms.

Verify

Xunzhuo Liu, @XunzhuoLiu, October 3, 2026, 01:25 UTC, status 2106193824091681126. The post names sizes from 0.6B to 27B and links the Hugging Face collection. The attached file is a video, so this page has no still. vLLM, @vllm_project, status 2106193438098256191, 01:23 UTC the same day, says the weights are Apache-2.0, load with Transformers, and answer 64 questions about a request in one forward pass.

The six model cards were opened October 5. The table below is those cards. Each index footnote says the Decision 2.0 cells are an independent reproduction with the official 0.2.1 kit on the released weights, and that other models' cells are a public board snapshot from September 28, 2026. The same line says training data was audited at row level against all Index test items. None of the six tables prints a Jev row. Kai's decision_config.json prints max_options 255 and does not print a question cap. We did not open that file on the other five, and we did not run the weights.

Compare

The September 28 board's Jev cell is 57.91, on the Decision Index page. Vega's own index cell is 56.5. The card does not put those two numbers in one table. The kit's tie band is 0.25. A gap of 1.41 is outside that band, and it is still a self-run next to a board snapshot.

Vega's highlight says the index cell is 11.2 points above Lux. The printed cells are 56.5 and 46.3, a difference of 10.2. Nox's highlight calls 63.6 statistically level with Decider 4B at 61.9 and does not print the interval. On transfer, Decider 4B is 55.5 and Nox is 52.3. Werdykt v1, a different export, prints Decision-2.0-Nox-4B at 71.64%. That row stays on the basal page.

Terms

Decision 2.0
Six Apache-2.0 checkpoints from vLLM Semantic Router, posted October 3, 2026. The cards name Kai, Eos, Sol, Nox, Lux, and Vega. A call returns a probability for every listed option and does not write a reply.
56.5
Decision-2.0-Vega-27B on the authors' Decision Index 0.2.1 reproduction, opened October 5, 2026. The card does not print a Jev row. The September 28 board's Jev cell is 57.91.
JevArena
The score printed on the Decision 2.0 cards for a frozen prompt set. The footnote says a missing or invalid answer counts as an error. The cards do not print how many prompts that set holds.

Sources

  1. Xunzhuo Liu, October 3
  2. vLLM, October 3
  3. Decision 2.0 collection
  4. Decision-2.0-Kai-0.6B
  5. Decision-2.0-Eos-0.8B
  6. Decision-2.0-Sol-2B
  7. Decision-2.0-Nox-4B
  8. Decision-2.0-Lux-9B
  9. Decision-2.0-Vega-27B