Vela 2.0's 9B card prints 41.63 on Decision Index 0.2.1

The Vela 2.0 9B card prints Decision Index 0.2.1 at 41.63 with Noul calibration on, and 41.09 with it off. The same evaluation file prints the Lux base at 46.23 and does not print a Jev row. Four Apache-2.0 checkpoints also return span offsets. We did not run the weights.

vLLM posted Vela 2.0 on October 7, 2026, at 04:30 UTC. Xunzhuo Liu posted the same release at 07:04 UTC, and a reply one minute later links the blog. The blog page is dated October 6. The models belong to KR Labs and vLLM Semantic Router. A call takes text, or named parts such as a request, a context, and an answer, plus questions whose labels were supplied in the request. The cards say the model returns Choice, Noul, Score, Span, or Set, and that it does not write a reply. The license line on each card is Apache-2.0. The earlier Decision 2.0 checkpoints are on the Decision 2.0 page. Other calls of this shape are on decision models like Jev.

The four cards were opened October 8. The 9B card’s results table prints Decision Index 0.2.1 at 41.63, 38 benchmarks, with noul_calibration on. The same card says the default with that option off is 41.09. The 9B evaluation file prints the family, and it prints Decision-2.0-Lux-9B at 46.23 in that index column. It does not print a Jev row. The September 28 board’s Jev cell, 57.91, stays on the Decision Index page.

ModelParametersContextSpan headsIndex, calibration onDefault
0.3B307M8,192routernonenone
0.8B756M16,384router and broad16.0116.01
4B4.2B16,384router and broad31.9131.63
9B7.9B16,384router and broad41.6341.09

The parameter column for 0.8B, 4B, and 9B is the figure each card prints. The 0.3B card’s specification table does not print a parameter count. It names a 22-layer ModernBERT encoder, hidden width 768, and an 8,192-token budget. The 307M figure is the blog’s size table. The 0.8B and 4B index cells, and the blank 0.3B cell, are the 9B evaluation file. The 4B card prints 31.91 and says the default is 31.63. The 0.8B card’s selected results do not print an index cell. The blog’s uncalibrated table prints 16.01, 31.63, and 41.09, and it says those scores keep 79%, 74%, and 89% of the harness bases.

That file says the harness reproduces the released Decision 2.0 cards within 0.1 and prints Eos 20.14, Nox 42.55, and Lux 46.23. The Decision 2.0 cards this desk opened October 5 print Eos 20.1, Nox 43.8, and Lux 46.3. Eos and Lux sit inside that 0.1 band. Nox’s printed cell, 43.8, does not. The blog’s retention line uses 42.55, not 43.8.

The 9B card says it is derived from Decision-2.0-Lux-9B, itself built on Qwen3.5-9B. The 4B card names Decision-2.0-Nox-4B and Qwen3.5-4B. The 0.8B card names Decision-2.0-Eos-0.8B and Qwen3.5-0.8B. The 0.3B license line names Decision-1.0-Kai-0.6B and Vela-1.0-Encoder-307M, the second of those under MIT, from mmBERT. The same line says the 0.3B tokenizer carries the Gemma Terms of Use. The three larger cards do not add that tokenizer line. Training data keep their own licenses, and the cards say some of those are CC-BY-SA. Reported GPU memory for the parameters, in FP32, is about 3 GB, 17 GB, and 32 GB. The 0.3B card says Torch or ONNX on CPU or GPU.

Choice on the 9B card is one of 2 to 255 supplied options. Score is 2 to 10 ordered levels. Noul is a yes/no probability. Set returns every label above a threshold, and the card says those probabilities do not have to sum to one. Span returns a label, a start, an end, the text, and a probability. Offsets are Unicode code points in the selected field, start inclusive and end exclusive. A span question’s entry in answers is also a Noul. The blog says that number is the grid’s highest word probability.

The decoder cards ship two span heads. The router head is trained on PII, unsupported claims, and toxic spans. The broad head is trained on open extraction. The 0.3B card ships the router head. The blog says the broad head is trained last, with the rest of the model frozen, and that router outputs stayed bitwise identical on 700 rows per decoder. Dispatch sends the question ids pii, halu, and toxic to the router head. Other labels go to the broad head. A head field overrides that.

On a decoder, the blog says each word in a span question is read from a second copy of the target. Under a causal mask, a word’s state has seen only the words to its left. In the second copy, placed after the labels, the word has seen the whole text and every label. Targets longer than 2,048 tokens are read in overlapping windows. The blog says reading that second copy, and pooling each label over its whole block, raises short-text PII F1 from 0.828 to 0.959 on development data, and long-document PII F1 from about 0.06 to 0.94. Those two figures are development data. They are not the test table below. The 0.3B encoder reads the questions and the text in one bidirectional sequence, so it does not use the second copy.

The 9B evaluation file’s router table uses the same rows as the Vela 1.0 specialist named in each line. Prompt attacks from unseen families are 0.989 AUC for the 9B, against Vela 1.0 Guard at 0.792. The paired interval is +15.4 to +23.9 points. Multilingual HateCheck is 0.855 against 0.646, interval +20.3 to +21.5. The mean AUC over 14 safety sets is 0.921 for the 9B and the 4B, 0.875 for the 0.8B, and 0.871 for the 0.3B, against 0.704 for GLiNER2.5-Decide. The file says those safety sets are trained task families, so the Decide comparison is not zero-shot. Fact-check macro-F1 on that router table is 1.000 for the 9B against Vela 1.0 FactCheck at 0.501. Toxic-span character F1 is 0.446 against an mmBERT span head at 0.451, and that interval includes zero.

PII in the family table uses shipped calibration. The 9B’s 8K-token document F1 is 0.940. The research-scorer line for the same cell is also 0.940, and its interval against Vela 1.0 at 0.908 includes zero, −1.1 to +7.0. Short-text PII on the research scorer is 0.995 for the 0.3B and 0.985 for the 9B, against Vela 1.0 at 0.976. The 0.3B file note, repeated on the 9B family table, says final release selection also considered test results. The shipped 8K figure of 0.929 uses a threshold floor lowered after that test effect. The test-blind floor of 0.001 gives 0.896, and the research scorer gives 0.894. Decoder checkpoints were selected on development rows only.

RAGTruth in the family table is 0.774 for the 9B, 0.770 for the 4B, 0.652 for the 0.8B, and 0.706 for the 0.3B. The file says the LettuceDetect v2 mmBERT encoder was re-scored on the same 2,700 rows at 0.743. The 9B’s paired interval against that encoder is +0.5 to +5.4, and the file prints the difference as +3.0 points. The blog prints +3.1 on the same interval. The same family table prints LettuceDetect v2 qwen-2b, a generative model that only does hallucination, at 0.817 on RAGTruth and at 0.921 on the 10,698-example hallucination set. The 9B on that 10,698-example set is 0.885, against Vela 1.0 at 0.875. The file says the 10,698 LettuceDetect numbers come from its model cards, while the RAGTruth numbers were re-scored.

Open extraction uses the broad head. ACL-Verbatim, held out of training, is 24.5 word-F1 for the 9B, 24.4 for the 4B, and 23.6 for the 0.8B. GLiFormer-large on the same harness is 7.0. An ACL-Verbatim ModernBERT trained on that dataset, evidence only, is 53.4. Zero-shot NER across seven sets is 43.5 for the 9B broad head. GLiFormer-large is 63.6 and GLiNER-large-v2.5 is 61.3. SQuAD v2 dev is 71.5 for the 9B, and the file says that train split is in the broad-head data, so the dev score is not zero-shot. fast-decisions is 62.5 for the 9B against GLiNER2.5-Decide at 62.9. The file says the 0.3B fast-decisions cell, 49.4, is not an independent zero-shot result for that checkpoint.

The blog’s full router request, one A40, is 0.09 seconds for the 0.3B, 0.13 seconds for the 0.8B, 0.40 to 0.49 seconds for the 4B, and 0.60 to 0.71 seconds for the 9B. The evaluation file’s latency column is a different measurement: one question, about 3,000 tokens of context, mean of 100 questions on one A40. That column prints 2,288 ms for the 9B broad head, 1,499 ms for the 4B, and 477 ms for the 0.8B.

Loading on the cards is AutoModel.from_pretrained with trust_remote_code=True. The 9B, 4B, and 0.8B install lines ask for transformers>=5.17. The 0.3B card’s install line asks for transformers>=4.57. The blog’s CPU quickstart asks for transformers>=5.17 on the 0.3B as well. The method is system_one. The blog says each model ships vela2_serve.py, a FastAPI server for the same request on POST /v1/systemone, and the printed example uses port 8001. The 9B card’s recorded example chooses health at confidence 0.977 and marks the span “6 grams” unsupported at 0.982. The 0.8B card’s recorded unsupported span on the same prompt is “to 6 grams” at 0.778. The 0.3B card’s recorded health confidence on that prompt is 0.51.

The blog says the decoders train in three stages from the released Decision 2.0 models. Stage 1 is 4,000 steps, with a KL term on replayed base-model rows. The router span head is next, on a frozen backbone. The broad head is last. The 0.3B encoder trains in one stage of 101,000 steps from Kai’s Choice trunk, about 6.9 hours on one AMD MI325X. The blog says a synthetic row is kept only when a blind re-label agrees, and that 71% do. The paper named on the blog, “Vela 2.0: Towards Open Foundation Routing Models,” is marked forthcoming. The blog’s last section says the next step is native signal-backend integration in vLLM Semantic Router, replacing one classifier per signal with one Vela 2.0 request. We did not read the router source for that backend. We did not run these weights.

This site's reading

Editorial notes evaluating claims against primary sources, contextualizing findings alongside related implementations, and defining technical terms.

Verify

vLLM, @vllm_project, October 7, 2026, 04:30 UTC, status 2107690067619930450. The post names four sizes, 0.3B to 9B, Apache-2.0, and links the Hugging Face collection. Xunzhuo Liu, @XunzhuoLiu, 07:04 UTC the same day, status 2107728799400133041, posts the same release. The attached file is a video, so this page has no still. A reply at 07:05 UTC, status 2107728921039085756, links the blog. A later reply, status 2107851245826474399, links a Hugging Face space. We did not open that space.

The blog page is dated October 6, 2026. The four model cards and the 9B EVALUATION.md were opened October 8. The index table below is that file's family table, with the default scores from the 9B and 4B cards and from the blog's uncalibrated table. The 0.8B card's selected results do not print an index cell. The 9B file leaves the 0.3B index cell blank. None of these tables prints a Jev row. We did not run the weights.

Compare

The 9B evaluation file puts 41.63 beside Decision-2.0-Lux-9B at 46.23 in one column, with noul_calibration on for the Vela cells. A separate line prints the 9B default, calibration off, at 41.09. The blog says that default keeps 89% of the harness base. The Decision 2.0 Lux card this desk opened October 5 prints 46.3. The same file says its harness reproduces the released cards within 0.1 and prints Eos 20.14, Nox 42.55, and Lux 46.23. Eos at 20.1 and Lux at 46.3 sit inside that band. Nox's printed cell is 43.8, which does not.

The September 28 board's Jev cell is 57.91, on the Decision Index page. The Vela file does not put 41.63 and 57.91 in one table. RAGTruth on the same file is a different test: 9B at 0.774, the LettuceDetect v2 mmBERT encoder at 0.743, and the generative qwen-2b row at 0.817.

Terms

Vela 2.0
Four Apache-2.0 checkpoints from KR Labs and vLLM Semantic Router, posted October 7, 2026. A call can return Choice, Noul, Score, Span, or Set. The model does not write a reply.
41.63
Vela-2.0-9B on Decision Index 0.2.1, 38 benchmarks, with noul_calibration on. The 9B card says the default with that option off is 41.09. The evaluation file prints the Lux base at 46.23 and does not print a Jev row.
Span
A Vela 2.0 answer that returns a label, character offsets, the text, and a probability. The 9B card says offsets are Unicode code points, start inclusive and end exclusive. The 0.3B card ships the router span head. The three larger cards also ship a broad span head.

Sources

  1. Xunzhuo Liu, October 7
  2. vLLM, October 7
  3. Vela 2.0 blog, October 6
  4. Vela 2.0 collection
  5. Vela-2.0-0.3B
  6. Vela-2.0-0.8B
  7. Vela-2.0-4B
  8. Vela-2.0-9B
  9. Vela-2.0-9B evaluation