Published
Updated
Strands Decider 2B scores 167 of 231 public JevBench tasks
This topic was created 8 days ago, and the information it contains may have evolved or changed since then.
Strands Labs posted Decider 2B on October 1, 2026, scoring 167 of 231 public JevBench tasks at a median 115 ms on an RTX 3090.
Strands Labs posted Strands Decider 2B on October 1, 2026.
The post calls it a small open-source decision model for local experiments, and it links strands-labs/strands-decider and the Hugging Face org StrandsAgents.
What the model is
The README says the first model in the family has 1.9 billion parameters. The torso is Qwen3.5-2B-Base. The language-model head is removed, so the checkpoint does not generate text. A pointer head of about a million parameters scores each option from one forward pass.
The torso is adapted with a rank-16 LoRA, and the head runs in fp32. The blog says this is the second architecture. The first used a slot head, and the authors say that one did worse. The reference checkpoint is v19.
The CLI is pip install strands-decider. The documented ask command loads
StrandsAgents/strands-decider-2B-hobson-v19 and accepts Choice, Noul, and Score
questions. strands-decider serve exposes POST /v1/systemone. An example in the
repository gates a weather tool inside a Strands agent.
The questions, the threshold, and the policy in that example were picked by hand.
Training is described as about 11 hours on one RTX 3090, or 1 hour 10 minutes on eight
H100s with FAST=1. The README says the data sources and their licences are listed in
the repository. We did not retrain it.
What v19 measures
The README’s performance table, and evaluation/README.md, print the same headline. On JevBench’s 231 public tasks, at the 3072-token window the run was preregistered at, v19 scores 167 of 231, 0.723. Brier is 0.342.
Expected calibration error is 0.052. On this repository’s split, easy, standard, and hard are 1.000, 0.875, and 0.505. Easy is 48 tasks, standard 72, hard 111.
Latency in that table is a median 115 ms and a 95th percentile of 299 ms per JevBench question on an RTX 3090 under WSL2. On an M3 Pro the warm median is 153 ms for tasks under 300 tokens, and 234 ms across all tasks. Served at the 4096-token default, the same notes print 168 of 231.
The authors call that one-task move noise. v20 reached 169 of 231 at 4096, missed four predictions, and did not replace v19.
Six retrains of the v17 recipe had a standard deviation of 3.2 tasks. The file says a single-run gap under about 10 tasks is unresolved. 167 against a published Jev count near 200 is larger than that band. This repository does not print Jev’s overall public accuracy.
It does print Jev at 0.27 on temporal_numeric, where v19 is also 0.267.
Where the rank is dated
evaluation/jevbench.md dates the ranking to the v1.4.2 board of September 25, 2026, 89 systems, using public-task accuracy. On that snapshot v19 is 50th overall and 3rd of 33 in the 2B class.
The blog repeats the 2B line and adds that, excluding models it calls just over 2B, the same snapshot is 1st of 30. The file is explicit that this is not the JevBench Score composite. Speed and cost were measured locally, and the sealed items are absent.
On that 2B list, decision-2b is 0.753, system-one-open is 0.732, v19 is 0.723, and
the board entry for decider-2b is 0.710, 164 of 231. decider-2b uses the same
Qwen3.5-2B-Base torso. The file says that board entry predates a v11 release which
scores 175 of 231 on their harness, eight tasks ahead of v19’s 167.
Of the 231 tasks, v19 is right on 19 that the board-entry decider-2b misses, and wrong
on 16 that the other gets.
The blog’s opening says the model can answer in tens of milliseconds. The latency section of the same post prints a median of around 115 ms on an RTX 3090. The graph of latency against token count is captioned as v18.
One limitation is in the evaluation README. With the state and the options held fixed, v19 gives the same answer to a changed question about 94% of the time. A question that reverses an obvious reading is often answered as if the reversal were not there.
On JevBench, 53 answers at confidence 0.9 or above were all right, at a mean claim of 0.957. The same note says the temperatures were fit on held-out classification, and that a threshold should be measured on the traffic it will see.
This site's reading
Editorial notes evaluating claims against primary sources, contextualizing findings alongside related implementations, and defining technical terms.
Verify
The blog is https://strandsagents.com/blog/introducing-strands-decider/. We opened it on October 2, 2026. The page says Published October 1, 2026.
It names Strands Decider 2B, links the GitHub repository and the Hugging Face org StrandsAgents, and describes a Qwen3.5-2B torso, a removed language-model head, a pointer head of just over a million parameters, and a rank-16 LoRA. The opening paragraph says answers come back in tens of milliseconds.
The latency section prints a median of around 115 ms on an RTX 3090, and around 153 ms for small tasks on an M3 MacBook. Figure 3 is captioned as v18, not v19.
The README at strands-labs/strands-decider, opened the same day, says the project is Apache 2.0 and points at a LICENSE file. It names the checkpoint StrandsAgents/strands-decider-2B-hobson-v19 and a local server on POST /v1/systemone.
The performance table prints v19 at 0.723, 167 of 231, Brier 0.342, ECE 0.052, tiers 1.000 / 0.875 / 0.505, RTX 3090 median 115 ms and 95th percentile 299 ms, and an M3 Pro warm median of 153 ms under 300 tokens and 234 ms on all tasks.
The same page says a 4096-token serve scores 168, and that a difference under about 10 tasks is unresolved.
evaluation/jevbench.md, opened the same day, dates the board position to JevBench v1.4.2 on September 25, 2026, 89 ranked systems. It prints v19 at public accuracy 0.723, overall rank 50th, and 3rd of 33 in the 2B class. Excluding three models the file calls just over 2B, it says first of 30.
It also says this run is the 231 public tasks, not the leaderboard composite, and that sealed items are absent. The nearest same-torso note says decider-2b's board entry is 164 of 231, while a later v11 scores 175 of 231 on their harness.
evaluation/README.md, opened the same day, says that with the state and options fixed, v19 gives the same answer to a changed question about 94% of the time. It prints temporal_numeric at 0.267 for v19. We did not run the model, and we did not re-open the live Benchmark Heaven page for this note.
Compare
The repository's 167 of 231 is public-task accuracy. The JevBench page records a different headline, the v1.5.4 composite, where Jev 1.13.0 is 72.1. This file does not print Jev's overall count on the 231. Drex 1.5 prints 201 of 231 for Jev on that public set.
Open-Jev prints 200 of 231. Those are other published cells.
On the 2B table in jevbench.md, decision-2b is 0.753, system-one-open is 0.732, v19 is 0.723, and the board entry for decider-2b is 0.710. The same file says decider-2b v11 scores 175 of 231 on their harness, eight tasks ahead of v19's 167. Clef is a hosted 27B and 9B pair with a different table.
Terms
- v19
- The reference Strands Decider recipe in the repository. At the 3072-token window it was preregistered at, the file prints 167 of 231. Served at 4096 tokens the same note prints 168. v20 scored 169 at 4096 and did not replace it.
- pointer head
- The small scoring head that replaces the language-model head. The blog says it is just over a million parameters and compares each option's hidden state with the hidden state at an answer position. The earlier slot head is described as worse.
- public 231
- JevBench's public tasks, 18 families, which this repository scores locally. The file says that number is not the JevBench Score composite, which also uses sealed items, speed, and cost.