AutoTrust scores GEV-26B-Decide at 62.48 on its own Decision Index run

AutoTrust posted GEV-26B-Decide on October 3, 2026, a Gemma 4 26B model with about 4B active. The card's own Decision Index 0.2.1 scoring prints an adaptive 62.48 against the September 28 board's Jev cell of 57.91. Adaptive thinking on Knowledge misses the kit's latency cap. A 1,754-question holdout moves from 73.3% to 83.4% when thinking is allowed. The project is not affiliated with TypeSafe.

AutoTrust posted GEV-26B-Decide on October 3, 2026. The model card says these are the same weights previously published as autotrust/JEV-Gemma4-26B-A4B. The base is google/gemma-4-26B-A4B-it, 26B parameters with about 4B active. The adapter, the head, and the calibration files are Apache-2.0. The base model stays under the Gemma 4 license. The card says the project is not affiliated with TypeSafe. The posts that morning are video, so this page has no still from them. The card does not name a person.

What they scored, and how long thinking takes

The card runs the kit’s score --edition 0.2.1 and prints an adaptive balanced skill of 62.48, raw 70.66, breadth 62.00, against the September 28 board’s Jev cell of 57.91. Knowledge and Reasoning is adaptive thinking on all ten benchmarks in that area, 33,856 scored requests. The other four areas stay System 1, taken from a finished run of 150,759 requests and 0 errors, stored as runs/jev-gemma4-26b-a4b.

Adaptive area skills on the card are Knowledge 0.602, Language 0.636, Retrieval 0.679, Tools 0.697, and Arts 0.415. System 1 latency is about 45 ms on one B200. The board’s cap, which the card restates, is a median at or under 1,000 ms on an RTX PRO 6000. The adaptive median on Knowledge is 13.4 seconds, and the 90th percentile is 33 seconds. Speculative decoding brings the median to 7.7 seconds. The card says that adaptive mode does not meet the Decision Index latency limit on Knowledge and Reasoning.

An older card for the same weights printed a System 1 index of 58.05 against Jev at 57.91 and JEV-27B at 53.30. The 62.48 figure does not replace 58.05. The October 3 post prints 53.3 for the lab’s JEV-27B-VL, 62.48 for this checkpoint, and Clef at 61.2, and it says all three are self-reported. The README table does not print 61.2. That 61.2 is the number on Fastino’s chart the same morning, captioned there as a self-reported Cloudflare score. 62.48 is not a name in the September 28 file. GLiDE’s 64.81 is another lab’s self-run of the same scorer.

A holdout, and the sets where thinking moves the answer

On 1,754 questions the card says are outside Decision Index, a threshold of 0.8 thinks on 47.8% of rows. System 1 scores 73.3%. The adaptive mix scores 83.4%. Always thinking scores 83.8%. Expected calibration error is 0.035 for System 1 and for the adaptive mix. The mix is half the System 1 probability and half the thinking probability.

AQuA-RAT goes from 68.1% to 89.0%, with thinking on 53.9% of that set. LogiQA goes from 55.3% to 80.3%, thinking on 63.7%. CommonsenseQA goes from 86.7% to 85.0%, thinking on 32.0%. The Knowledge area skill moves from 0.429 to 0.602.

Inside it, GPQA Diamond goes from 42.9% to 78.6%, CRUXEval from 67.5% to 90.7%, and GSM8K from 97.6% to 99.1% with thinking on 5% of items. HLE goes from 8.4% to 17.8%, against a chance line of 16.4%, and the card says System 1’s confident answers there were 1.6% correct. ChessBench stays flat, 23.7% to 23.6%.

The card says to leave thinking off for classification, retrieval, and tool routing. BFCL goes from 94.4% to 96.0% with thinking on 7%. BANKING77 macro-F1 goes from 88.0 to 85.0, and the training split of that set overlaps the run. CLINC150 with 150 options is 95.5% in System 1. The same labels placed among 255 options are 89.0%.

The post’s four puzzles are 200 generated items each, with a think budget of 8,192 tokens. Minesweeper goes from 21.0% to 86.0% (chance 25). Wordle goes from 52.0% to 100 (chance 12.5). Connect Four goes from 54.0% to 99.5% (chance 14.5). Sudoku goes from 76.0% to 99.5% (chance 11.1).

The same card prints whole games where thinking does not help. Snake collects 3 pieces of food in 18 steps without thinking, and 11 pieces in 90 steps with it, about 21 seconds a step. Six games of Connect Four go from a 1-5 record to 1-4-1. 2048 scores 1,476 against 1,016. Flappy Bird clears 0 pipes either way.

Screens and the request

Computer use, System 1, thinking off, one B200: 95% of 60 tasks, about 85 ms a click. The card prints the same 95% for JEV-27B-VL at about 260 ms and for JEV-9B at about 200 ms. Boxes with no element text fall to 15%. A robot arm is 40% of 20 scenes at 61 ms a decision, against 75% at 239 ms for JEV-27B-VL and 50% at 163 ms for JEV-9B. After a grasp, 8 of 9 cubes land in the tray. Eight direct motor commands score 0 of 10. VL-RewardBench, 1,247 pairs, is 78.4% against 78.3% for JEV-27B-VL on the protocol the card names.

A needle check is 10 of 10 at 4K, 32K, 64K, and 128K tokens, with median latency 0.15, 1.6, 5.0, and 17.8 seconds. Above 128K is untested. The native context is 262,144. The API is POST /v1/decide, with thinking off, auto, or on, and a default threshold of 0.8. System 1 throughput on the card is 257 decisions a second at 64 clients on one B200. The card says the BANKING77 and CLINC150 training splits were included, that no test split of those sets was, and that MMMU is not a valid score here.

We did not run the checkpoint.

This site's reading

Editorial notes evaluating claims against primary sources, contextualizing findings alongside related implementations, and defining technical terms.

Verify

AutoTrust, @AutoTrustAI, October 3, 2026, 05:29 UTC, status 2106255213854409172. The post is four videos and no still, so this page has no share image. The text prints Minesweeper 21% to 86%, Connect Four 54% to 99.5%, Wordle 52% to 100%, and Sudoku 76% to 99.5%. It also prints a Decision Index line of 53.3 for the lab's JEV-27B-VL, 62.48 for GEV-26B-Decide, and Clef at 61.2, and it says all of those are self-reported. A companion post at 04:23 UTC is status 2106238621301063751.

The Hugging Face README at autotrust/GEV-26B-Decide says the weights were previously published as autotrust/JEV-Gemma4-26B-A4B. The base is google/gemma-4-26B-A4B-it, 26B total, about 4B active. Adapter, head, and calibration are Apache-2.0. The base stays under the Gemma 4 license. The card says the project is not affiliated with TypeSafe. We did not find a personal author name on the card.

The card's own `score --edition 0.2.1` prints adaptive balanced skill 62.48, raw 70.66, and breadth 62.00, against the board's Jev cell of 57.91. Knowledge and Reasoning uses adaptive thinking on all ten benchmarks, 33,856 scored requests. The other four areas are System 1 from a complete run of 150,759 requests and 0 errors, stored as runs/jev-gemma4-26b-a4b. Adaptive area skills are Knowledge 0.602, Language 0.636, Retrieval 0.679, Tools 0.697, and Arts 0.415.

The card says the board's latency cap is a median at or under 1,000 ms on an RTX PRO 6000. System 1 is about 45 ms on one B200. The adaptive median on Knowledge is 13.4 seconds, 90th percentile 33 seconds. Speculative decoding brings that median to 7.7 seconds. Adaptive mode does not meet the latency limit on Knowledge and Reasoning. The README table does not print 61.2. An older card printed a System 1 index of 58.05 against Jev 57.91 and JEV-27B at 53.30. 62.48 does not replace 58.05.

We did not run the model, and we did not re-download the September 28 board file. 62.48 is the card's self-score.

Compare

62.48 is not a name in the September 28 board file, which lists Jev at 57.91. Fastino's GLiDE blog prints 64.81 on another self-run of the same scorer. The public kit's tie band is 0.25. Both gaps against 57.91 sit outside it, and neither run is a board row.

The tweet's 61.2 for Clef matches the number on Fastino's October 3 chart, which captions Cloudflare's scores as self-reported. The Clef model card this desk read does not print that composite. The tweet's 53.3 is the lab's JEV-27B-VL, in line with the older card's 53.30, not the board's 57.91.

The 1,754-question holdout is outside Decision Index. Thinking on 47.8% of those rows moves accuracy from 73.3% to 83.4%, against 83.8% if the model always thinks. CommonsenseQA in that set goes from 86.7% to 85.0%. BANKING77 macro-F1 goes from 88.0 to 85.0, and the card says the training split of that set was included.

Terms

GEV-26B-Decide
AutoTrust's October 3, 2026 checkpoint. About 4B active parameters on google/gemma-4-26B-A4B-it. Previously released as JEV-Gemma4-26B-A4B. Adapter, head, and calibration are Apache-2.0. The base stays under the Gemma 4 license. Not affiliated with TypeSafe.
62.48
The card's adaptive balanced skill from its own Decision Index 0.2.1 scoring, against the September 28 board's Jev cell of 57.91. Raw 70.66, breadth 62.00. Not a row in that board file. Adaptive mode misses the kit's latency cap on Knowledge and Reasoning.

Sources

  1. AutoTrust, October 3
  2. AutoTrust, earlier the same morning
  3. GEV-26B-Decide