Published
Updated
Maincode's Matilda card prints 59.26 on a self-run of Decision Index 0.2.1
Opened October 8, 2026, the Matilda card prints Decision Index 0.3 at 60.16, and the model API lists private false. Maincode posted Matilda on September 30, a 26.1B model under an Apache 2.0 line, and a self-run of the Decision Index 0.2.1 kit at 59.26 against the board's published Jev cell of 57.91. The card's median is 56.6 ms per request. The post's text says 56.8 ms. The run section still says the weights are private. We did not run the model.
Maincode posted on September 30 that it had released Matilda, which the card calls Australia’s first Jev-style model. The license line is Apache 2.0. The card prints 26.1 billion parameters, about 49 GiB in bf16, and a 1.3 million-parameter head that reads the final hidden state as one of 255 answer codes. The weights page we opened is Maincode/matilda-jev-v1. Dave Lemphers replied in the thread that this was the “top Jev spot” and linked that repo. The number under that phrase is a Decision Index score.
The card scores the model with the official Decision Index 0.2.1 kit. Matilda is 59.26, raw 68.89, breadth 58.06. The Jev (hosted) cells on the same card are 57.91, 68.09, and 57.08. The comparison table reprints board entries “as published on the leaderboard on 2026-09-28”: Surogate Rune 26B-A4B v3 at 57.44, Decider chat at 57.33, AutoJev-27B at 56.40. Those three match the September 28 board file already on this desk. The card says “We are submitting the model to the board,” so 59.26 is the author’s kit run.
Chance-corrected area scores, times 100, are on the card. Matilda: Knowledge and Reasoning 43.52, Language 66.36, Retrieval and Classification 61.34, Tools and Automation 77.72, Arts and Human Taste 43.65. Jev on the same lines: 51.40, 62.02, 55.42, 75.09, 37.66. Matilda’s Knowledge and Reasoning cell is the lower of the two.
The card counts 150,317 evaluation requests, every request answered, none truncated. The post’s text says the median is 56.8 ms. The card says 56.6 ms per request, on one AMD Instinct MI355X, bf16, one request at a time. It also says the board remeasures latency on its own RTX PRO 6000. An exact-copy check against 155,390 suite requests left 0 remaining copies.
Context on the card is up to 131k tokens on GPUs of at least 200 GB, and 32k on 96 GB cards. Inputs are text or JSON. Optional images go up to four, and the card says decision training was text-only. The model is English only. It returns decisions and does not generate text.
The post calls the release open-sourcing, and the license line is Apache 2.0. The run section still says: “While the weights are private, authenticate with hf auth login before downloading.” The clone URL printed on the card is github.com/Maincode/matilda-jev. The download id in the snippet is Maincode/matilda-jev. The citation URL is huggingface.co/Maincode/matilda-jev, without the -v1 on the page we opened. The local server in the snippet is POST /v1/systemone at localhost:8000, model matilda-jev-latest. Training is described as eight MI355X GPUs, full-weight bf16, on the MC-2 cluster in Australia. The sample curl response is an example on the card. We did not run the model.
Opened October 8, 2026, the same card’s results table says Decision Index 0.3 (public), over 150,317 requests, and prints Matilda-Jev at 60.16, raw 70.04, breadth 59.21. The area cells are 46.84, 64.80, 61.07, 79.07, and 46.22. The line says the score is awaiting official submission. A search of that card for the private-weights sentence returned no match. The model API lists private false and gated false. lastModified on that response is 2026-10-07. The download snippet still includes hf auth login. The 59.26 line above stays the September 30 self-run of edition 0.2.1. We did not run the model.
This site's reading
Editorial notes evaluating claims against primary sources, contextualizing findings alongside related implementations, and defining technical terms.
Verify
Post by @MaincodeAU, September 30, 2026, 09:36 UTC. The card calls Matilda Australia's first Jev-style model, Apache 2.0, 26.1 billion parameters, about 49 GiB in bf16, with a 1.3 million-parameter 255-way answer-code head on the final hidden state.
The self-score uses the official Decision Index 0.2.1 kit. Matilda is 59.26, raw 68.89, breadth 58.06. The card's Jev (hosted) cells are 57.91, 68.09, and 57.08. The comparison table reprints the leaderboard as published on 2026-09-28: Surogate Rune 26B-A4B v3 57.44, Decider chat 57.33, AutoJev-27B 56.40. The card says the model is being submitted, so 59.26 is not a row on that file yet.
Chance-corrected area scores, times 100. Matilda: Knowledge and Reasoning 43.52, Language 66.36, Retrieval and Classification 61.34, Tools and Automation 77.72, Arts and Human Taste 43.65. Jev on the same card: 51.40, 62.02, 55.42, 75.09, 37.66.
The card counts 150,317 evaluation requests, every one answered, none truncated. The post's text says the median is 56.8 ms. The card says 56.6 ms per request, one AMD Instinct MI355X, bf16, one request at a time. It says the board remeasures latency on its own RTX PRO 6000.
An exact-copy check against 155,390 suite requests left 0 copies. Context is up to 131k tokens on GPUs of at least 200 GB, and 32k on 96 GB cards. Inputs are text or JSON. Optional images go up to 4. The card says decision training was text-only. English only. Decisions only, no generated text.
The license line is Apache 2.0, and the post calls the release open-sourcing. The run section still says the weights are private and asks for `hf auth login` before download. The clone URL is github.com/Maincode/matilda-jev. The download id in the snippet is Maincode/matilda-jev. The page we opened is matilda-jev-v1. The citation URL on the card has no -v1. The sample server is POST /v1/systemone at localhost:8000, model matilda-jev-latest. Training is described as 8 MI355X GPUs, full-weight bf16, on the MC-2 cluster in Australia. The sample curl response is an example. Dave Lemphers's reply says "top Jev spot" and links the same Hugging Face repo. We did not run the model.
Compare
The 57.91 next to Matilda is the September 28 Decision Index board cell for Jev 1.13.0. Nace's page prints Drex 1.5 at 58.28. That 58.28 is not this board file, and 59.26 is a self-run the card says is still being submitted.
JevBench v1.5.4, the page we opened September 30, still lists Jev 1.13.0 at 72.1, official rank 3. That composite prices speed and cost. It is a different scale from 59.26.
Terms
- 59.26
- Matilda on Maincode's self-run of the Decision Index 0.2.1 kit, posted September 30. The card's Jev (hosted) cell is 57.91. The card says the model is being submitted, so this is not yet an independent board row.
- 56.6 ms
- The card's median per request, one MI355X, bf16, one request at a time. The post's text says 56.8 ms. The card says the board remeasures latency on its own RTX PRO 6000.
- 150,317
- Evaluation requests in the card's self-run. The card says every request was answered and none were truncated.