Open decision models start from an open language model

Open decision checkpoints start from a public language model because a closed API does not release the weights a decision head has to read. Jeff v1.3 prints a mean of 30.4% for untrained Qwen3.5-0.8B and 92.8% with an adapter, across 13 of its own tests.

The open checkpoints on this desk sit on a language model someone else already trained. Jev and OpenAI’s Decisions API answer the same kinds of questions and do not publish a weight file. A decision head has to read parameters. A closed API serves the answer and keeps the file.

Fastino’s guide of August 5, 2026 says a closed family such as Claude or GPT cannot be fine-tuned in the ordinary way, because the weights are not public. Fine-tuning, in that guide, continues training on a smaller set. It changes weights the pretraining run already set, and the guide says the base sets the ceiling.

The public base is already weak at the decision format. Unsloth’s documentation, posted with the October 7 repository note, prints the gap for one epoch of LoRA with a Clef head.

Modeltyped-decisionsBANKING77CLINC150Holdout
Qwen3.5-0.8B36% to 73%7% to 74%19% to 76%78%
Qwen3.5-2B33% to 78%1% to 58%1% to 62%81%

The same page calls the starting band 30% to 37%, roughly chance, and the trained peak 78%. A second table on that page is the holdout accuracy, the memory, and the time: Qwen3.5-0.8B at 78% on 4 GB in 42 minutes, Qwen3.5-2B at 81% on 8 GB in 40 minutes, Llama 3.2 3B at 79% on 4.1 GB in 30 minutes, Gemma 4 E4B at 77% on 14.4 GB in 49 minutes, and a further fine-tune of Laya at 77% on 2.5 GB in 10 minutes. The doc says the language model reads the prompt once. A small head, the same design as Clef, scores every option. It does not write text. Daniel Han’s post the same afternoon says the local train fits in 3 GB. The doc’s smallest row is the Laya fine-tune at 2.5 GB.

The Unsloth post the same afternoon prints 20.7% to 74.3% for Qwen3.5-0.8B across three decision benchmarks, on 4 GB of VRAM. The mean of that row’s three columns, typed-decisions, BANKING77, and CLINC150, is 20.7 and 74.3. The 78% holdout is a separate column.

Jeff’s current card is jeff-base, v1.3, a fine-tune of Qwen3.5-0.8B under Apache 2.0. Across 13 adapter test sets the card’s mean goes from 30.4% for the untrained Qwen, to 38.6% for the v1.3 base alone, to 92.8% with the adapter. On a general panel of 4,599 questions the v1.3 base is 78.6% and the v1.2 base is 78.8%. The card tells a zero-shot user to stay on v1.2. v1.3 puts the options before the input so a repeated question can cache the prefix, and the base alone then falls apart on long lists: legal clauses, 100 options, drop from 66.0% on v1.2 to 7.4% on v1.3. With the adapter those two are 85.7% and 83.6%. The limitations line says a model this small does not reason.

The older Jeff card still prints a Gemma row on that same 4,599-question panel: Gemma 4 E2B at 62.5% untuned and 81.6% after the fine-tune, next to Qwen3.5-0.8B at 45.3% untuned and 78.7% as Jeff v1.2. That card says the base mix was trained on one workstation GPU, the 0.8B in about two hours, and that no closed model was the teacher for the base. Some of the adapter data on that card was written by a hosted Qwen.

Together’s September 23 tutorial states the step in one sentence: take an existing language model and turn it into a model that specializes in classification. The base they pick is Qwen3.5 4B. The table totals 37,840 examples. The next sentence says 38,340. The paragraph above the table says the sample is 38,000 questions. The run is priced at about $17 and about 25 minutes. The Tev1 card says the fine-tune keeps Qwen’s next-token head and should return one letter. Its development numbers, 880 of 1,000 and 300 of 300, are sets the card says were used while the model was built. The Tev1 page has the hosted price.

Kev ships the adapter and downloads the base. The v0.1.0 release note says the tarball does not contain Qwen2.5-0.5B. The first load fetches it from the Hub. The prototype card counts 9.3 million trainable parameters, 1.9% of that backbone. The current README’s smaller checkpoints freeze a Qwen3.5 base and train a rank-16 LoRA plus a pointer head. The Qwen3-generation 8B card, which that file marks as the previous generation, says the 4B base scores 0.688 on the same MMLU items and 0.787 on PAWS with a letter readout and no decision training. The default learning rate, 2e-4, trained those down to 0.60 to 0.66 and 0.56 to 0.71. Dropping the learning rate to 5e-5 recovered most of it. With the public examples held equal, that card’s out-of-domain gain from 0.6B to 4B is 14 to 19 points, and from 4B to 8B is 1 to 7. The lower learning rate is there to keep the base’s score on those two sets. The Kev page has the comparison with Jev.

Cloudflare’s Clef post describes an earlier experiment that read DiffusionGemma log probabilities, then says Clef uses a different backbone. The backbone is Qwen. Qwen3.8-27B and Qwen3.5-9B stay frozen. A routing head and rank-256 adapters train on top. At inference the model prefills and scores the legal schema choices in parallel. The post says that is faster than Jev and faster than the base Qwen models. The post names Qwen as the replacement and gives no further reason. The Clef page has the latency table.

An encoder is still a pretrained language model, picked for a different forward pass. Laya Studio’s ModernBERT page, marked updated September 23, says ModernBERT is Apache 2.0, which is why a derivative such as Laya can ship under the same license. The decision head is trained from scratch on top of the 395 million parameter encoder. The page names three things the design needs from the backbone: bidirectional attention, so each option marker can see the whole state and the other options; a masked-token pretraining objective, so the vector at a [MASK] already summarizes the context around it; and fast inference at a small batch. The encoder page reports 39.5 ms for one question and 158.6 ms for ten, on a T4. The same page says the untuned checkpoints score 0.362 on typed-decisions, under a 0.461 majority, and 0.425 on Banking77 against a published Jev figure of 0.870. Laya, WaterSheep, fsdecide, and NIRNAY are that encoder line. NIRNAY starts from Laya.

Liquid’s October 7 post trains Open d1 from Liquid’s own checkpoints. d1-3B starts from LFM2.5-VL-3B, a decoder vision-language model, after the text weights are averaged with LFM2.5-2.6B. d1-omni-600M starts from LFM2.5-Encoder-350M, a bidirectional encoder, and then adds audio and vision. The post says these models do not produce tokens. The closing list says the weights can be downloaded and fine-tuned without restrictions. The LICENSE file quoted on the Open d1 page conditions commercial use on annual revenue under $10,000,000.

Qwen3, Qwen3.5, and Qwen3.8 are the usual decoder file on this desk: Clef, Tev1, Nimble, imajev, Kev, Strands Decider, Decision 2.0, the larger Vela checkpoints, JEV-27B-VL, Open-Jev, Eikos, and PolicyLM through BidirLM from Qwen3-1.7B-Base. Unsloth’s Llama and Gemma rows put the same head on two other bases. basal-1.5 uses Bielik. Vela’s 0.3B checkpoint, on the Vela page, is an mmBERT encoder, and its tokenizer carries the Gemma Terms of Use. The StartLux page does not name a base. Drex DLM names NVIDIA Efficient-DLM-8B and says the block tensor names match Qwen3.

The same pages that recommend a public base also mark where a small one stops. Laya’s encoder page says a model of a few hundred million parameters does not stand in for a large language model on a multi-step policy, a written explanation, or a label list the size of Banking77. Jeff’s v1.3 card says to use the adapter, and that a model at this size does not reason. We did not train a head or load any of these weights.

This site's reading

Editorial notes evaluating claims against primary sources, contextualizing findings alongside related implementations, and defining technical terms.

Verify

Pages opened for this note on October 8, 2026: Fastino's August 5 fine-tuning guide, Unsloth's decision-model doc, the Jeff v1.3 card at mstrasser/jeff-base and the v1.2 card it supersedes, Together's September 23 tutorial, the Tev1 GitHub README and the Tev1-4B card, the Kev README, the v0.1.0 release note, and the Kev-8B Qwen3 card, Cloudflare's Clef blog, Laya Studio's ModernBERT and encoder pages (both marked updated September 23, 2026), and Liquid's Open d1 blog.

Jeff's 30.4%, 38.6%, and 92.8% are the v1.3 card's mean over 13 adapter test sets. The 4,599-question panel, 78.6% against 78.8%, is a different row on that card. Unsloth's 36% to 73% is typed-decisions for Qwen3.5-0.8B. The 78% holdout and the Llama and Gemma rows are the same doc's second table. The Unsloth post prints 20.7% to 74.3% for Qwen3.5-0.8B. The mean of typed-decisions, BANKING77, and CLINC150 on the doc is that pair. Together's table totals 37,840 examples. The next sentence on that page says 38,340. The paragraph above the table says the sample is 38,000 questions. The tutorial prices the run at about $17 and about 25 minutes. Kev's 9.3M trainable parameters and 1.9% are the 0.5B prototype card. The MMLU 0.688 and PAWS 0.787 lines are the Qwen3-generation 8B card, which that page marks as a previous generation.

Laya's 0.362 against a 0.461 majority, and Banking77 at 0.425 against a published Jev 0.870, are the encoder page. The Apache-2.0 sentence is the ModernBERT page. Liquid's backbone paragraph is the October 7 blog. The revenue condition is the LICENSE file already quoted on the Open d1 page. We did not train a head, load a checkpoint, or rerun a table.

Compare

The untrained column is the comparison these pages share. The percentages are not one test. Jeff's 30.4% is a mean over 13 adapter sets. Unsloth's 36% is typed-decisions for one 0.8B run. Laya's 0.362 is the author's typed-decisions reading of the untuned checkpoints, set under a 0.461 majority. Together's 880 of 1,000 is a development set the card says was used while the model was built.

A Gemma row and a Qwen row can share a recipe and still be different bases. Unsloth prints Llama 3.2 3B at 79% and Gemma 4 E4B at 77% beside Qwen. Jeff's older card prints Gemma 4 E2B at 62.5% untuned and 81.6% after the fine-tune, on its five-benchmark panel. The v1.3 card does not reprint that Gemma row.

Terms

Backbone
The pretrained language-model checkpoint a decision model keeps. The decision head reads its hidden states or its option logits.
Decision head
The small stack trained to score a fixed set of options. Clef, Kev, Laya, and Unsloth each publish their own version of it.
Prefill
One forward pass over the input. A decision model stops there and does not generate the answer token by token.
Open weights
A checkpoint whose parameters can be downloaded. Fastino's August 5 guide separates that from a release that also includes training data and training code.

Sources

  1. Fastino, how to fine-tune open weights
  2. Unsloth, train a decision model
  3. Unsloth, October 7 post
  4. Jeff v1.3 base card
  5. Jeff v1.2 card
  6. Together, train your own Jev for $17
  7. togethercomputer/tev1
  8. Tev1-4B model card
  9. jaredpalmer/kev
  10. Kev 0.5B release
  11. Kev-8B Qwen3 card
  12. Cloudflare, Introducing Clef
  13. Laya Studio, ModernBERT
  14. Laya Studio, encoder vs decoder
  15. Liquid, Open d1