Published
Jared Palmer's Kev family copies Jev's API on Qwen, and trails it on new sources
Cognition's VP of engineering published Kev, an Apache 2.0 family of Jev-like models you train and serve yourself. On data the models never trained on, Kev-8B scored 79.6% against Jev 85.7%. JevBench listed kev 0.6B at 66.7. A Qwen3.5 generation followed the same day.
Jared Palmer, VP of engineering at Cognition, posted a Kev family on September 20. The first public note, two days earlier, was Kev-0.5B on Qwen2.5-0.5B. The update is 0.6B, 4B, and 8B on Qwen3, same LoRA plus small pointer head, Apache 2.0. The repo is jaredpalmer/kev. Weights sit in a Hugging Face collection.
The API matches TypeSafe’s System One shape. Point the official Python SDK at a local server and change base_url. Choice, noul, and score share one request. Questions see the state and cannot see each other.
On data Kev never trained on, Palmer posted Kev-8B at 79.6% and Jev at 85.7%. He wrote that Jev still leads most of his benches, and that Kev beat Jev on some logic-recognition tasks, with 4B the most interesting size in that family. A later post in the thread puts the remaining gaps at knowledge (MMLU 70 versus 90) and day-precision date arithmetic (60 versus 93). The weekend of hill-climbing on Modal cost $228. Most training runs were $10 to $35.
Serving, from the same post: Kev-4B on a 32 GB Mac in bf16, about 300 ms for five questions, about 40 ms on an H100. Repeated documents hit a KV cache and run 2 to 2.5 times faster. Train times on one H100: 4B in 40 minutes, 8B in 83 minutes.
That night Palmer posted that a Qwen3.5 family was up, “about +7 or 8% on MMLU Pro.” The README now ships Kev-0.8B, 4B, and 9B on Qwen3.5 bases. Its new-source table, development then test: Kev-9B 0.812 / 0.837, Kev-4B 0.794 / 0.832, Jev 0.857 on development. Brier on that split is 0.291 / 0.243 for Kev-9B against Jev 0.211. The README says the Jev column is not a controlled architecture comparison, because nobody outside TypeSafe knows which datasets Jev trained on. No Jev outputs were used for training.
Qwen3.5 mixes attention with Gated DeltaNet layers that ignore attention masks. The README’s workaround is one row per question, with the state computed once and reused. On a Mac, those recurrent layers have no fast kernels yet. Median model time in bf16 on an M5, five questions: Kev-4B 779 ms, against 174 ms for the Qwen3 Kev-4B checkpoint. The notes say to serve the Qwen3 weights on Apple Silicon if latency matters.
JevBench is the independent row. Florian S posted kev 0.6B as the strongest Kev variant in that suite, then at rank 9. The table we read lists kev 0.6B at 66.7 (rank 11), kev 0.5B 63.1, kev 4B 62.2, kev 8B 58.3, all marked research previews, measured from commit 20fa626 on an RTX 3090. Jev 1.13.0 is 75.4 on that page. We did not run either system.
This site's reading
Editorial notes evaluating claims against primary sources, contextualizing findings alongside related implementations, and defining technical terms.
Verify
Palmer's September 20 family post (quoted from the September 18 0.5B prototype) is the primary source. GitHub README (jaredpalmer/kev, Apache 2.0) is the spec. Qwen3 family in that post: 0.6B, 4B, 8B. Out of domain, data Kev never trained on: Kev-8B 79.6%, Jev 85.7%. Kev-4B on a 32 GB Mac in bf16 about 300 ms for five questions, about 40 ms on an H100. Train times on one H100: 4B 40 minutes, 8B 83 minutes. Later in the thread: MMLU 70 versus 90, day-precision date arithmetic 60 versus 93, weekend Modal bill $228, most runs $10 to $35. A same-day follow-up says a Qwen3.5 family was pushed, about +7 or 8% on MMLU Pro. The README now ships 0.8B / 4B / 9B on Qwen3.5 bases; new-source development/test for Kev-9B is 0.812 / 0.837 against Jev 0.857 on development. Florian S posted kev 0.6B on JevBench as the strongest Kev variant then, rank 9. The table we read lists kev 0.6B at 66.7 (rank 11), kev 0.5B 63.1, kev 4B 62.2, kev 8B 58.3, all as research previews from commit 20fa626 on an RTX 3090. We did not train or serve Kev. TypeSafe's Master Customer Agreement section 2.3(f) forbids publishing benchmarks of the Services; the Jev column is reported as published.
Compare
SemIf is the closest open row on JevBench (74.6, 0.7 behind Jev) and reads option logits from Qwen3.5-4B. Kev copies POST /v1/systemone with a LoRA plus pointer head, so the TypeSafe SDK works with a base_url change. jeff copies the same HTTP shape on GLiFormer and trails Jev on public labels (AG News 75.5% versus 90.5%, JevBench 66.9). Laya is the ModernBERT-sized encoder (JevBench 70.1). Nimble is a 9B LoRA on 324 synthetic labels. LocalJev prompts JSON probabilities and says that is not a logit read. Palmer's new-source table is his own suite (decision-v7 / transfer-v4), not JevBench. JevBench's Kev rows are independent of that table and use the Qwen3 checkpoints.
Terms
- Kev
- Jared Palmer's Apache 2.0 family of Jev-like decision models. Rank-16 LoRA plus a pointer head on Qwen bases. Serves POST /v1/systemone.
- Pointer head
- Kev's scoring layer. It compares each option's </opt> hidden state with the question's <decide> hidden state, then softmaxes those scores.
- New sources
- Palmer's out-of-domain split: datasets and policy rule types Kev was not trained on. The September 20 post puts Kev-8B at 79.6% there against Jev 85.7%.