basal-1.5 scores 72.12% on Werdykt, and Jev 1.13.0 scores 81.64%

Remek Kinas posted basal-1.5 on October 4, 2026. Werdykt v1, exported that day, puts the 4.5B at 72.12% and Jev 1.13.0 at 81.64%. A closed set of 3,000 new Polish decisions is 93.1% against basal-1.0 at 87.4%. The site schedules the weights for October 5. Opened October 8, 2026, GitHub main titles the README basal 1.5 and says the models are on Hugging Face from 2026-10-05.

Remek Kinas posted basal-1.5 on October 4, 2026, at 06:58 UTC. The site describes three Apache-2.0 checkpoints for Polish and English, built on Bielik from SpeakLeash. basal-1.5 is 4.5B. basal-1.5-max is 11B. basal-1.5-mini is 1.5B, distilled from max. A request sends a state and questions whose answers were listed in advance. The server returns a probability for each option and does not write a reply. The process is basal-serve, and the path is POST /v1/systemone. Other calls of this shape are on decision models like Jev.

The site says the weights go on Hugging Face on October 5, 2026, at 09:00 Polish time. On October 4 the blog’s link to tag v1.5.0, and the commit it names, returned HTTP 404. The Hugging Face API for Remek/basal-1.5-4.5B returned 401. GitHub main still describes basal-1.0. The numbers below are the ones the October 4 site, blog, and JSON already print.

Opened October 8, 2026, that same README is titled basal 1.5. The models table says the weights are on Hugging Face from 2026-10-05. The Werdykt table on that README prints basal-1.5 at 0.721 and Jev 1.13.0 at 0.816, the October 4 export rounded to three places.

Werdykt v1 is their hidden set: 5,000 decisions, 10 categories in each language, about 250 questions a cell. The export is dated 2026-10-04. Unread answers count as wrong. Open weights ran on one H100 80 GB, except where the page says otherwise. API models, including Jev, went out over the network. Cost for an open model is H100 time at $3.95 an hour, taken from a full load. Cost for an API model is what the provider reported. The questions are not published. The page says none of them appear in basal’s training data.

ModelPL+EN95% intervalp50USD / 1,000Engine
Jev 1.13.081.64%80.70% to 82.58%286.4 ms0.0710API
basal-1.5-max, 11B77.28%76.24% to 78.28%67.1 ms0.1081SGLang
basal-1.5, 4.5B72.12%71.02% to 73.26%33.8 ms0.0477SGLang
basal-1.0, 4.5B66.04%64.70% to 67.18%26.5 ms0.1353fast
basal-1.5-mini, 1.5B60.66%59.40% to 61.92%14.2 ms0.0535fast

Jev’s Polish score in that export is 82.04%, and English is 81.24%. p95 is 378.8 ms. The JSON marks the row independent and records no protocol deviation. basal-1.5 is 72.44% in Polish and 71.80% in English. The SGLang rows for 4.5B and max raise the context limit from 4,096 to 32,768 tokens so the long documents fit. In fast mode on the same test the medians are 31.2 ms and 44.7 ms. A separate HTTP load, one question, fast, bf16, one H100, is 13.7 ms median for the 4.5B, p95 23.8 ms, 64.4 decisions a second with 32 clients. The homepage’s 13.7 ms line is that HTTP load. Werdykt’s median for the 4.5B stays 33.8 ms.

The closed Polish test is a different file. 3,000 new decisions, written after the weights were frozen and scored once: basal-1.5 93.1%, max 93.3%, mini 91.0%, basal-1.0 4.5B 87.4%, basal-1.0 1.5B 88.9%. The post’s gap for the 4.5B is 5.7 points, 95% interval +4.5 to +6.8. On the older Polish decision set, 7,081 cases, basal-1.5 is 92.3%, max 94.0%, mini 87.7%, basal-1.0 4.5B 88.6%, and basal-1.0 1.5B 85.1%. On Polish general knowledge, 3,079 cases, basal-1.5 is 74.1%, max 79.4%, mini 66.7%, basal-1.0 4.5B 73.7%, and basal-1.0 1.5B 65.6%. The blog says the general-knowledge gain for the 4.5B and the mini is not statistically significant. The post also prints 68.8% against 67.1% on 10 public English tasks, and 76.2% against 74.0% on a public English decision benchmark. Those two rows are not in the Werdykt JSON’s family table.

SOAM, state once, ask many, reads the state once and scores every question from that pass. On one H100, five questions about Polish decisions with a median of 336 tokens take 37.8 ms on the 1.5 engine and 111.7 ms on the 1.0 engine. Ten questions take 80.3 ms and 222.0 ms. The blog says fp32 answers match the answers from separate queries, and that SOAM does not speed a single question. The measurement skips HTTP. It is a different clock from the 13.7 ms homepage load.

The new request fields are multi, act, evidence, and facts. multi returns an independent probability for each label. act picks the cheaper of a mistaken action and a handoff, and the site marks it experimental. evidence returns a span with character offsets. On 200 documents the 4.5B’s token-F1 against a reference quote is 0.53, and the median moves from 40 ms to 85 ms. The site marks that experimental, and says the span engine is PyTorch. facts set to auto fills Polish dates and amounts in code before the model reads the state. The blog says the mini missed a 5% error target: observed error on a held-out test was 5.3%, so the shipped advice is to use the 1% threshold for that size.

Checked engine notes on the site, for the 4.5B: vLLM offline agreement with the fp32 reference is 0.994, at 147 decisions a second. SGLang is 0.995 and 104.6 a second, and the same answer as the basal engine on 98.3% of Werdykt questions, 97.0% above 8,000 tokens. On an Apple M4 Pro the blog prints 587 ms median for MLX 8-bit and 581 ms for 4-bit, with choice agreement 99.2% and 97.0% against bf16. Ollama and llama.cpp GGUF on that Mac are about half a second, agreement 0.994 for Q8_0 and 0.980 for Q4_K_M. ollama run by itself returns free text. The typed reply comes from basal-serve in front of it.

We did not run the models. The September 29 basal-1.0 post is the README on GitHub main. Its numbers are a separate release. Opened October 8, 2026, that file is titled basal 1.5.

This site's reading

Editorial notes evaluating claims against primary sources, contextualizing findings alongside related implementations, and defining technical terms.

Verify

Remek Kinas, @KinasRemek, October 4, 2026, 06:58 UTC, status 2106639959004176440. The post is in Polish. It names three sizes, max 11B, opt 4.5B, and mini 1.5B. The closed Polish test is 93.1% against 87.4% for basal-1.0, a gap of 5.7 points, 95% interval +4.5 to +6.8, on 3,000 new decisions written after the weights were frozen and scored once. The same post prints 92.3% against 88.6% on the basal-1.0 decision test, 68.8% against 67.1% on 10 public English tasks, and 76.2% against 74.0% on a public English decision benchmark. Five questions on one H100 take 37.8 ms instead of 111.7 ms. The post says fp32 answers match separate queries. The attached file is a video, so this page has no still.

basal.si5.pl, the blog at /blog/basal-1-5/, and the benchmark page were opened October 4. The JSON at /assets/benchmarks/werdykt-v1.json is dated 2026-10-04 and names the models in the table below. GitHub main, opened the same day, still describes basal-1.0. The blog's link to tag v1.5.0 and to commit b99c0e96 returned HTTP 404. The Hugging Face API for Remek/basal-1.5-4.5B returned 401. The site says the weights go on Hugging Face on October 5, 2026, at 09:00 Polish time. We did not download weights, and we did not run a model.

Compare

On Werdykt v1, Jev 1.13.0 is 81.64%, interval 80.70% to 82.58%, p50 286.4 ms, $0.0710 per 1,000 decisions, through the API. The JSON marks that row independent and records no protocol deviation. basal-1.5 at 72.12% and basal-1.5-max at 77.28% are the SGLang rows. Both raise the context limit to 32,768 tokens so long documents fit. Cygnet is 78.52% at p50 39.6 ms, and the page says that row uses a field-name adapter and a 16,384-token cap, with longer inputs counted as errors. Winnow-12B Q8 is 78.50% at p50 126.2 ms, with the context set to 65,536 as in its README. Decision-2.0-Nox-4B is 71.64% at p50 863.8 ms on an H100 NVL, a different machine from the H100 80 GB used for the other open weights. Those three are neighbors on the same export. They are not a second measurement of basal.

The closed Polish test of 3,000 decisions does not include Jev. The 93.1% is basal-1.5 against basal-1.0 at 87.4%. JevBench's official harmonic mean is a different ledger. Bespoke Nimble's 292 of 324 stays on its own holdout.

Decision-2.0-Nox-4B's own card, opened October 5, prints JevArena 63.6 and a Decision Index cell of 43.8. That card is on the Decision 2.0 page. The 71.64% on this export stays the Werdykt cell.

Terms

basal-1.5
Remek Kinas's Apache-2.0 decision family for Polish and English. Three sizes, 1.5B, 4.5B, and 11B, on a Bielik base. The server is basal-serve, POST /v1/systemone.
Werdykt
basal's hidden benchmark, v1, exported October 4, 2026. 5,000 decisions, 10 categories in Polish and in English, 250 questions a cell. The questions are not published.
SOAM
State once, ask many. basal-1.5 reads the state once and scores every question in the request from that pass. On the H100 comparison, five questions take 37.8 ms instead of 111.7 ms.

Sources

  1. Remek Kinas, October 4
  2. basal-1.5
  3. basal-1.5 blog
  4. Werdykt
  5. Werdykt v1 JSON
  6. rkinas/basal