Eikos-27B prints 82.9 on its JevBench hard row, and 2.4% error once confidence is 0.90

This topic was created 10 days ago, and the information it contains may have evolved or changed since then.

Caio Vicentino released Eikos-4B and Eikos-27B under MIT. His harness, same items for every system: public JevBench hard tier 82.9, 72.1, Jev 73.0, Laya 35.1. At confidence 0.90 or higher, error on 7,140 items is 2.4% for both Eikos sizes and 5.9% for Jev. A later chess screenshot has Eikos at 53.9% over 388 games. A paper-trading file opened at 09:20 UTC on September 26, cycle 626 of 864, has Jev equity at $10,051.22 and Eikos at $9,653.19. Opened September 30, the model card's RuleArena row is 0.50, 0.50, and a blank Jev cell. The repo's only JSON file is a list of eleven label ids.

Caio Vicentino posted on September 23 that Eikos, a pair of open decision models, was coming in 4B and 27B. The September 24 post says the weights are up, MIT, with the same headline figures. The repo is caiovicentino/eikos. Models are caiovicentino1/Eikos-4B and caiovicentino1/Eikos-27B, each with FP8, INT4, and MLX builds. The training set is caiovicentino1/eikos-decisions. He says the name is Aristotle’s word for the probable, and that the intended job is applying a stated rulebook to a case in finance and trade.

The README’s table is his harness, the same items for every system. Public JevBench hard tier: Eikos-27B 82.9, Eikos-4B 72.1, Jev 73.0, Laya 35.1. Those cells do not print a denominator. When the system is at least 90% confident, error on 7,140 items across 6 suites is 2.4% for both Eikos sizes, 5.9% for Jev, and 37.3% for Laya. A probe that hides the decision inside 64k tokens, averaged over three positions, is 88.3 for the 27B and 74.2 for the 4B. Jev’s cell is a dash, with the note that the API limit is 32k. Laya’s cell says the context is 512 to 1k tokens. The post says eight batteries against Jev and Laya. The confidence row in the README says six suites.

On September 30, 2026 the git tree at main contained one JSON file, release_tools/known_label_errors.json. It lists eleven TAT-QA ids. That file is a label list. Eight batteries and six suites remain two printed counts.

The 4B recipe starts from Qwen/Qwen3.5-4B. LoRA rank 64, one epoch, learning rate 1e-4. The loss is soft cross-entropy on the option-letter logits, with the option order permuted, plus an auxiliary rationale loss at weight 0.3. The released checkpoints use temperature 1. The 27B run is a separate train_final.sh target. This README does not name that base model. Answers come from letter logits in one pass over the shared state. More than 26 options go through a tournament. POST /v1/systemone is the HTTP shape, and a session can append state instead of resending it. The README says vLLM 0.30 or newer is required: on 0.11, batched long requests dropped 3 to 6 points against the PyTorch reference.

The file says no evaluation data was used for training. Decontamination against the public JevBench items is an 8-gram check, and the snapshot also holds out FinQA-derived items and the trade-rule families reserved for the eval. The public dataset leaves out 2% of the rows they trained on, so a retrain is close and not bit-identical. A quantized build ships only if accuracy stays within 1 point of bf16, ECE within 0.01, and at least 97% of answers are unchanged. Code is MIT. The fine-tune deltas are MIT. The base models are Apache 2.0. The data card is CC BY 4.0. The notes say the models apply rules they are given, do not predict prices, and are not legal, tax, or investment advice. They also say a single pass does not do long multi-step reasoning, and they point at RuleArena on the model card for that limit. We did not run evaluation/eval_vllm_suite.py.

On September 30, 2026 we opened the Eikos-27B model card. It names the 27B base as Qwen/Qwen3.8-27B. The card’s RuleArena NBA row, balanced accuracy, is 0.50 for the 4B, 0.50 for the 27B, and a blank cell for Jev. The limitations line says that on RuleArena, NBA salary-cap rulebooks of about 25k tokens, the model does not discriminate.

The release-builds heading on that card says 7 suites and 7,371 items. The chart note says 7,140 items are common to all four systems. Those counts sit beside the README’s 6 suites and 7,140 items.

Later on September 24 Vicentino posted Eikos Chess. Eikos-27B and Jev each pick a legal move. The how-it-works page says neither model searches ahead. Both receive the same request: the board, and one question, “Which move is best?”, with every legal move as an option. Openings come from Lichess’s public list, each played twice with the colors swapped. Games end by the rules of chess, or as a draw at 200 half-moves. A missing answer is left out of the ranking. The score is wins plus half the draws.

His 22:54 UTC screenshot shows 388 games in 23 rounds. Eikos-27B has 42 wins, Jev has 12, and 334 games are draws. The page scores that as 53.9% for Eikos. The Elo gap is +27 for Eikos, with a 95% interval of +14 to +40. Eikos’s wins split 20 as White and 22 as Black. Jev’s split 6 and 6. Average answer time is 3.5 seconds for Eikos and 288 ms for Jev. Endings on that image: insufficient material 170, threefold repetition 164, checkmate 54. One finished round of 32 games on the same image is Eikos 1, draws 31, Jev 0, in 5 minutes 22 seconds. We opened the page later and the ranking widget was still loading, so these counts are the screenshot’s.

The same day he posted eikos-arena, MIT. Eikos-27B and Jev each start with $10,000 of paper money on 14 Hyperliquid perpetuals. Every five minutes both see the same snapshot and answer 28 questions: long, flat, or short on each market, and whether the price is higher in five minutes. Fills use Hyperliquid’s bid, ask, fees, and funding. No order goes to the exchange. The run lasts 72 hours, from 05:12 UTC on September 24 to 05:12 UTC on September 27. The page says profit over three days is mostly luck, and that the Brier score on the up-or-down question is the probability check. Always answering 50% scores 0.250. Full logs are added when the run ends.

The dataset file live/state.json, published at 06:20 UTC on September 25, is cycle 302 of 864, about 25 hours in. Equity is $9,922.72 for Eikos and $10,125.97 for Jev. Brier is 0.2771 on 4,214 Eikos forecasts and 0.2505 on 3,976 Jev forecasts. The directional hit rate is 49.5% and 44.9%. Median latency is 3,227 ms and 385 ms. The snapshot’s missed count is 0 for Eikos and 17 for Jev. Closed trades are 98 and 27. Fees are $54.15 and $12.77.

The same file, opened again on September 26, has a published timestamp of 09:20 UTC. It is cycle 626 of 864, about 52 hours in. Equity is $9,653.19 for Eikos and $10,051.22 for Jev. Brier is 0.2716 on 8,750 Eikos forecasts and 0.2500 on 8,470 Jev forecasts. The directional hit rate is 50.2% and 45.9%. Median latency is 3,369 ms and 381 ms. Missed is 0 for Eikos and 20 for Jev. The trades field is 135 and 31, and each side still has 13 open positions. Fees are $66.92 and $13.49. The run still ends at 05:12 UTC on September 27. We did not replay a chess game or a fill.

This site's reading

Editorial notes evaluating claims against primary sources, contextualizing findings alongside related implementations, and defining technical terms.

Verify

The September 24 post, 00:46 UTC, says the weights are open and links the repo. The September 23 post, 15:09 UTC, announced the same headline numbers before the link. The README table, labelled "our harness, same items for every system": JevBench public hard tier, Eikos-27B 82.9, Eikos-4B 72.1, Jev 73.0, Laya 35.1.

Error when confidence is at least 0.90: 2.4%, 2.4%, 5.9%, 37.3%, on 7,140 items and 6 suites. Accuracy with the decision hidden in 64k tokens, mean of 3 positions: 88.3 and 74.2. Jev's cell says the API limit is 32k.

Laya's cell says 512 to 1k context. The post says eight batteries. The README's confidence row says six suites. The hard-tier cells do not print a denominator. The 4B recipe starts from Qwen/Qwen3.5-4B, LoRA rank 64, one epoch, learning rate 1e-4, temperature 1.

On September 30, 2026 the git tree at main contained one JSON file, release_tools/known_label_errors.json, eleven TAT-QA ids. That file is a label list. It is not a dump of the suite table. Eight batteries and six suites remain two printed counts.

The 27B target is a separate train_final.sh argument; this page does not name that base. Decontamination is 8-gram against public JevBench items. The release omits 2% of the rows they trained on.

vLLM older than 0.30 dropped 3 to 6 points on long batched items in their note. Quantized builds ship if accuracy stays within 1 point, ECE within 0.01, and at least 97% of answers match bf16. Code is MIT.

Fine-tune deltas are MIT. Base weights are Apache 2.0. Data is CC BY 4.0. We did not run eval_vllm_suite.py. TypeSafe's Master Customer Agreement section 2.3(f) forbids customers from publishing benchmarks of the Services.

The same day we opened the Eikos-27B model card. It names the 27B base as Qwen/Qwen3.8-27B. The RuleArena NBA row, balanced accuracy, is 0.50 for the 4B, 0.50 for the 27B, and a blank cell for Jev. The limitations line says that on RuleArena, NBA salary-cap rulebooks of about 25k tokens, the model does not discriminate.

The release-builds heading says 7 suites and 7,371 items. The chart note says 7,140 items are common to all four systems. Those counts sit beside the README's 6 suites and 7,140 items.

We did not ask TypeSafe whether the Jev column is permitted. The September 24 chess screenshot at 22:54 UTC shows 388 games, Eikos 42 wins, Jev 12 wins, 334 draws, score 53.9%, Elo +27 with a 95% interval of +14 to +40, and answer times of 3.5 s and 288 ms.

Endings on that image: insufficient material 170, threefold repetition 164, checkmate 54. The how-it-works page says both models get every legal move, openings are played twice with colors swapped, and a missing answer is left out. We opened the page later and the ranking widget was still loading.

The arena snapshot live/state.json is timestamped 06:20 UTC on September 25, cycle 302 of 864. Equity is $9,922.72 and $10,125.97. Brier is 0.2771 on 4,214 forecasts and 0.2505 on 3,976.

Hit rate is 49.5% and 44.9%. Median latency is 3,227 ms and 385 ms. Missed is 0 and 17. The same file, opened again on September 26, has published time 09:20 UTC, cycle 626 of 864.

Equity is $9,653.19 and $10,051.22. Brier is 0.2716 on 8,750 forecasts and 0.2500 on 8,470. Hit rate is 50.2% and 45.9%. Median latency is 3,369 ms and 381 ms. Missed is 0 and 20. Trades are 135 and 31, with 13 open positions on each side.

Fees are $66.92 and $13.49. The run ends 05:12 UTC on September 27. We did not replay a game.

Compare

Florian's JevBench v1.4.1 scores the public 534 and a sealed 308 with a harmonic mean. This README scores a public hard tier on the author's runner and prints 73.0 for Jev, which matches the 81/111 figure Open-Jev published for that public hard slice (72.97%).

The README does not print 111, so the match is a reading, not a count we recomputed. Laya at 35.1 on this row is far from Laya's 70.1 composite on the older JevBench geometric mean. Those are different formulas and, here, a hard-tier cell.

The 64k probe has no Jev number to put beside 88.3.

Terms

0.90
The README's act-on-it line. Among decisions the system takes at confidence 0.90 or higher, error is 2.4% for Eikos-4B and Eikos-27B, 5.9% for Jev, and 37.3% for Laya, on 7,140 items across 6 suites.
letter logits
Eikos reads the probability of each option from the letter tokens, in one forward pass over the shared state. More than 26 options go through a tournament. The served route is POST /v1/systemone.
chess score
On the Eikos Chess page, wins plus half the draws. The September 24 screenshot scores 388 games as 53.9% for Eikos-27B.

Sources

  1. Caio Vicentino, weights are up
  2. Caio Vicentino, Eikos announcement
  3. caiovicentino/eikos
  4. caiovicentino1/eikos-decisions
  5. Caio Vicentino, Eikos Chess
  6. Caio Vicentino, 388 chess games
  7. Eikos Chess
  8. Caio Vicentino, Eikos Arena code
  9. caiovicentino/eikos-arena
  10. Arena snapshot, live/state.json
  11. Eikos-27B model card