fsdecide scores 91.3% on ForkScape review, and Jev scores 56.3%

ForkScape posted fsdecide on October 5, 2026, a 149M ModernBERT trained on the app's own decisions. On 80 review items it scores 91.3% against jev-latest at 56.3%. On the authors' 1,248-request general set, jev-latest is 90.5% and fsdecide is 60.5%. That set is not Florian's JevBench.

ForkScape posted a comparison on October 5, 2026, at 03:26 UTC. The model is fsdecide, a 149M ModernBERT-base encoder. The comparison doc says it was trained on about 34,000 labelled decisions from five questions the app already asks, and that anything close to a bench item was dropped from training. The repository holds 450 hand-labelled items. The README says the license is MIT. The local server is fsdecide/serve.py on port 8790, and the path is POST /v1/systemone. Other calls of this shape are on decision models like Jev.

The stored report is dated 2026-10-05. Grading is code against the hand labels. jev-latest and jev-preview went out over the API. fsdecide is the row named sys1:base-v11e4s21. No suite records an abstention.

Suitenfsdecidejev-latest
Triage8097.5% (91.3% to 99.3%)92.5% (84.6% to 96.5%)
Review8091.3% (83.0% to 95.7%)56.3% (45.3% to 66.6%)
Browse16088.8% (82.9% to 92.8%)81.3% (74.5% to 86.5%)
Router7095.7% (88.1% to 98.5%)100% (94.8% to 100%)
Command intent60100% (94.0% to 100%)78.3% (66.4% to 86.9%)

The doc says the review gap and the command-intent gap sit outside those intervals, and the browse gap does not. Triage and router overlap as well. On review, jev-latest’s confusion list is pass called fail, 35 times. The same section says unbacked claims were 15 of 15. At a confidence gate of 0.8, jev-latest’s review accuracy is 88.5% on 26 of the 80 items. fsdecide’s is 100% on 66 of 80.

Browse is scored a second way. An unsafe allow is a payment, destructive, or auth action called none or send. An over-ask is a harmless action called hard. jev-latest has 0 unsafe allows out of 85 and 11 over-asks out of 40. fsdecide has 1 of 85 and 2 of 40. The doc describes that one allow as typing into a card-number field on a pay page, called harmless.

jev-latest’s run is 282,702 input tokens and $0.0119. The summary page rounds the bill to $0.012. p50 in the report is 158 ms to 167 ms for jev-latest. fsdecide’s p50 is 50 ms to 56 ms on four suites and 146 ms on command intent. The summary page prints 44 ms to 202 ms for the laptop row. Those two ranges are not the same column. jev-preview costs the same $0.0119 and stays within about a point of jev-latest on every suite. Its review row is 57.5%, with 34 pass-to-fail confusions.

The same doc runs both models on a file it calls JevBench-mini: 1,248 frozen requests, 240 cases, 12 domains. It says the package is Apache-2.0. This is not the JevBench board maintained by Florian. The picture on the post labels this file as a public leaderboard.

fsdecidejev-latestOpen-Jev-9BUniform guess
Coherence65.5 (63 to 68)82.3 (80 to 85)83.989.4
Accuracy60.5%90.5%88.4%39.8%

The doc says coherence is not a quality score on its own, because a model that ignores its input scores 89.4. fsdecide’s weak rows in that write-up are negation, where a yes/no question and its negation sum to one in 3% of tests, the same question asked as yes/no or as a choice at 17%, and non-English input at 21%. Batch independence is 100. The post’s 91%, 56%, and 78% are the review and command-intent cells rounded. The 90.5% and 60.5% are this general file. We did not rerun either set.

This site's reading

Editorial notes evaluating claims against primary sources, contextualizing findings alongside related implementations, and defining technical terms.

Verify

ForkScape, @forkscape, October 5, 2026, 03:26 UTC, status 2106949207466844212. The post prints review 91% against 56%, command intent 100% against 78%, and a general set at 90.5% against 60.5%. The photo on that post is the share image. It draws review at 56.3%, triage at 97.5%, command intent at 100%, and a panel it labels JevBench with Jev at 90.5% on 1,248 general decisions. The speech bubble about 50% plus 50% equaling 103% is a drawing.

The comparison doc, the stored report, and the repository README were opened the same day. The report's review row is 91.3% (83.0% to 95.7%) for sys1:base-v11e4s21 and 56.3% (45.3% to 66.6%) for jev-latest, 80 items, no abstentions. The README says MIT. We did not run decision-bench, and we did not open the release tarball named in the doc.

Compare

jev-preview, in the same report, is within about one point of jev-latest on every suite. Review is 57.5% (46.6% to 67.7%), and the confusion count is 34 pass-to-fail instead of 35. Both Jev rows have 0 unsafe allows out of 85 and 11 over-asks out of 40.

The 1,248-request file is the doc's JevBench-mini. The picture calls it a public leaderboard. Florian's JevBench is a different ledger. The doc says a model that ignores its input scores 89.4 on the coherence column, above jev-latest at 82.3. The measured negation line in that doc is 3% of tests, which is the rate at which a yes/no question and its negation sum to one.

Terms

fsdecide
ForkScape's 149M ModernBERT-base encoder, release fsdecide-base-v11e4s21. The comparison doc says it was trained on about 34,000 labelled decisions from five app questions, and that the 450 bench items were kept out of training. The local server speaks POST /v1/systemone.
91.3%
fsdecide on the 80-item review suite in the October 5, 2026 report. jev-latest on the same items is 56.3%. The interval is 83.0% to 95.7% against 45.3% to 66.6%.
90.5%
jev-latest accuracy on the authors' 1,248-request general set. fsdecide on that file is 60.5%. The doc says the set has 240 cases in 12 domains. It is not Florian's JevBench.

Sources

  1. ForkScape, October 5
  2. dreamerron/decision-bench
  3. fsdecide versus Jev
  4. Run report, October 5