Published
Updated
WaterSheep fine-tunes ModernBERT and serves /v1/systemone with a multi-label question
Samrat Dutta posted WaterSheep on October 3, 2026. The Apache-2.0 repo fine-tunes ModernBERT-base and serves POST /v1/systemone on port 8766, with noul, choice, score, and multi. The README prints 77.8% on an in-distribution split, ECE 0.026, and 61.2% on held-out datasets, ECE 0.043. Those two rows do not print a question count, and the page has no Jev column. We did not run it.
Samrat Dutta posted WaterSheep on October 3, 2026, at 18:29 UTC, as a reply in someone else’s thread. The post says it is a ModernBERT-base fine-tune that scores every option in one pass, serves Jev’s POST /v1/systemone, adds multi-label, and is Apache 2.0. The post has no photo. Other calls of this shape are on decision models like Jev.
The README says the server is watersheep --serve on port 8766, and that a TypeSafe client can point at it with TYPESAFE_BASE_URL=http://127.0.0.1:8766. Any API key value works on that local process. The README says the project is independent of TypeSafe. Question types are noul, choice, score, and multi. Multi returns every label above a threshold. The README does not print the default threshold. The license section says Apache 2.0. The Hugging Face card samratduttaofficial/WaterSheep says version 0.1.0 and checkpoint id watersheep-20260928-125452. The base line on both pages is answerdotai/ModernBERT-base.
The evaluation table prints 77.8% and ECE 0.026 on an in-distribution test split, and 61.2% and ECE 0.043 on held-out datasets. Those two rows do not print how many questions they cover. A second table is per dataset. GoEmotions is 22.4% of 2,000, ECE 0.023, marked as another split of data that was in training. HateCheck is 75.1% of 2,000, ECE 0.139, marked as not in training. LegalBench CUAD audit rights is 86.3% of 1,216, ECE 0.041, marked as not in training. Prompt injection is 91.4% of 116, ECE 0.079, marked as another split. There is no Jev column on either table.
Training is a decision head on ModernBERT-base, public datasets named in NOTICE, and synthetic decisions from Qwen3.5-4B, with one temperature per question type. The limitations say English only, long inputs are truncated, ratings are the weaker type, and the model is not for a medical, legal, financial, or hiring decision on its own. We did not install it, and we did not open the demo space. Charly Poly’s encoder bake-off is a different ModernBERT comparison, and it does have a Jev cell.
This site's reading
Editorial notes evaluating claims against primary sources, contextualizing findings alongside related implementations, and defining technical terms.
Verify
Samrat Dutta, @TheSamratDutta, October 3, 2026, 18:29 UTC, status 2106451707517649177. The post is a reply in another person's thread. It describes WaterSheep as a ModernBERT-base fine-tune that scores every option in one pass, serves POST /v1/systemone, adds multi-label, and is Apache 2.0. It links https://github.com/SamratDuttaOfficial/WaterSheep. The post has no attached photo. We did not run the server.
The README, opened October 4, 2026, says the model answers yes/no, single-choice, rating, and multi-label questions, with a probability for every option. The server command is watersheep --serve, and the printed port is 8766. The client line is TYPESAFE_BASE_URL=http://127.0.0.1:8766. The README says any local API key works, and that the project is independent of TypeSafe. The license section says Apache 2.0. The training section names answerdotai/ModernBERT-base plus a decision head, public datasets, and synthetic decisions from Qwen3.5-4B, with a temperature per question type.
The evaluation table prints 77.8% and ECE 0.026 on an in-distribution test split, and 61.2% and ECE 0.043 on held-out datasets. Neither row prints a question count. The benchmark table is a separate grid. Rows we read include goemotions 22.4% of 2,000, ECE 0.023, marked other split; hatecheck 75.1% of 2,000, ECE 0.139, marked no; legal CUAD audit rights 86.3% of 1,216, ECE 0.041, marked no; prompt injection 91.4% of 116, ECE 0.079, marked other split. The page has no Jev column.
The Hugging Face card samratduttaofficial/WaterSheep, opened the same day, says Apache-2.0 and version 0.1.0, checkpoint id watersheep-20260928-125452. The base model line is answerdotai/ModernBERT-base. The limitations on the README say English only, long inputs are truncated, rating answers are less accurate, and the model is not for a high-stakes decision on its own. We did not open the demo space, and we did not install the package.
Compare
Jev's three question types are Choice, Score, and Noul. WaterSheep's README adds a fourth, type multi, for every label above a threshold. The README does not print that threshold's default.
The README has no Jev column, so this page does not invent one. Charly Poly's Banking77 run is macro-F1 for hosted Jev against ModernBERT-large-zeroshot, 0.782 against 0.712. Laya is another encoder fine-tune, and its Banking77 cell is 0.425 against a published Jev 0.870. WaterSheep's 77.8% and 61.2% do not say which datasets, and they are not that Banking77 pair.
Ollaya also serves POST /v1/systemone, on port 11435, and its README does print a Jev cell. The ports are different processes. NIRNAY's card also says POST /v1/systemone, from a Laya base, with a Banking77 file this desk opened. WaterSheep's evaluation table is the one without a Jev key.
Terms
- WaterSheep
- Samrat Dutta's Apache-2.0 fine-tune of ModernBERT-base. The README serves POST /v1/systemone on port 8766 and adds a multi-label question type.
- 0.1.0
- The version string on the Hugging Face card samratduttaofficial/WaterSheep, checkpoint id watersheep-20260928-125452.