Published
Updated
NIRNAY 450M prints 0.8792 on 3,080 Banking77 rows and says the Jev cell was not measured here
Gautam Kishore posted NIRNAY 450m on October 3, 2026, an Apache-2.0 Laya fine-tune from Eulogik. The phase_b JSON is 0.8792 on 3,080 Banking77 test rows, Brier 0.208, fitted ECE 0.045. The post's headline 0.8656 is the phase_a file, set against a Jev figure of 0.64. The card also prints Jev at 0.803, and the post says that column was not measured in the repo. The eval JSON has no Jev field.
Gautam Kishore posted NIRNAY 450m on October 3, 2026, at 07:10 UTC. The post says Banking77 is 0.8656 against a published Jev figure of 0.64, that the model was fine-tuned on the train split, and that Jev was zero-shot. He writes that this is Louis Grenard’s Laya thesis: a checkpoint you specialise on your own labels. His bio says he is CEO of Eulogik. The catalog page is decision models like Jev.
The headline 0.8656 is the phase_a file. eval/banking77_phase_a.json has accuracy 0.8655844155844156 on 3,080 rows. A reply one minute later prints ECE 0.089 raw, 0.045 fitted, and Brier 0.208 for phase_b. Those three numbers match eval/banking77_phase_b.json: accuracy 0.8792207792207792, Brier 0.20832792665271307, ECE raw 0.08908705256427277, ECE fitted 0.0454289691801037, n 3080. Neither file has a Jev field.
The model card’s table still prints Jev 1.13.0 at 0.803 on those 3,080 cases. A later sentence on the same card says the Jev figures were not measured in the repo, because the authors had no API access. The October 3 reply says the same thing in the other direction: they never measured Jev. The post’s 0.64 is not linked. Laya’s README prints a published Jev Banking77 of 0.870, 72 labels against 77. The NIRNAY card’s Julia-1 cell is 0.64 on a 72-label pilot of 100 items. The 0.64 in the post and the 0.64 on Julia’s row are the same digits, and the post calls the figure Jev’s.
The card’s speed line is 209 ms on an M4 GPU and 361 ms on CPU, batch 1, measured 2026-09-30. The reply lists 310 to 478 ms locally, slower than Laya’s 33 ms. The card places that 310 to 478 ms range on Jev API medians, not on the local clock. Context is 512 tokens. The base line is Laya 421M plus about 30M, 7,000 phase A steps and 50 RLCD steps on one Mac. The license line is Apache-2.0. The server line is POST /v1/systemone. The card prints a JevBench public slice of 231 items at 0.550 overall, easy 0.875, original 0.569, and hard 0.396, and it sets 0.55 beside Jev at 72.1. On this desk, 72.1 is the v1.5.0 composite. We did not download the checkpoint, and we did not call Jev for this comparison.
This site's reading
Editorial notes evaluating claims against primary sources, contextualizing findings alongside related implementations, and defining technical terms.
Verify
Gautam Kishore, @gautamkishore, October 3, 2026, 07:10 UTC, status 2106280618573389954. The post says NIRNAY 450m is an open-weight 450M decision model, and that Banking77 is 0.8656 against Jev's published 0.64. It says the checkpoint was fine-tuned on the train split and that Jev answered zero-shot. It points at Louis Grenard's Laya work. The post has no attached photo. We did not run the checkpoint.
A reply at 07:10:28 UTC, status 2106280707970798001, prints ECE 0.089 raw and 0.045 fitted, Brier 0.208, and 92% accuracy above 0.9 confidence for phase_b. A reply at 07:10:38 UTC, status 2106280750786158609, lists 512-token context, English banking, 310 to 478 ms locally, and a 466 MB checkpoint, and says the authors never measured Jev's API. A reply at 07:10:48 UTC, status 2106280794226655554, says the base is Laya 421M plus about 30M, trained for 7,000 phase A steps and 50 RLCD steps on one Mac, Apache-2.0, and links the model card. A separate post at 07:11 UTC, status 2106281006919733551, says the card reports raw and fitted ECE and ships per-bucket temperatures. Kishore's bio says he is CEO of Eulogik.
The model card is https://huggingface.co/eulogik/nirnay-450m. We opened it on October 4, 2026. The phase_b row is 0.8792 on 3,080 Banking77 test cases, Brier 0.208, ECE raw 0.089 and fitted 0.045. The phase_a row is 0.8656, Brier 0.240, ECE raw 0.110 and fitted 0.033. The card's Jev 1.13.0 cell is 0.803 on the same 3,080 cases, marked zero-shot, and a later sentence says Jev figures on that card were not measured in the repo. Julia-1 is 0.64 on that card's 72-label pilot, n=100. The speed line says 209 ms on M4 MPS and 361 ms on CPU, measured 2026-09-30, and it places 310 to 478 ms on Jev API medians. The card prints JevBench public, 231 items, at 0.550 overall, easy 0.875, original 0.569, and hard 0.396. Context is 512 tokens. The server line is POST /v1/systemone. The license line is Apache-2.0. We did not download the weights.
eval/banking77_phase_b.json, opened the same day, has accuracy 0.8792207792207792, brier 0.20832792665271307, ece_raw 0.08908705256427277, ece_fitted 0.0454289691801037, n 3080, and no Jev field. eval/banking77_phase_a.json has accuracy 0.8655844155844156, n 3080, and no Jev field. The post's 0.8656 matches the phase_a accuracy. The calibration reply's 0.089, 0.045, and 0.208 match phase_b.
TypeSafe's Master Customer Agreement section 2.3(f) forbids customers from publishing benchmarks of the Services. The Jev cells on this page are the ones the card and the post print. This desk did not call Jev for them.
Compare
The October 3 post prints 0.8656 against "Jev's published 0.64" and does not link the page that prints 0.64. The card's table prints Jev 1.13.0 at 0.803 on the same 3,080 rows, and the card also says that Jev column was not measured in the repo. The two JSON files have no Jev key. Laya's README, on this desk, prints a published Jev Banking77 of 0.870, with 72 labels against 77, and routed Laya at 0.425. The card's Julia-1 cell is 0.64 on a 72-label pilot of 100 items. Those are separate printed figures. The post calls 0.64 Jev's. The card calls 0.64 Julia-1's.
Charly Poly's Banking77 figure is macro-F1 0.782 against ModernBERT-large at 0.712. Cloudflare's BANKING77 row is macro-F1. NIRNAY's 0.8792 is accuracy on the authors' test file. The three pages are not one run.
The reply puts 310 to 478 ms on the local run. The card puts 209 ms and 361 ms on the local run, measured 2026-09-30, and puts 310 to 478 ms on Jev API medians. Laya's author timing on this desk is a different machine. JevBench's medians are a different task mix.
The card's 0.550 on 231 public items sits beside a Jev figure of 72.1 in the card's prose. On this desk, 72.1 is the JevBench v1.5.0 composite, not a fraction of 231. Open-Jev's 231-task slice is a different write-up, with the 9B at 179 of 231. The card's hard tier is 0.396. We did not check that the 231 rows are the same file.
Terms
- NIRNAY 450M
- Eulogik's Apache-2.0 decision checkpoint at eulogik/nirnay-450m. The card says it is a Laya 421M base plus about 30M, and that it serves POST /v1/systemone.
- phase_b
- The Banking77 JSON at eval/banking77_phase_b.json. Accuracy 0.8792, n 3080, Brier 0.208, ECE raw 0.089, ECE fitted 0.045. The file has no Jev field.