Published
A 612-row banking-feed test puts Jev under two Gemini rows on accuracy
Peter posted a categorization run on 612 sanitized rows from a private banking-feed set of about 33,000. Jev 1.13 scored 36.1% in 6 seconds at $0.023. Gemini 2.5 Flash scored 37.3% in 95 seconds. Gemini 3.7 Flash scored 42.5% in 78 seconds. The post names no repository and no category list.
Peter (@LordMarket22) posted a categorization test on September 22. The task is banking-feed labels from a product he says is already in use. The slice is 612 rows, described as a sanitized and harder subset of an evaluation set of about 33,000 rows.
Gemini 2.5 Flash is about $0.088, 95 seconds, 37.3%. Gemini 3.7 Flash is about $0.180, 78 seconds, 42.5%. Jev 1.13 is $0.023, 6 seconds, 36.1%.
Gemini 3.7 Flash is the high score. Jev is close to Gemini 2.5 Flash on accuracy and far from it on the clock. The post’s summary is about 4 times cheaper and 16 times faster than that Gemini 2.5 run. Dividing the printed figures gives about 3.8 times on cost and about 15.8 times on time. The asterisk says the Gemini dollar amounts are estimates from the same text-token budget at public rates, and that thinking tokens are left out. Jev’s $0.023 is printed without that asterisk.
The post does not list the categories, say who labelled the rows, or link a repository. A 15-second video is attached. We did not transcribe it. One reply calls the test useful and adds no number. We did not rerun the rows.
Banking77 is a different set. Poly’s macro-F1 for Jev was 0.782 against a zero-shot ModernBERT, and Kumar’s 300-item slice had Jev at 76.0%, behind gpt-5.4-mini and Luna. Those accuracies sit in the seventies and eighties because the label space and the items are the public intent set. Peter’s figures sit in the thirties and forties on a private feed, with Gemini 3.7 Flash six points above Jev and Gemini 2.5 Flash about one point above it.
This site's reading
Editorial notes evaluating claims against primary sources, contextualizing findings alongside related implementations, and defining technical terms.
Verify
Primary post is @LordMarket22, display name Peter, September 22, 2026, 12:30 UTC. The post says the task is banking-feed categorization from a product already serving customers, on a sanitized and harder subset of an evaluation set of about 33,000 rows. The printed table is 612 rows. Gemini 2.5 Flash about $0.088, 95 seconds, 37.3%. Gemini 3.7 Flash about $0.180, 78 seconds, 42.5%. Jev 1.13 $0.023, 6 seconds, 36.1%. The asterisk says Gemini costs use the same text-token budget at public rates and exclude thinking tokens. The post says Jev was roughly 4 times cheaper and 16 times faster than the Gemini 2.5 run. 0.088 divided by 0.023 is about 3.8, and 95 divided by 6 is about 15.8. No category list, no label source, no repository, and no row file appear in the post. A reply adds no figures. The post includes a 15-second video. We did not transcribe it. We did not rerun the 612 rows.
Compare
Banking77 on this desk is a public 77-way intent set. Kumar's 300-item slice had Jev at 76.0% against gpt-5.4-mini at 78.7% and Luna at 81.7%. Poly's macro-F1 was 0.782 against ModernBERT-large at 0.712. Bryo's supplier-mail table is 1,565 rows and a Gemini baseline on a different label scheme. Peter's 612 rows are a private feed, accuracy is in the mid-30s for Jev and Gemini 2.5, and the higher Gemini row is 42.5%. The sets do not share categories.
Terms
- banking feed
- In this post, rows from a product that categorizes bank transactions. The 612 rows are described as a sanitized, harder subset of an evaluation set of about 33,000. The categories are not listed.