Published
Bryo ran 1,565 business emails: Jev at 96.4%, Gemini a point or two higher
Nikhil Mudholkar, CTO of Bryo AI, routed 1,565 German and English supplier emails into 10 categories. Jev scored 96.4% against Gemini 3.5 Flash-Lite at 97.5% and Gemini 3.8 Flash at 98.5%. Jev was about 10 and 22 times cheaper per thousand emails, and its high-confidence band had no errors on this set.
Nikhil Mudholkar, CTO of Bryo AI, posted on September 17 that Jev lost to Gemini on an email classification benchmark, and that he still wants to put it in production.
The task was routing business mail from industrial suppliers into 10 categories, including orders, quote requests, and invoice disputes. The set is 1,565 German and English emails: 1,201 real, 364 written by a model to cover rare categories. Mostly German. Same text and the same category instructions for three models.
Overall accuracy he reports:
- Jev: 96.4%
- Gemini 3.5 Flash-Lite: 97.5%
- Gemini 3.8 Flash: 98.5%
Gemini led by 1.1 to 2.1 points. Weighting each category equally widens that gap: 92.0% / 94.6% / 96.9%.
He wants to ship the confidence number. 737 Jev predictions at 99% confidence or higher all matched the reference label. Below 70%, nearly half were wrong. The most confident 85.5% of Jev’s answers had zero errors on this set. That, he writes, is how you decide which mail can be routed automatically and which mail needs a person.
Cost per 1,000 emails in his runs: Jev $0.08, Flash-Lite $0.80, Flash $1.79. He calls that roughly 10 times and 22 times cheaper. At ordinary inbox volumes the dollar gap is small; the confidence band is the reason he is still interested. Jev never took longer than 1.3 seconds.
Two limits sit at the end of the thread. Jev does not take non-text input, so attachments were excluded, which he calls the biggest blocker for production. And 364 of the labels came from synthetic mail written to match chosen categories.
This site's reading
Editorial notes evaluating claims against primary sources, contextualizing findings alongside related implementations, and defining technical terms.
Verify
The eight-post thread is the source. Mudholkar reports 1,565 emails (1,201 real, 364 AI-written for rare categories), mostly German, 10 routing labels, same instructions for three models. Overall accuracy: Jev 96.4%, Gemini 3.5 Flash-Lite 97.5%, Gemini 3.8 Flash 98.5%. Equal weight per category: 92.0% / 94.6% / 96.9%. He says 737 Jev predictions at 99% confidence or higher all matched the reference label, that nearly half of answers below 70% were wrong, and that the most confident 85.5% of Jev's answers had zero errors on this set. Cost per 1,000 emails in his runs: Jev $0.08, Flash-Lite $0.80, Flash $1.79. Latency: Jev never took longer than 1.3 seconds. We did not rerun the set. Attachments were excluded. 364 labels came from synthetic mail.
Compare
Ryan Vogel classified 1,500 personal emails and posted no comparator and no error rate. Every timed 777 writing judgments and later scored 12 passages against Claude Fable 5.1. This thread is the first named-company accuracy table on this desk with a frontier-model baseline, a confidence band, and a cost per thousand. TechCrunch later rounded the cost gap to 10 to 20 times and skipped the 96.4% figure. The 364 synthetic mails and the missing attachments are the limits he names for production.
Terms
- Reference-label accuracy
- Share of emails where the model's top category matches the label Mudholkar assigned. The 96.4% figure is this metric on 1,565 items, including 364 synthetic mails.
- Confidence band
- A slice of answers above a probability cutoff. On this set he reports zero errors in Jev's most confident 85.5%, and that nearly half of answers below 70% were wrong.
- Macro accuracy
- Accuracy with each of the 10 categories weighted equally. He reports 92.0% for Jev, 94.6% for Flash-Lite, and 96.9% for Flash.