One post scores Jev at 100 and Laya's English checkpoint at 92, on 100 items

码农暖爸 posted a 100-item English judgment set. Laya's English checkpoint is 92 correct and 8 wrong. Jev is 100 correct. The post names four scene types and says the set is not a full benchmark. The item list is a reply image.

码农暖爸 posted on September 23 a second look at Laya’s English checkpoint against Jev. The set is 100 English judgment items. Laya is scored 92 correct and 8 wrong. Jev is scored 100 correct. The post says that is 8 points on this set, and that the set is not a full benchmark. It is meant to show a gap on items the author picked because small models miss them.

The post names four scenes and gives one example of each. A feature description does not say the name: “The sweet red fruit with a leafy green cap on top.” A negation tells the model what to skip: “Skip the grapes and citrus, give me the other one.” The sentence mentions grapes and asks for the orange: “The recipe picture has grapes, but the fruit to serve is orange.” Padding wraps the request: “Hey there, I think I’m done with berries, please pass the purple bunch on the counter.” The full list is in the reply’s report image. We did not transcribe that image into per-item scores.

An earlier post the same day, at 01:15 UTC, says Laya is fast because it runs locally, so a speed comparison with a cloud call to Jev mixes network time into the clock. That post says Jev looked more accurate on the author’s round and does not print a count. Neither post links a repository. We did not rerun the 100 items.

Laya’s published Banking77 table is a different denominator: macro F1 0.425 against a published Jev figure of 0.870, 77 labels against 72, and the README says that Jev number was not measured in the Laya repository. laya-mlx checks whether an Apple Silicon port matches upstream Laya on 63 questions. It does not score hosted Jev. The 100-item post is the author’s own items, with the 92 and the 100 printed in the text.

This site's reading

Editorial notes evaluating claims against primary sources, contextualizing findings alongside related implementations, and defining technical terms.

Verify

Primary post is @Delroy715, display name 码农暖爸, September 23, 2026, 07:49 UTC. Chinese text, with a video. It says the set is 100 English judgment items, Laya's English checkpoint 92 correct and 8 wrong, Jev 100 correct, and that this is not a full benchmark.

Four scene names, each with one English example: feature alignment ("The sweet red fruit with a leafy green cap on top."), negation ("Skip the grapes and citrus, give me the other one."), mention versus intent ("The recipe picture has grapes, but the fruit to serve is orange."), and conversational padding ("Hey there, I think I'm done with berries, please pass the purple bunch on the counter.").

The next post attaches a report image. An earlier post the same day, 01:15 UTC, says Laya is local so a speed comparison with cloud Jev is uneven, and prints no count. No repository.

We did not transcribe the image into per-item scores and we did not rerun the 100 items.

Compare

Laya's own Banking77 table is macro F1 0.425 against a published Jev figure of 0.870, on 77 labels versus 72, and the README says that Jev number was not measured in the Laya repo. This post is 92/100 against 100/100 on the author's items.

laya-mlx matches upstream Laya on 63 validation questions and does not score hosted Jev. The denominators are different.

Terms

100 English judgments
The count in the September 23 post. Laya's English checkpoint is scored 92 correct and 8 wrong. Jev is scored 100 correct. The author says the set is not a full benchmark.
four scene types
Feature descriptions, negation, mention versus the intended target, and conversational padding. The post gives one English example of each. The item list is the reply's report image.

Sources

  1. 码农暖爸, 100 judgments
  2. Earlier same-day note