Published
Updated
JevAdvBench's abstract says one appended opinion flips 12.1% of jev-1.13.0 decisions
This topic was created 10 days ago, and the information it contains may have evolved or changed since then.
Jianyi Hu and coauthors submitted arXiv:2609.31142 on September 25. The abstract describes 812 typed questions over 66 scenarios and 9,744 single-edit variants, run on jev-1.13.0. Rewording stays within 1.2 percentage points of a re-run baseline. One unverified opinion appended to the state flips 12.1% of decisions. We read the abstract and a September 28 post, not the 19 tables.
Isao Takaesu posted on September 28 a pointer at JevAdvBench. The paper is arXiv:2609.31142, version 1, submitted September 25. The authors on the abstract page are Jianyi Hu, Hangtao Zhang, Yi Liu, Yeqi Zeng, Li Zeng, Xianlong Wang, Rui Wang, and Leo Yu Zhang. The license line is the arXiv.org perpetual, non-exclusive license 1.0. The page lists 33 pages, 13 figures, and 19 tables. We read the abstract and the post. We did not transcribe the tables, and we did not open JevAdvBench.github.io/JevAdvBench. The post has no photo we saved.
The abstract says the suite is 812 typed questions over 66 scenarios, plus 9,744 single-edit variants, run on jev-1.13.0. Rewording stays within 1.2 percentage points of the re-run baseline. Fields that sit outside the schema never reach the model. One unverified opinion appended to the state flips 12.1% of decisions. The abstract says that rate is statistically tied with the strongest injected command, which it puts at 10.1%. The same condition pushes 38% of confident answers below the 0.8 confidence threshold.
This is a different paper from decision-hijack, arXiv:2609.28613, which reports validated attack success on 510 InjecAgent cases moving from 1.8% to 3.5%. The 12.1% figure here is the abstract’s appended-opinion rate on its own 812 questions.
This site's reading
Editorial notes evaluating claims against primary sources, contextualizing findings alongside related implementations, and defining technical terms.
Verify
The post is @bbr_bbq, September 28, 2026, 11:29 UTC. It points at the paper and has no photo we saved. The abstract page is arXiv:2609.31142v1. Authors listed there are Jianyi Hu, Hangtao Zhang, Yi Liu, Yeqi Zeng, Li Zeng, Xianlong Wang, Rui Wang, and Leo Yu Zhang. The paper was submitted 2026-09-25. The license on the abstract page is arXiv.org perpetual, non-exclusive license 1.0, not CC BY. The abstract page says 33 pages, 13 figures, and 19 tables. We did not transcribe those tables. The project site https://JevAdvBench.github.io/JevAdvBench/ was not opened.
Abstract figures we are using: 812 typed questions over 66 scenarios; 9,744 single-edit variants; the model is jev-1.13.0. Rewording stays within 1.2 percentage points of the re-run baseline. Fields outside the schema never reach the model. One unverified opinion appended to the state flips 12.1% of decisions. The abstract says that rate is statistically tied with the strongest injected command, at 10.1%. The same condition pushes 38% of confident answers below the 0.8 confidence threshold.
Compare
decision-hijack, arXiv 2609.28613, is a different paper. On 510 InjecAgent cases it moves validated attack success on jev-1.13.0 from 1.8% to 3.5%. The 12.1% here is an appended opinion on this abstract's 812 questions.
The judge-cascade paper, arXiv 2609.26550, measures JEV as a grader of other models' answers. It does not report this 12.1% flip.
Terms
- 12.1%
- The abstract's rate at which one unverified opinion, appended to the state, flips a jev-1.13.0 decision. The same sentence says this is statistically tied with the strongest injected command at 10.1%.
- 812
- Typed questions in the abstract, across 66 scenarios. The abstract also counts 9,744 single-edit variants. We did not read the 19 tables behind those counts.
- 1.2 percentage points
- How far rewording stays from the abstract's re-run baseline on jev-1.13.0. Fields outside the schema, the abstract says, never reach the model.