Published
Updated
JEV-27B-VL finishes 15 of 20 pick-and-place scenes
AutoTrust posted a JEV-27B-VL follow-up on October 4, 2026. The model card's October 3 section prints 15 of 20 pick-and-place scenes, and 95% of 60 browser tasks when the element text is included. A six-benchmark mean from September 27 is 84.07 against a hosted Jev cell of 83.85. That JevBench column is not the official board.
AutoTrust posted a follow-up on October 4, 2026, at 10:28 UTC. The model is JEV-27B-VL, an Apache-2.0 adapter on Qwen3.8-27B. The October 1 clip introduced it. System 1 is POST /v1/decide. A request can be yes/no, a choice of 2 to 256 options, or a score from 0 to 5, and the state can include images. The reply is a probability for every option. System 2 is the unmodified Qwen checkpoint, and that path writes text. The card’s serve script asks for one GPU with 80 GB. Other calls of this shape are on decision models like Jev.
The card’s section dated October 3 describes two control loops. In the robot loop, each step looks at a top camera and answers whether the target is left or right of the gripper, and whether it is above or below. The arm halves its step when an answer flips. The card prints 75% of 20 random scenes completed. Every cube that was grasped ended in the tray, 15 of 15. Each miss was a grasp 3.0 to 3.7 cm off the cube. Time per decision is about 240 ms in the prose and 239 ms in the table. Asked to pick one of 8 motor commands directly, the same section says the model completed 0 of 10 scenes.
The browser loop draws a numbered box on each clickable element. With the element text included, the card prints 95% of 60 tasks completed, 3 to 7 clicks each, about 0.26 s per click. With numbered boxes and no element text, the same model is at 10%. The note says it often declares the task complete too early. The colour-swatch click is the example that has to be decided from the screenshot alone.
| JEV-27B-VL | JEV-9B | GEV-26B-Decide | |
|---|---|---|---|
| Pick and place, 20 scenes | 75% | 50% | 40% |
| Browser, boxes plus element text, 60 tasks | 95% | 95% | 95% |
| Browser, numbered boxes only | 10% | 37% | 15% |
| Time per robot-arm decision | 239 ms | 163 ms | 61 ms |
The October 4 post names RoboCasa kitchen tasks, a drawer, a microwave, a cup moved to a sink, and Space Invaders at 40 of 40 with no damage. The October 3 section we opened does not print those lines. The card points demo code at the JEV-9B tree under vl/demos, and other write-ups at yuhai-china/JEV-27B-DEMO. We did not open that repository.
A second table on the card is six public sets, scored in percent. The card says AutoTrust ran JEV-27B and hosted Jev 1.13 on September 27, 2026, and that the remaining rows are copied from the NeoHorse-Jev-4B card. Text decisions are attributed to JEV-27B. The card says an image check on 1,000 answer-checking items moved the mean probability by 0.010, with 99.8% on the same side of 0.5.
| Model | JevBench | Kev | OpenJev text | Nimble | VitaminC | MASSIVE-en | Mean |
|---|---|---|---|---|---|---|---|
| JEV-27B / JEV-27B-VL | 88.70 | 83.75 | 73.89 | 92.91 | 77.46 | 87.71 | 84.07 |
| Jev 1.13, hosted | 87.18 | 85.52 | 72.96 | 91.84 | 78.46 | 87.14 | 83.85 |
Hosted Jev is ahead on the Kev column and on VitaminC. The six-group mean differs by 0.22 points. The column headed JevBench is this September 27 run. It is not Florian’s board. The page opened October 4 still describes v1.5.6, and the Jev cell this desk records from v1.5.0 is 72.1.
On 800 human-labelled items, with 2, 4, 8, and 16 options, the card prints hosted Jev at 0.890, 0.801, 0.782, and 0.769. JEV-27B is 0.876, 0.784, 0.767, and 0.740. A separate calibration block, on the held-out test_set_30k of jev-distill-corpus-v3, prints ECE 0.0009 and a KL to hosted Jev of about 0.017. That set is the one used to match Jev’s probabilities. The card also prints recommendation, agent-judge, and reward tables. This page does not transcribe them. We did not run the model.
This site's reading
Editorial notes evaluating claims against primary sources, contextualizing findings alongside related implementations, and defining technical terms.
Verify
AutoTrust, @AutoTrustAI, October 4, 2026, 10:28 UTC, status 2106692811542425880. The post names RoboCasa kitchen tasks, a robosuite arm, MuJoCo, a drawer, a microwave, a cup moved to a sink, and Space Invaders at 40 of 40 with no damage. The attached file is a video, so this page has no still. The October 1 post, status 2105476752449679686, is the introduction clip.
The model card was opened October 5. Its section dated October 3 prints the pick-and-place table and the browser table below. That section does not print Space Invaders, and it does not print the drawer, microwave, or cup tasks. The six-benchmark table says AutoTrust ran JEV-27B and hosted Jev 1.13 on September 27, 2026, and that the other rows are copied from the NeoHorse-Jev-4B card. We did not run the model, and we did not open the demo repository the card names.
Compare
On the six-benchmark mean the gap is 84.07 to 83.85. Hosted Jev is ahead on the Kev column, 85.52 to 83.75, and on VitaminC, 78.46 to 77.46. The card's JevBench column, 88.70 against 87.18, is this September 27 run. Florian's board, opened October 4, is v1.5.6, and the Jev cell recorded there is 72.1 on v1.5.0. Those are different ledgers.
The pressure table is 800 human-labelled items. Hosted Jev's published cells are 0.890, 0.801, 0.782, and 0.769 at 2, 4, 8, and 16 options. JEV-27B is 0.876, 0.784, 0.767, and 0.740. The ECE of 0.0009 is on jev-distill-corpus-v3, the set used to match Jev's probabilities, not on the six public benchmarks.
Terms
- JEV-27B-VL
- AutoTrust's Apache-2.0 adapter on Qwen3.8-27B. System 1 is POST /v1/decide: yes/no, a choice of 2 to 256 options, or a score from 0 to 5, over text or images. System 2 is the unmodified Qwen checkpoint writing text.
- 15 of 20
- Pick-and-place scenes completed in the card's October 3 section. The card says every cube that was grasped ended in the tray, 15 of 15, and that each miss was a grasp 3.0 to 3.7 cm off the cube.
- 84.07
- Six-group mean on the card's September 27 table. The hosted Jev cell on that table is 83.85. The JevBench column inside the mean is not Florian's official board.