Published
Updated
Step-Jev scores 19.6% when the reward grades every search step
On October 9, 2026 hosted jev-1.13.0 scored the repository's no-step test state incorrect at probability 1.0. The five-step fixture in the same test was correct at 0.05. On a Moon-landing question written here, a no-search answer of 1965 was incorrect at 1.0, and a repeated NASA search stayed positive at 0.53. real_jev_gate.py's rollout files are not in the repository.
laxman (@llmluthor) posted Step-Jev on October 8, 2026, at 17:20 UTC. The code is bandr-ai/bandits, main at 9d6f5dcdbf00. That commit merges pull request 142, feat/jev-step-reward, at 16:32 UTC the same day. The GitHub license field was empty when we opened the repository on October 9.
The write-up in step-rl-results.md is dated October 7 and marked as one seed per arm. The policy is Qwen3-4B-Instruct-2507, revision cdbee75f, LoRA rank 32, served by vLLM. Each turn may call search, open, or find once. A trajectory stops at 12 turns, and the last turn has to answer. The corpus is BrowseComp-Plus with its Qwen3-Embedding-8B index. The 830 questions are hash-split into 415 for training and 415 held out. Evaluation is 60 of the held-out questions, eight rollouts each. Nothing held out is used to train or tune the judges. Accuracy is a strict normalized string match against the gold answer. The note calls that a proxy, not the official LLM grader.
Training is Dr. GRPO with no KL term. One update is a batch of 16 training questions, eight rollouts each, learning rate 1e-4. Each arm sees the same questions in the same order. The note says the study used about $50 to $55 of Modal compute. group_advantages subtracts the group mean and does not divide by the group’s spread.
The post’s four cells, and the same cells in the note:
| Arm | What is rewarded | Accuracy | Tool calls |
|---|---|---|---|
| No RL | nothing | 15.6% | 3.6 |
| Outcome, ground truth | string match | 22.9% | 4.1 |
| Jev final | outcome Jev, P(correct) | 5.4% | 10.5 |
| Jev final + step | outcome Jev, plus step Jev | 19.6% | 6.5 |
Jev final stopped at 8 of 10 updates because of a 60-minute cap. The note says its accuracy fell on every update. Repeats on that arm are 15%. Jev final plus step is +14.2 points against Jev final, 95% interval +8.1 to +20.4, and -3.3 points against the string-match key, interval -8.3 to +1.5. Evidence seen on the step arm is 0.318 against 0.200 for the string-match key.
group_rewards in bandits_jev/grpo.py is the branch. For jev_final and jev_final_step the outcome term is jev_outcome, the outcome judge’s P(correct), not the gold label. jev_final stops there. jev_final_step also builds a StepEvent for each tool call and passes the list to shape_rollout. A repeated action, matched on tool name plus case-folded text, is stored as -repeat_penalty before shaping. The default penalty is 0.5. shape_rollout then clamps a repeat to at most 0, so a repeat can lower the return and cannot raise it. reward_vector caps the positive step mass at step_cap and the negative step mass at the same magnitude, separately. The note says the runs used step weight 0.3 and cap 1, and that those values were not tuned.
A rollout with no tool events never enters the step loop. shape_rollout then adds nothing, and the return is the outcome probability alone. judged_step_quality does return -1 when a rollout never searched, and the step2 arm uses that function. jev_final_step does not. The note’s hosted arm matches that path: the policy answered at once, the step judge had nothing to score, and the outcome judge still paid about 0.70 for answers that were all wrong. Accuracy on that arm is 0%, with 0.0 tool calls. Hosted Jev as both the final judge and the step judge is also 0% accuracy, 0.3 tool calls, and evidence seen 0.046.
The small judges are LoRAs on Qwen3.5-4B-Base. step_judge.py asks “What did this observed search action accomplish?” and offers positive, neutral, and negative. The training score is P(positive) minus P(negative), from one forward pass over letter logits. outcome_judge.py asks “Is the agent’s final answer to the research question correct?” and keeps the actions plus the last three tool results. The note prints step Jev v1 at AUC 0.78, and 78 of 100 on a fresh audit. Retraining on gold-evidence labels from their own rollouts, v2, prints AUC 0.786. The outcome Jev prints AUC 0.92 on 461 held-out answers, trained on 591 of 777 labelled finals from the training pool. Judge inputs do not contain the gold answer.
Hosted Jev is a different client. RemoteJev.choose posts one Choice to https://api.typesafe.ai/v1/systemone with model jev-latest. The step arm sends the same state, question, and options as the small step judge. The final arm sends full_outcome_state, which keeps every step and cuts results so the state stays near 80,000 characters. The note prints hosted Jev, zero-shot, at step AUC 0.891 and final-answer AUC 0.751. The step baseline used in training is 0.20 for the small step Jev and -0.17 for the hosted one. Adapters from the study are written to the Modal volume jev-runs, under /runs/step-rl/grpo/. We did not find a Hugging Face repository for those weights.
Three more arms sit on the same page and use the gold string match as the outcome, so they are not the no-answer-key result. Summing step-Jev scores, step v1, prints 20.6% accuracy, 7.4 tool calls, and 26% repeats. Averaging them only inside groups where every rollout is wrong, step v2, prints 15.8% and 2.8 calls. An oracle that adds the share of gold evidence documents the rollout saw prints 17.3% accuracy and evidence seen 0.350. The note says a sum rewards extra searches and a mean rewards stopping early.
The hosted check in the note is python recipes/jev/scripts/real_jev_gate.py. It needs JEV_API_KEY or TYPESAFE_API_KEY, then scores 480 held-out rollouts in work/step-rl/rollouts/base.json and 1,172 step rows in judge-v2-heldout.jsonl. Questions come from tasks.jsonl. .gitignore lists work/. On October 9 the three raw URLs on main returned HTTP 404. The script cannot score those rollouts from the public tree, so this desk did not recompute hosted AUC 0.751, step AUC 0.891, or the claim that instant wrong answers were rated about 70% likely correct.
The states a reader can rebuild are the ones the tests assert. On October 9 this desk sent them as RemoteJev does: one Choice, model jev-latest, to https://api.typesafe.ai/v1/systemone. The response model was jev-1.13.0.
test_jev_remote.py freezes a rollout with no steps and no answer: “Research question: Q?” and “Final answer: (no answer)”. Jev called it incorrect at 1.0, confidence 1.0, 355 input tokens and 31 output tokens, desk clock 466 ms. The same file freezes five browser.search steps whose observations are “r0” through “r4”, and the final text “Answer: Paris”. Jev called that incorrect at 0.95, correct 0.05, confidence 0.9, 574 and 31, 419 ms. Placeholder searches did not make the named answer look supported.
test_step_judge.py freezes one search, action ”{}”, observation “reply”. Jev called it neutral at 0.79, positive 0.04, negative 0.17, confidence 0.69, 404 input tokens and 38 output tokens, 645 ms. The score the trainer uses, P(positive) minus P(negative), is -0.13. A result with no source in it did not count as progress.
The same builders were then used on a question written here: “In what year did a human first walk on the Moon?” No search, and the answer 1965, was incorrect at 1.0, confidence 1.0, 366 input tokens and 31 output tokens, 454 ms. One NASA result, and the answer 1969, was correct at 1.0, confidence 1.0, 414 and 31, 430 ms. The first search that returned the Apollo 11 line was positive at 1.0, confidence 1.0, 451 input tokens and 38 output tokens, 436 ms. Repeating that search left positive at 0.53 and negative at 0.47, neutral at 0.0, confidence 0.3, 499 and 38, 423 ms. The step score falls from 1.0 to 0.06.
An empty or unsupported final answer was P(correct) 0.0, both on the test string “(no answer)” and on 1965 with no search. A final answer was P(correct) 1.0 when the state contained the NASA line for 1969, and 0.05 when the searches were the placeholders “r0” through “r4”. A step was positive at 1.0 when the observation named the Apollo 11 page, and neutral when the observation was the word “reply”. The soft case was the repeat: the label stayed positive, at 0.53, with confidence 0.3. jev_final_step would still clamp that repeat to at most 0, so the leftover 0.06 would not be paid. A step reward in that arm also never starts when the event list is empty.
The +14.2 point gain, and the hosted arm at 0% with answers rated about 70% likely correct, remain one seed whose rollouts are not in the public tree. On the states this desk could actually send, jev-1.13.0 paid for a sourced year and refused an unsupported one.
This site's reading
Editorial notes evaluating claims against primary sources, contextualizing findings alongside related implementations, and defining technical terms.
Verify
Primary post is @llmluthor, display name laxman, October 8, 2026, 17:20 UTC, status 2108246054500258256. The post names Step-Jev and says step rewards fixed a search agent that learned to fool a final-answer judge. The attachment is a video, about 41 seconds. The still on this page is the chart on reply 2108246058858054100.
That reply prints four cells for a Qwen3-4B search agent on BrowseComp-Plus, Dr. GRPO, one GPU, about $50: no training 15.6%, judge on the final answer 5.4%, judge on the final answer and every step 19.6%, real answer key 22.9%. Reply 2108246064218435863 says the hacked agent went from 3.6 to 10.5 searches a question, 15% of them repeats, and that adding step rewards gave +14.2 points, 95% interval +8.1 to +20.4. Reply 2108246068542816540 names a step Jev and a final Jev with AUC 0.92, and says a hosted judge was hacked in about three updates by answering with no search. Those answers were rated about 70% likely correct and were all wrong. Reply 2108250202775969827, 17:36 UTC, links github.com/bandr-ai/bandits.
We read main at 9d6f5dcdbf00, committed October 8, 2026, 16:32 UTC, message "Merge pull request #142 from bandr-ai/feat/jev-step-reward". The GitHub license field on that repository was empty. The numbers in the body are from recipes/jev/docs/step-rl-results.md, status date 2026-10-07, one seed per arm. The reward branches are group_rewards in bandits_jev/grpo.py. Question strings are step_judge.QUESTION and outcome_judge.QUESTION. The hosted client is RemoteJev in jev_remote.py, which posts to https://api.typesafe.ai/v1/systemone. real_jev_gate.py reads a local work/step-rl directory that is not in the repository. We did not run that script, and we did not train the LoRAs.
The hosted check named in step-rl-results.md is `JEV_API_KEY=... python recipes/jev/scripts/real_jev_gate.py`. The script reads work/step-rl/tasks.jsonl, rollouts/base.json, and judge-v2-heldout.jsonl. work/ is in .gitignore. On October 9 those three raw URLs on main returned HTTP 404. The script was not run. Those files are the 480 rollouts and 1,172 steps the note's hosted AUC would be recomputed from.
The same day this desk posted the states the unit tests freeze, with RemoteJev's Choice body, model jev-latest. The response model was jev-1.13.0. test_jev_remote.py's no-step state is "Research question: Q?" and "Final answer: (no answer)": incorrect 1.0, correct 0.0, confidence 1.0, 355 input tokens, 31 output tokens, 466 ms. The five-step state in that file, observations "r0" through "r4", final "Answer: Paris": correct 0.05, incorrect 0.95, confidence 0.9, 574 and 31, 419 ms. test_step_judge.py's one search, tool browser.search, action "{}", observation "reply": neutral 0.79, positive 0.04, negative 0.17, confidence 0.69, 404 input tokens and 38 output tokens, 645 ms. The step score, P(positive) minus P(negative), is -0.13.
A second set used the same two question strings and the same state builders, on a question written here: "In what year did a human first walk on the Moon?" No search and the answer 1965: incorrect 1.0, confidence 1.0, 366 input tokens, 31 output tokens, 454 ms. One NASA result and the answer 1969: correct 1.0, confidence 1.0, 414 and 31, 430 ms. The first Apollo 11 search: positive 1.0, confidence 1.0, 451 input tokens and 38 output tokens, 436 ms. The same search repeated: positive 0.53, negative 0.47, neutral 0.0, confidence 0.3, 499 and 38, 423 ms. The step score falls from 1.0 to 0.06. jev_final_step clamps a repeat to at most 0, so that 0.06 would not be paid.
Compare
The 19.6%, 5.4%, 15.6%, and 22.9% cells are one seed of one search agent. Accuracy is a strict normalized string match on 60 held-out questions, eight rollouts each. The paired gap of +14.2 points is Jev final plus step against Jev final, interval +8.1 to +20.4. Against the string-match answer key the same arm is -3.3 points, interval -8.3 to +1.5. Hosted Jev, zero-shot, is a different arm on that page: final-answer AUC 0.751, step AUC 0.891, accuracy 0%.
recipes/jev/docs/launch-results.md, dated 2026-09-28, is a different measurement. It scores a Qwen3.5-4B-Base LoRA at 79.1% against TypeSafe jev-1.13.0 at 66.8% on 1,920 AgentProcessBench steps. That file is not the BrowseComp-Plus table.
The note's hosted AUC, 0.751 on final answers and 0.891 on steps, is what real_jev_gate.py prints from files that are not in the repository. This desk did not recompute those two numbers, or the line that instant wrong answers were rated about 70% likely correct.
On the states the tests publish, an empty final answer was P(correct) 0.0, five placeholder searches plus "Answer: Paris" were P(correct) 0.05, and a search whose result is the word "reply" was neutral. On the Moon-landing question, 1965 with no search was also P(correct) 0.0, and 1969 after a NASA line was 1.0. The repeat of that NASA search stayed positive at 0.53. The published rollouts behind the 70% line were not in either set.
Terms
- Step Jev
- A Qwen3.5-4B-Base LoRA that scores one observed search step. The score used in training is P(positive) minus P(negative).
- Outcome Jev
- A Qwen3.5-4B-Base LoRA that reads the question, the actions, the last tool results, and the final answer, and returns P(correct).
- Dr. GRPO
- In this repository, the advantage is the rollout reward minus the mean reward of the other rollouts for the same question. The code does not divide by the group's spread.