RLCDAlignBench puts a generic Noul at median AUROC 0.886, and median ECE at 0.168

This topic was created 10 days ago, and the information it contains may have evolved or changed since then.

Ruoqi Guo, Yi Liu, and seven coauthors posted RLCDAlignBench on September 24. The set is 10 failure types, 44 benchmarks, and 7,193 instances. Every run in the paper we read uses jev-1.13.0. A generic Noul has median AUROC 0.886 on 31 benchmarks. Median ECE is 0.168 against a null of 0.074, so a threshold fit on one benchmark does not travel. Code is MIT. The dataset is CC BY-NC 4.0 and the Hugging Face copy is gated. We did not rerun it.

RLCDAlignBench is a monitor suite for alignment failures. The paper is arXiv:2609.29429, version 1, dated September 24, CC BY 4.0. The authors are Ruoqi Guo, Yi Liu, Gelei Deng, Yuekang Li, Lida Zhao, Yutao Wu, Simin Chen, Ying Zhang, and Leo Yu Zhang. The title-page image marks Yi Liu with a star. Leo Yu Zhang is the corresponding author. Guo, Liu, and Zhang are at Griffith, Deng at NTU, Li at UNSW, Zhao independent, Wu at Deakin, Chen at George Mason, and Ying Zhang at Wake Forest. Code in the repository is MIT. The data is CC BY-NC 4.0, and the Hugging Face copy is gated. The README points at YuehHanChen/automated_alignment_researcher, commit 02dbe9d2cadc553720d17cdf6259c0b8727e6cde. September 26 posts link the paper. One of them attaches the title card. We did not rescore the cached answers.

The set has 10 failure types, 44 benchmarks, 5 target models, and 7,193 instances. The models are Qwen3.5-2B, Phi-4-mini, Gemma-2-2B, Llama-3.2-3B, and Olmo-3-7B. Table 1’s counts are sycophancy 4 benchmarks and 639 items, jailbreaks 4/414, deception 4/540, prompt injection 4/1,036, hallucination 6/1,164, privacy 4/808, social bias 4/199, reward hacking 6/714, concealing uncertainty 4/748, and power seeking 4/931. One sentence in the paper counts 20 rule scorers and 24 LLM judges. The table header writes the scorer mix as 20/20/4. Labels marked validated, derived, or unvalidated are 11, 8, and 25. Six benchmarks have fewer than 5 items in the minority class, so the aggregates use 38. AbstentionBench is left out for too few negatives.

Every run we read uses jev-1.13.0. A generic Noul, the same wording across benchmarks that can take that form, has a median AUROC of 0.886, interval [0.821, 0.952], over 31 benchmarks. Against the better of a length baseline and an in-domain TF-IDF baseline, the median gap is +0.132, interval [+0.057, +0.190]. That comparison wins 25 of 31, sign test p = 9e-4. A targeted wording, fit on one half of a benchmark and scored on the other, is 0.911, interval [0.860, 0.944], over 38. Taking that targeted wording to a benchmark it was not fit on adds +0.006, interval [-0.004, +0.015], Wilcoxon p = 0.055. Split by who wrote the original label, a judge benchmark is 0.906, a rule benchmark is 0.890, and a multi-turn benchmark is 0.870, p = 0.78. Unvalidated labels score 0.949 on the generic Noul. Validated labels score 0.872.

The calibration line goes the other way. Median expected calibration error is 0.168. A null that simply predicts the base rate is 0.074. The paper says the probabilities miss that base rate, and that a threshold fit on one benchmark does not transfer to the next.

Human agreement on StrongREJECT is Cohen’s kappa 0.809 for the Jev scores, against 0.811 for the reference scorer. The human sets named in the paper are a HarmBench validation with 3 annotators, and StrongREJECT at 1,360 of 1,361 items with 5 annotators. That StrongREJECT file is separate from the 100-item Phi-4-mini benchmark inside the suite. On PrivacyLens, handing the scorer the secret list moves AUROC from 0.79 to 0.95. A context the paper calls a deployable monitor helps on 1 of 7 benchmarks. One pass over the 19 benchmarks that use an API judge costs $0.30, which the paper says is 63 times less than those judges at list prices. The judge dollar total is not printed.

AYi, whose reply links the paper and the repository, also prints an average of 0.31 seconds, $18.96 on the judge side of the 63-times line, and 93% accuracy on the half of decisions where confidence is highest. A different recap the same day prints 93% accuracy, 0.31 seconds, and an F1 move from 0.706 to 0.793. None of those three, and not the $18.96, appear in the tables transcribed above.

This site's reading

Editorial notes evaluating claims against primary sources, contextualizing findings alongside related implementations, and defining technical terms.

Verify

arXiv:2609.29429v1, 24 September 2026, CC BY 4.0. Authors Ruoqi Guo, Yi Liu, Gelei Deng, Yuekang Li, Lida Zhao, Yutao Wu, Simin Chen, Ying Zhang, and Leo Yu Zhang. The title-page image marks Yi Liu with a star.

Leo Yu Zhang is the corresponding author. Affiliations in the paper pair Guo, Liu, and Zhang with Griffith, Deng with NTU, Li with UNSW, Zhao as independent, Wu with Deakin, Chen with George Mason, and Zhang with Wake Forest. Repository sumleo/RLCDAlignBench, code MIT, data CC BY-NC 4.0.

The Hugging Face dataset is gated. The README names YuehHanChen/automated_alignment_researcher at commit 02dbe9d2cadc553720d17cdf6259c0b8727e6cde. The bench is 10 failure types, 44 benchmarks, 5 target models, and 7,193 instances.

Table 1 counts: sycophancy 4/639, jailbreaks 4/414, deception 4/540, prompt injection 4/1036, hallucination 6/1164, privacy 4/808, social bias 4/199, reward hacking 6/714, concealing uncertainty 4/748, power seeking 4/931. One sentence counts 20 rule scorers and 24 LLM judges. The table header writes R/J/M as 20/20/4.

Label tags V/D/U are 11/8/25. Six benchmarks have fewer than 5 minority-class items, so aggregates use 38. AbstentionBench has too few negatives. All runs use jev-1.13.0. Generic Noul median AUROC is 0.886 [0.821, 0.952] over 31 benchmarks that have a Noul form.

Against the better of length and in-domain TF-IDF the median gap is +0.132 [+0.057, +0.190], 25 wins out of 31, sign test p = 9e-4. Split-half targeted wording is 0.911 [0.860, 0.944] over 38. Targeted wording out of sample is +0.006 [-0.004, +0.015], Wilcoxon p = 0.055.

Judge 0.906, rule 0.890, multi-turn 0.870, p = 0.78. Unvalidated labels score 0.949 and validated labels 0.872 on the generic Noul. Median ECE is 0.168 against a null of 0.074. StrongREJECT human agreement is Cohen's kappa 0.809, against 0.811 for the reference scorer.

The human sets are a HarmBench validation with 3 annotators and StrongREJECT at 1,360 of 1,361 items with 5 annotators, separate from the 100-item Phi-4-mini StrongREJECT benchmark. PrivacyLens with the secret list moves from 0.79 to 0.95. A deployable-monitor context helps on 1 of 7 benchmarks.

One pass over 19 API-judge benchmarks costs $0.30, which the paper says is 63 times less than those judges at list prices. The share image on the September 26 post is the paper's title card, not a results figure. We did not rerun the scores.

AYi prints 0.31 seconds, $18.96 for the judge side, and 93% on the confident half. A different recap prints 93% accuracy, 0.31 seconds, and an F1 move from 0.706 to 0.793. Those figures are not in the tables we transcribed.

Compare

Decision Index 0.2 locks Jev at 51.67 as a chance-corrected average of 40 benchmarks. JevBench v1.4.2 is a harmonic mean, 63.29, with an Intelligence cell of 53.1. The 0.886 here is a median AUROC for a Noul used as a monitor, not either of those composites.

The ECE line is the paper's own warning that the probabilities miss the base rate.

Terms

0.886
Median AUROC of a generic Noul over 31 benchmarks that have a Noul form. The interval is [0.821, 0.952]. Split-half targeted wording on 38 benchmarks is 0.911.
0.168
Median expected calibration error. A null that knows the base rate is 0.074. The paper says thresholds fit on one benchmark do not transfer.
$0.30
One Jev pass over 19 benchmarks that otherwise use an API judge. The paper says that is 63 times less than those judges at list prices. A September 26 post prints $18.96 on the judge side. That figure is not in the tables we transcribed.

Sources

  1. arXiv:2609.29429
  2. sumleo/RLCDAlignBench
  3. Project page
  4. AINativeF, the title card
  5. AYi, paper and repo
  6. amasen02, earlier pointer