REVIEW 1 major objections 5 minor 95 references
Sharding Prevents LLM Oversight Failures and Adversarial Exploitation
T0 review · 1 major / 5 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read The paper claims that decision load — the number of verdicts one LLM call must return — is the binding constraint in model-based oversight, and that sharding verdicts across calls recovers expert agreement and blocks presentation attacks…
desk verdict The empirical evidence that sharding helps is real; the claim that it isolates decision load from compute is not supported by the reported budget accounting. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the capacity account — a factorization of per-decision attention: for a failed decision $k$ assigned to a call with $n$ verdicts and budget $b$, the detection probability is $\Pr[s_k \text{ high} \mid y_k = 1] = \rho_k \, \alpha(n,b)$, with $\alpha(n,b) = g(b/n)\, h(n)$, where $\rho_k$ is how recoverably the failure is exposed by the evidence, $g$ is nondecreasing in per-decision budget, and $h$ is nonincreasing in decision load, sitting near $1$ below a judge-specific capacity $n^\ast$ and falling past it. This identity separates 'not enough compute' from 'too many verdicts': raising the per-call budget grows $g$ but leaves $h$ depressed, while sharding reduces $n$ and restores $h$ at fixed total budget. The experimental arms — solo, full-budget solo, pooled opinions, and sharded — are constructed so that the sharded-minus-full-budget contrast varies only the load per call ($K$ against $K/S$) while holding model, evidence, total budget, and per-decision budget fixed, which is what lets the paper attribute the gain to division rather than to compute.
What would settle it
Re-run the PaperBench load sweep at equal billed cost: compute the API-billed input-plus-output tokens for the whole sharded panel (each call re-sends the full context) and cap the full-budget solo call at that same total token bill, then see whether the pooled kappa gap of +0.142 the paper reports persists or collapses. The paper's own cost appendix reports input-token costs differing by roughly a factor of 23 between group sizes, so this is a directly checkable quantity, and if the gap vanishes at equal billing the load-only conclusion fails.
Extended reading notes
Core claim
Across expert-labeled datasets in three domains, the paper claims that decision load — the number of verdicts one call must return — degrades oversight accuracy that additional single-call budget cannot recover, and that sharding restores it at fixed total budget. In its strongest form the claim is that a single holistic call can spend the panel's full combined budget and still scrutinize each decision less than a divided panel; a sharded weaker judge can outperform a more capable holistic judge and match it even when the stronger judge receives the full budget. The same unsupported fraction of verdicts that drives the accuracy loss is directionally selectable: a best-of-N presentation attack that changes only how the artifact is framed raises over-acceptance of genuinely unmet criteria severalfold on grading (and raises a code verifier's cleared-defect rate from 0.12 to 0.44), and sharding keeps over-acceptance low as the adversary's search widens. The defense is largest when many criteria require one shared understanding of a large submission (up to 0.261 reduction in over-acceptance), small when each decision can be evaluated from a local slice, and null when decisions are independent. Sharding does not defend attacks that persuade the judge on each criterion individually; there, one-sided opposition on top of the sharded base limits a five-round adaptive attack's cumulative surface to 0.39–0.56, against 0.96 for sharding alone at the weaker tier.
Load-bearing premise
That the full-budget solo call and the sharded panel really spend the same total budget with only the division of verdicts different — the paper's headline contrast isolates decision load only if that budget accounting is exactly right.
Editorial extensions
If this is right
- Oversight pipelines should report decision load alongside token and tool budgets, because one holistic call can spend as much as a divided panel while scrutinizing each decision less.
- Task structure can partly substitute for judge capability: a sharded weaker judge can outperform a stronger holistic judge, so pipeline designers can trade model tier for decision grouping.
- Best-of-N presentation selection is a practical threat to overloaded judges — an adversary changes only the framing and severalfolds the acceptance of genuinely unmet criteria — so unsharded single-call oversight is exploitable at the presentation layer.
- Wherever sharding reduces baseline error it closes the attack surface, keeping over-acceptance low as the adversary's search widens; where decisions are independent or within a judge's capacity, sharding gives little or nothing.
- Load-exploiting attacks and per-criterion persuasion attacks need different defenses: division for the first, opposition on the sharded base for the second.
Reading between the lines
- Under ordinary API billing the budget-equality framing is fragile: because each shard re-sends the full context, the sharded panel's billed input tokens exceed the full-budget solo call's by a factor the paper's own appendix puts near 23 at some group sizes, so the real-world cost of the accuracy and security gains is higher than the fixed-budget framing suggests — the direction of the effect is u
- The capacity account makes a testable prediction for future model tiers: as judges' capacity $n^\ast$ grows, the group size at which sharding pays off should rise, and the load deficit should eventually disappear on routine rubrics; re-running the same load sweep on a newer model would confirm or bound the mechanism.
- The attack-defense asymmetry suggests a design rule: match the defense structure to the attack's unit of exploitation — division against overload, balanced evidence against per-criterion persuasion — so a pipeline that is safe in load-based evaluations can still be vulnerable to memo-style persuasion attacks.
- A practical engineering rule follows from the observed capability dependence: choose shard group size by model tier and decision difficulty (capacity-sized grouping rather than one verdict per call), which the paper's cost comparison shows is far cheaper per unit of agreement.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper studies LLM-based oversight when a single call must return many verdicts. It claims that agreement with expert labels falls as the number of decisions per call grows; that giving the same call a larger budget does not reverse this; that partitioning decisions into separate calls ('sharding') improves agreement while holding model, evidence, total budget, and per-decision budget fixed; that a best-of-N presentation attack can exploit the overload and sharding mitigates it; and that only debate-style opposition handles adaptive per-criterion persuasion. Evidence is drawn from expert labels on PaperBench, JudgmentBench, and ROBoto2, plus code, cyber, synthetic, and structural-control datasets.
Significance. The question is important for scalable oversight and LLM-as-judge pipelines. If the equal-budget claim held, the paper would establish decision load as a distinct bottleneck and offer a simple, practical intervention with security implications. The manuscript has genuine strengths: external expert labels in three domains, cluster-resampled intervals as the conservative inferential unit, explicitly disclosed scope conditions, a ground-truth-fixed attack construction, independent stronger-model re-validation of the cyber labels, and a capacity account tested on a controlled synthetic grid with a falsifiable prediction about group size and judge capability. However, the absence of a genuinely budget-matched control currently leaves the central claim unproven in the reported grading experiments.
major comments (1)
- [§2.1, Appendix A.2, A.4, B.2] The 'equal total budget' assertion is not satisfied by the implemented grading arms. Section 2.1 defines the arms so that sharded, pooled opinions, and full-budget solo 'spend the same total budget SB,' and claims the contrast (sharded minus full-budget solo) varies decision load alone. In the PaperBench load sweep (Appendix A.4), each sharded call is capped at 1,500 output tokens and re-sends the full evidence, while the full-budget call is a single context pass with an output cap of min(8000,1500+120K). At K=116 the sharded panel uses 116 context passes and a nominal 174,000 output tokens versus one context pass and 8,000 output tokens for the full-budget call; per decision the sharded judge receives 1,500 output tokens versus about 69 for the full-budget judge. The cost table in Appendix B.2 reports input-token totals that differ by roughly a factor of 23 between shard group sizes (50,524 tokens at g=32 versus 1,155,404 at g=1), so total token cost is not fixed across arms under ordinary billing. The paper defines a pooled-opinions arm that matches the sharded panel's true total token count while keeping all K decisions per call, but no results from this arm are reported anywhere. The central conclusion that 'decision load, not total compute, is the binding constraint' (Section 3, Table S19) and the security claims in Sections 4-5 therefore lack the required control. Please run and report the pooled-opinions arm, or equalize total and per-decision token budgets across arms, and rerun the central contrasts.
minor comments (5)
- [Abstract, §3] The statement that the full-budget single judge 'does not improve' should be qualified: Table S19 shows the full-budget arm exceeding holistic at K=8 (0.708 versus 0.631), so the claim is specifically about high-load conditions.
- [§2.1] The term 'budget' is used ambiguously: for grading it appears to mean the output-token cap, but re-sending the full context to every shard means input-token cost differs across arms; the paper should specify whether budget means output tokens, input-plus-output tokens, or dollar cost.
- [Appendix A.2] The full-budget grading arm adds a step-by-step instruction in addition to the larger token cap, so it is a budget-plus-instruction control rather than a pure budget manipulation; the synthetic control isolates this, but the text should state the distinction plainly.
- [Appendix B.4, B.5] In the text under review, the headings 'JudgmentBench Tables' and 'ROBoto2 Table' are placeholders with no content; the corresponding results appear later in B.8 and Table S13, so the appendix numbering should be cleaned up.
- [§7] The capacity account introduces h(n) and n* as free constructs fit on one synthetic dataset; the manuscript should make explicit that these are descriptive quantities not estimated from the expert datasets, to avoid implying a parameter-free derivation.
Circularity Check
No significant circularity found; central results are externally anchored and the one self-citation is non-load-bearing.
full rationale
The paper's load-bearing empirical claims are checked against external reference labels and fixed ground truth, not against its own fitted quantities. Agreement-with-expert results use expert-graded datasets (PaperBench, JudgmentBench, ROBoto2); the best-of-N attack holds ground truth fixed and denies the adversary access to labels; the code-verification results use hidden tests or an oracle dataset that the paper itself re-validates with an independent stronger refuter, explicitly removing a potential same-family label circularity in Appendix C. None of these central contrasts reduces to an input by construction. The capacity account in Section 7 is a post-hoc multiplicative fit, alpha(n,b) = g(b/n)h(n), fitted to the same synthetic grid it is used to interpret; the paper labels this as a fit ('a multiplicative fit explains 91 percent of the grid's variance'), presents it as an interpretation rather than as a first-principles prediction, and does not use it as evidence for the empirical claims. A fit explaining a dataset is not a derived prediction, so this does not constitute fitted-input-as-prediction circularity. The only self-citation (Nayebi, 2026, in Related Work) is contextual and non-load-bearing: the paper does not invoke it to justify sharding or to forbid alternative explanations. The equal-budget control is asserted in Section 2.1, and any mismatch in actual token accounting is a validity concern about the control, not a circularity in the derivation chain. Overall, the derivation is self-contained with respect to its external benchmarks, and no circular step merits a score above the low range.
Assumptions & free parameters
free parameters (5)
- Full-budget token cap parameters =
min(8000, 1500 + 120K)
- Default shard group size g =
4
- Load factor h(n) =
1.20, 0.99, 0.84 at n=8, 16, 32
- Budget factor g(b/n) =
0.92, 0.96, 1.01, 1.11 at r=16, 32, 64, 128
- Capacity threshold n* =
not quantified
assumptions (5)
- domain assumption Expert labels are the correct reference for oversight decisions.
- domain assumption Decisions are atomic, independent units that can be partitioned without loss.
- domain assumption The full-budget solo call receives the same total budget as the sharded panel.
- domain assumption Best-of-N adversary can select presentations by observing the judge's acceptance count without label access.
- ad hoc to paper Judge attention capacity is representable as a multiplicative factor h(n) that depends only on load.
invented entities (1)
-
latent per-decision attention capacity h(n)
Cite this review
Pith. "Pith review of Sharding Prevents LLM Oversight Failures and Adversarial Exploitation." pith.science (2026). https://pith.science/paper/SDJZWXYH
@misc{pith2026260806422,
author = {Pith},
title = {Pith review of: Sharding Prevents LLM Oversight Failures and Adversarial Exploitation},
year = {2026},
howpublished = {\url{https://pith.science/paper/SDJZWXYH}},
note = {Machine review of arXiv:2608.06422}
}
read the original abstract
Giving an LLM judge more compute does not necessarily make it check more requirements. When one call must return many verdicts, some decisions become weakly grounded in the evidence, even when that call receives the same token or tool budget as a panel of separate calls. Across expert-graded research replications, legal work, and clinical-trial assessments, agreement with experts falls as the number of verdicts per call grows. We identify sharding as the intervention that mitigates this failure in model-based oversight. Sharding partitions the requirements into smaller groups, assigns each group to a separate call, and aggregates the verdicts. Against a single call with the panel's full budget, sharding improves agreement while holding the model, evidence, total budget, and per-decision budget fixed. Overall, we find that a sharded weaker judge can outperform a more capable holistic judge and match that judge even when the latter receives the panel's full budget. Additionally, we find that sharding exhibits robustness against adversaries. A best-of-N adversary can hold the underlying work fixed, vary only its presentation, and increase an overloaded judge's acceptance of genuinely unmet criteria severalfold. Wherever sharding reduces baseline error, it removes this adversarial advantage, keeping over-acceptance low even as the adversary's search widens. Sharding does not address attacks that persuade the judge separately on each criterion rather than exploiting overload. In that setting, we find that debate-style opposition on top of sharding withstands such adaptive re-optimization.
Figures
Figures from the paper (3 more)
Reference graph
Works this paper leans on
-
[1]
Learning to give checkable answers with prover-verifier games, 2021
Cem Anil, Guodong Zhang, Yuhuai Wu, and Roger Grosse. Learning to give checkable answers with prover-verifier games, 2021
2021
-
[2]
Bowman, Jeeyoon Hyun, Ethan Perez, Edwin Chen, Craig Pettit, Scott Heiner, Kamil \.e Luko s i \=u t \.e , Amanda Askell, Andy Jones, Anna Chen, et al
Samuel R. Bowman, Jeeyoon Hyun, Ethan Perez, Edwin Chen, Craig Pettit, Scott Heiner, Kamil \.e Luko s i \=u t \.e , Amanda Askell, Andy Jones, Anna Chen, et al. Measuring progress on scalable oversight for large language models, 2022
2022
-
[3]
Weak-to-strong generalization: Eliciting strong capabilities with weak supervision
Collin Burns, Pavel Izmailov, Jan Hendrik Kirchner, Bowen Baker, Leo Gao, Leopold Aschenbrenner, Yining Chen, Adrien Ecoffet, Manas Joglekar, Jan Leike, Ilya Sutskever, and Jeffrey Wu. Weak-to-strong generalization: Eliciting strong capabilities with weak supervision. In Proceedings of the 41st International Conference on Machine Learning, 2024
2024
-
[4]
Supervising strong learners by amplifying weak experts, 2018
Paul Christiano, Buck Shlegeris, and Dario Amodei. Supervising strong learners by amplifying weak experts, 2018
2018
-
[5]
A coefficient of agreement for nominal scales
Jacob Cohen. A coefficient of agreement for nominal scales. Educational and Psychological Measurement, 20 0 (1): 0 37--46, 1960
1960
-
[6]
Zico Kolter
Jeremy Cohen, Elan Rosenfeld, and J. Zico Kolter. Certified adversarial robustness via randomized smoothing. In Proceedings of the 36th International Conference on Machine Learning, pages 1310--1320, 2019
2019
-
[7]
Limits to scalable evaluation at the frontier: LLM as judge won't beat twice the data
Florian Eddie Dorner, Vivian Nastl, and Moritz Hardt. Limits to scalable evaluation at the frontier: LLM as judge won't beat twice the data. In International Conference on Learning Representations, 2025. URL https://proceedings.iclr.cc/paper_files/paper/2025/hash/4264ee4376776907c0b87ed70b959585-Abstract-Conference.html
2025
-
[8]
Baek, Subhash Kantamneni, and Max Tegmark
Joshua Engels, David D. Baek, Subhash Kantamneni, and Max Tegmark. Scaling laws for scalable oversight, 2025
2025
Show all 95 references
-
[9]
Factored verification: Detecting and reducing hallucination in summaries of academic papers, 2023
Charlie Fadeeva George and Andreas Stuhlm \"u ller. Factored verification: Detecting and reducing hallucination in summaries of academic papers, 2023
2023
-
[10]
Criteria-based LLM relevance judgments
Naghmeh Farzi and Laura Dietz. Criteria-based LLM relevance judgments. In Proceedings of the 2025 ACM SIGIR International Conference on the Theory of Information Retrieval, 2025. URL https://www.cs.unh.edu/ dietz/papers/farzi2025criteria.pdf
2025
-
[11]
Scaling laws for reward model overoptimization
Leo Gao, John Schulman, and Jacob Hilton. Scaling laws for reward model overoptimization. In Proceedings of the 40th International Conference on Machine Learning, pages 10835--10866, 2023
2023
-
[12]
Chandra, Ponnurangam Kumaraguru, Douwe Kiela, Ameya Prabhu, Matthias Bethge, and Jonas Geiping
Shashwat Goel, Joschka Str \"u ber, Ilze Amanda Auzina, Karuna K. Chandra, Ponnurangam Kumaraguru, Douwe Kiela, Ameya Prabhu, Matthias Bethge, and Jonas Geiping. Great models think alike and this undermines AI oversight. In Proceedings of the 42nd International Conference on M...
2025
-
[13]
Klassen, and Lucy Lu Wang
Anthony Hevia, Sanjana Chintalapati, Veronica Ka Wai Lai, Nguyen Thanh Tam, Wai-Tat Wong, Terry P. Klassen, and Lucy Lu Wang. ROBOTO2 : An interactive system and dataset for LLM -assisted clinical trial risk of bias assessment. In Proceedings of the 2025 Conference on Empirica...
2025
-
[14]
AI safety via debate, 2018
Geoffrey Irving, Paul Christiano, and Dario Amodei. AI safety via debate, 2018
2018
-
[15]
Bowman, Tim Rockt \"a schel, and Ethan Perez
Akbir Khan, John Hughes, Dan Valentine, Laura Ruis, Kshitij Sachan, Ansh Radhakrishnan, Edward Grefenstette, Samuel R. Bowman, Tim Rockt \"a schel, and Ethan Perez. Debating with more persuasive LLMs leads to more truthful answers. In Proceedings of the 41st International Conf...
2024
-
[16]
Prover-verifier games improve legibility of LLM outputs, 2024
Jan Hendrik Kirchner, Yining Chen, Harri Edwards, Jan Leike, Nat McAleese, and Yuri Burda. Prover-verifier games improve legibility of LLM outputs, 2024
2024
-
[17]
SHADE-Arena : Evaluating sabotage and monitoring in LLM agents, 2025
Jonathan Kutasov, Yuqi Sun, Paul Colognese, Teun van der Weij, Linda Petrini, Chen Bo Calvin Zhang, John Hughes, Xiang Deng, Henry Sleight, Tyler Tracy, Buck Shlegeris, and Joe Benton. SHADE-Arena : Evaluating sabotage and monitoring in LLM agents, 2025
2025
-
[18]
Selective weak-to-strong generalization
Hao Lang et al. Selective weak-to-strong generalization. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 40, 2026
2026
-
[19]
Scalable agent alignment via reward modeling: A research direction, 2018
Jan Leike, David Krueger, Tom Everitt, Miljan Martic, Vishal Maini, and Shane Legg. Scalable agent alignment via reward modeling: A research direction, 2018
2018
-
[20]
LLMs cannot reliably judge (yet?): A comprehensive assessment on the robustness of LLM -as-a-judge, 2025
Songze Li, Chuokun Xu, Jiaying Wang, Xueluan Gong, Chen Chen, Jirui Zhang, Jun Wang, Kwok-Yan Lam, and Shouling Ji. LLMs cannot reliably judge (yet?): A comprehensive assessment on the robustness of LLM -as-a-judge, 2025
2025
-
[21]
Let's verify step by step, 2023
Hunter Lightman, Vineet Kosaraju, Yura Burda, Harri Edwards, Bowen Baker, Teddy Lee, Jan Leike, John Schulman, Ilya Sutskever, and Karl Cobbe. Let's verify step by step, 2023
2023
-
[22]
Mitigating self-preference by authorship obfuscation
Taslim Mahbub et al. Mitigating self-preference by authorship obfuscation. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 40, 2026
2026
-
[23]
McKenzie et al
Ian R. McKenzie et al. STACK : Adversarial attacks on LLM safeguard pipelines. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 40, 2026
2026
-
[24]
ARVO : Atlas of reproducible vulnerabilities for open source software, 2024
Xiang Mei, Pujan Aurangzeb, Weiteng Chen, Dongyeop Kim, Meng Xu, and Ruoyu Sun. ARVO : Atlas of reproducible vulnerabilities for open source software, 2024
2024
-
[25]
Julian Michael, Salsabila Mahdi, David Rein, Jackson Petty, Julien Dirani, Vishakh Padmakumar, and Samuel R. Bowman. Debate helps supervise unreliable experts, 2023
2023
-
[26]
FActScore : Fine-grained atomic evaluation of factual precision in long form text generation
Sewon Min, Kalpesh Krishna, Xinxi Lyu, Mike Lewis, Wen-tau Yih, Pang Wei Koh, Mohit Iyyer, Luke Zettlemoyer, and Hannaneh Hajishirzi. FActScore : Fine-grained atomic evaluation of factual precision in long form text generation. In Proceedings of the 2023 Conference on Empirica...
2023
-
[27]
Rossi, Seunghyun Yoon, and Hinrich Schuetze
Ali Modarressi, Hanieh Deilamsalehy, Franck Dernoncourt, Trung Bui, Ryan A. Rossi, Seunghyun Yoon, and Hinrich Schuetze. NoLiMa : Long-context evaluation beyond literal matching. In Proceedings of the 42nd International Conference on Machine Learning, volume 267 of Proceedings...
2025
-
[28]
Intrinsic barriers and practical pathways for human- AI alignment: An agreement-based complexity analysis
Aran Nayebi. Intrinsic barriers and practical pathways for human- AI alignment: An agreement-based complexity analysis. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 40, 2026
2026
-
[29]
LieCraft : A multi-agent framework for evaluating deceptive capabilities in language models
Matthew Lyle Olson et al. LieCraft : A multi-agent framework for evaluating deceptive capabilities in language models. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 40, 2026
2026
-
[30]
The effects of reward misspecification: Mapping and mitigating misaligned models
Alexander Pan, Kush Bhatia, and Jacob Steinhardt. The effects of reward misspecification: Mapping and mitigating misaligned models. In International Conference on Learning Representations, 2022
2022
-
[31]
Bowman, and Shi Feng
Arjun Panickssery, Samuel R. Bowman, and Shi Feng. LLM evaluators recognize and favor their own generations, 2024
2024
-
[32]
Confirmation bias: A challenge for scalable oversight
Gabriel Recchia et al. Confirmation bias: A challenge for scalable oversight. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 40, 2026 a
2026
-
[33]
FindTheFlaws : Annotated errors for detecting flawed reasoning and scalable oversight research
Gabriel Recchia et al. FindTheFlaws : Annotated errors for detecting flawed reasoning and scalable oversight research. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 40, 2026 b
2026
-
[34]
Verbosity bias in preference labeling by large language models, 2023
Keita Saito, Akifumi Wachi, Koki Wataoka, and Youhei Akimoto. Verbosity bias in preference labeling by large language models, 2023
2023
-
[35]
Self-critiquing models for assisting human evaluators, 2022
William Saunders, Catherine Yeh, Jeff Wu, Steven Bills, Long Ouyang, Jonathan Ward, and Jan Leike. Self-critiquing models for assisting human evaluators, 2022
2022
-
[36]
SWE-Bench Pro : Can AI agents solve long-horizon software engineering tasks?, 2025
Scale AI . SWE-Bench Pro : Can AI agents solve long-horizon software engineering tasks?, 2025. arXiv preprint
2025
-
[37]
PaperBench : Evaluating AI 's ability to replicate AI research, 2025
Giulio Starace, Oliver Jaffe, Dane Sherburn, James Aung, Jun Shern Chan, Leon Maksin, Rachel Dias, Evan Mays, Benjamin Kinsella, Wyatt Thompson, Johannes Heidecke, Amelia Glaese, and Tejal Patwardhan. PaperBench : Evaluating AI 's ability to replicate AI research, 2025
2025
-
[38]
Jonathan A. C. Sterne, Jelena Savovi \'c , Matthew J. Page, Roy G. Elbers, Natalie S. Blencowe, Isabelle Boutron, Christopher J. Cates, Hung-Yuan Cheng, Mark S. Corbett, Sandra M. Eldridge, et al. RoB 2: A revised tool for assessing risk of bias in randomised trials. BMJ, 366:...
2019
-
[39]
Christiano
Nisan Stiennon, Long Ouyang, Jeffrey Wu, Daniel Ziegler, Ryan Lowe, Chelsea Voss, Alec Radford, Dario Amodei, and Paul F. Christiano. Learning to summarize with human feedback. In Advances in Neural Information Processing Systems, volume 33, pages 3008--3021, 2020
2020
-
[40]
JudgeBench : A benchmark for evaluating LLM -based judges
Sijun Tan, Siyuan Zhuang, Kyle Montgomery, William Yuan Tang, Alejandro Cuadron, Chenguang Wang, Raluca Ada Popa, and Ion Stoica. JudgeBench : A benchmark for evaluating LLM -based judges. In International Conference on Learning Representations, 2025. URL https://proceedings.i...
2025
-
[41]
Large language models are not fair evaluators, 2023
Peiyi Wang, Lei Li, Liang Chen, Zefan Cai, Dawei Zhu, Binghuai Lin, Yunbo Cao, Qi Liu, Tianyu Liu, and Zhifang Sui. Large language models are not fair evaluators, 2023
2023
-
[42]
Ziegler, Nisan Stiennon, Ryan Lowe, Jan Leike, and Paul Christiano
Jeff Wu, Long Ouyang, Daniel M. Ziegler, Nisan Stiennon, Ryan Lowe, Jan Leike, and Paul Christiano. Recursively summarizing books with human feedback, 2021
2021
-
[43]
Does context matter? ContextualJudgeBench for evaluating LLM -based judges in contextual settings
Austin Xu, Srijan Bansal, Yifei Ming, Semih Yavuz, and Shafiq Joty. Does context matter? ContextualJudgeBench for evaluating LLM -based judges in contextual settings. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics, pages 9541--9564, ...
2025 doi
-
[44]
JudgmentBench : Comparing rubric and preference evaluation for quality assessment
Russell Yang, Ruishi Chen, Pierce Kelaita, Riya Ranjan, Sibo Ma, Charles Dickens, Matthew Guillod, Megan Ma, and Julian Nyarko. JudgmentBench : Comparing rubric and preference evaluation for quality assessment. arXiv preprint arXiv:2605.25240, 2026
2026 arXiv
-
[45]
Gonzalez, and Ion Stoica
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica. Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena . In Advances in Neural Information Processin...
2023
-
[46]
Evaluating judges as evaluators: The JETTS benchmark of LLM -as-judges as test-time scaling evaluators, 2025
Yilun Zhou, Austin Xu, Peifeng Wang, Caiming Xiong, and Shafiq Joty. Evaluating judges as evaluators: The JETTS benchmark of LLM -as-judges as test-time scaling evaluators, 2025
2025
-
[47]
Zico Kolter, and Matt Fredrikson
Andy Zou, Zifan Wang, Nicholas Carlini, Milad Nasr, J. Zico Kolter, and Matt Fredrikson. Universal and transferable adversarial attacks on aligned language models, 2023
2023
-
[48]
2025 , eprint =
Starace, Giulio and Jaffe, Oliver and Sherburn, Dane and Aung, James and Chan, Jun Shern and Maksin, Leon and Dias, Rachel and Mays, Evan and Kinsella, Benjamin and Thompson, Wyatt and Heidecke, Johannes and Glaese, Amelia and Patwardhan, Tejal , title =. 2025 , eprint =
2025
-
[49]
Sterne, Jonathan A. C. and Savovi. BMJ , volume =
-
[50]
and Wang, Lucy Lu , booktitle =
Hevia, Anthony and Chintalapati, Sanjana and Lai, Veronica Ka Wai and Tam, Nguyen Thanh and Wong, Wai-Tat and Klassen, Terry P. and Wang, Lucy Lu , booktitle =
-
[51]
Yang, Russell and Chen, Ruishi and Kelaita, Pierce and Ranjan, Riya and Ma, Sibo and Dickens, Charles and Guillod, Matthew and Ma, Megan and Nyarko, Julian , journal =
-
[52]
2024 , eprint =
Tian, Minyang and Gao, Luyu and Zhang, Shizhuo Dylan and Chen, Xinan and Fan, Cunwei and Guo, Xuefei and Haas, Roland and Ji, Pan and Krongchon, Kittithat and Li, Yao and others , title =. 2024 , eprint =
2024
-
[53]
2024 , eprint =
Mei, Xiang and Aurangzeb, Pujan and Chen, Weiteng and Kim, Dongyeop and Xu, Meng and Sun, Ruoyu , title =. 2024 , eprint =
2024
-
[54]
Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing , pages =
Min, Sewon and Krishna, Kalpesh and Lyu, Xinxi and Lewis, Mike and Yih, Wen-tau and Koh, Pang Wei and Iyyer, Mohit and Zettlemoyer, Luke and Hajishirzi, Hannaneh , title =. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing , pages =
2023
-
[55]
Arora, Rahul K. and Wei, Jason and Soskin Hicks, Rebecca and Bowman, Preston and Quinonero-Candela, Joaquin and Tsimpourlas, Foivos and Sharman, Michael and Shah, Meghan and Vallone, Andrea and Beutel, Alex and Heidecke, Johannes and Singhal, Karan , title =. 2025 , eprint =
2025
-
[56]
2018 , eprint =
Irving, Geoffrey and Christiano, Paul and Amodei, Dario , title =. 2018 , eprint =
2018
-
[57]
2018 , eprint =
Leike, Jan and Krueger, David and Everitt, Tom and Martic, Miljan and Maini, Vishal and Legg, Shane , title =. 2018 , eprint =
2018
-
[58]
2018 , eprint =
Christiano, Paul and Shlegeris, Buck and Amodei, Dario , title =. 2018 , eprint =
2018
-
[59]
2024 , eprint =
Kirchner, Jan Hendrik and Chen, Yining and Edwards, Harri and Leike, Jan and McAleese, Nat and Burda, Yuri , title =. 2024 , eprint =
2024
-
[60]
2021 , eprint =
Anil, Cem and Zhang, Guodong and Wu, Yuhuai and Grosse, Roger , title =. 2021 , eprint =
2021
-
[61]
and Stoica, Ion , title =
Zheng, Lianmin and Chiang, Wei-Lin and Sheng, Ying and Zhuang, Siyuan and Wu, Zhanghao and Zhuang, Yonghao and Lin, Zi and Li, Zhuohan and Li, Dacheng and Xing, Eric and Zhang, Hao and Gonzalez, Joseph E. and Stoica, Ion , title =. Advances in Neural Information Processing Sys...
-
[62]
2023 , eprint =
Wang, Peiyi and Li, Lei and Chen, Liang and Cai, Zefan and Zhu, Dawei and Lin, Binghuai and Cao, Yunbo and Liu, Qi and Liu, Tianyu and Sui, Zhifang , title =. 2023 , eprint =
2023
-
[63]
and Feng, Shi , title =
Panickssery, Arjun and Bowman, Samuel R. and Feng, Shi , title =. 2024 , eprint =
2024
-
[64]
2023 , eprint =
Saito, Keita and Wachi, Akifumi and Wataoka, Koki and Akimoto, Youhei , title =. 2023 , eprint =
2023
-
[65]
Proceedings of the 40th International Conference on Machine Learning , pages =
Gao, Leo and Schulman, John and Hilton, Jacob , title =. Proceedings of the 40th International Conference on Machine Learning , pages =
-
[66]
and Hyun, Jeeyoon and Perez, Ethan and Chen, Edwin and Pettit, Craig and Heiner, Scott and Luko
Bowman, Samuel R. and Hyun, Jeeyoon and Perez, Ethan and Chen, Edwin and Pettit, Craig and Heiner, Scott and Luko. Measuring Progress on Scalable Oversight for Large Language Models , year =. 2211.03540 , archivePrefix =
-
[67]
and Rockt
Khan, Akbir and Hughes, John and Valentine, Dan and Ruis, Laura and Sachan, Kshitij and Radhakrishnan, Ansh and Grefenstette, Edward and Bowman, Samuel R. and Rockt. Debating with More Persuasive. Proceedings of the 41st International Conference on Machine Learning , year =
-
[68]
, title =
Michael, Julian and Mahdi, Salsabila and Rein, David and Petty, Jackson and Dirani, Julien and Padmakumar, Vishakh and Bowman, Samuel R. , title =. 2023 , eprint =
2023
-
[69]
Factored Verification: Detecting and Reducing Hallucination in Summaries of Academic Papers , year =
Fadeeva George, Charlie and Stuhlm. Factored Verification: Detecting and Reducing Hallucination in Summaries of Academic Papers , year =. 2310.10627 , archivePrefix =
-
[70]
2023 , eprint =
Lightman, Hunter and Kosaraju, Vineet and Burda, Yura and Edwards, Harri and Baker, Bowen and Lee, Teddy and Leike, Jan and Schulman, John and Sutskever, Ilya and Cobbe, Karl , title =. 2023 , eprint =
2023
-
[71]
2022 , eprint =
Saunders, William and Yeh, Catherine and Wu, Jeff and Bills, Steven and Ouyang, Long and Ward, Jonathan and Leike, Jan , title =. 2022 , eprint =
2022
-
[72]
and Stiennon, Nisan and Lowe, Ryan and Leike, Jan and Christiano, Paul , title =
Wu, Jeff and Ouyang, Long and Ziegler, Daniel M. and Stiennon, Nisan and Lowe, Ryan and Leike, Jan and Christiano, Paul , title =. 2021 , eprint =
2021
-
[73]
Educational and Psychological Measurement , volume =
Cohen, Jacob , title =. Educational and Psychological Measurement , volume =
-
[74]
International Conference on Learning Representations , year =
Pan, Alexander and Bhatia, Kush and Steinhardt, Jacob , title =. International Conference on Learning Representations , year =
-
[75]
Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics , year =
Malaviya, Chaitanya and Lee, Subin and Chen, Sihao and Sieber, Elizabeth and Yatskar, Mark and Roth, Dan , title =. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics , year =
2024
-
[76]
, title =
Stiennon, Nisan and Ouyang, Long and Wu, Jeffrey and Ziegler, Daniel and Lowe, Ryan and Voss, Chelsea and Radford, Alec and Amodei, Dario and Christiano, Paul F. , title =. Advances in Neural Information Processing Systems , volume =
-
[77]
Zico and Fredrikson, Matt , title =
Zou, Andy and Wang, Zifan and Carlini, Nicholas and Nasr, Milad and Kolter, J. Zico and Fredrikson, Matt , title =. 2023 , eprint =
2023
-
[78]
Zico , title =
Cohen, Jeremy and Rosenfeld, Elan and Kolter, J. Zico , title =. Proceedings of the 36th International Conference on Machine Learning , pages =
-
[79]
Proceedings of the 41st International Conference on Machine Learning , year =
Burns, Collin and Izmailov, Pavel and Kirchner, Jan Hendrik and Baker, Bowen and Gao, Leo and Aschenbrenner, Leopold and Chen, Yining and Ecoffet, Adrien and Joglekar, Manas and Leike, Jan and Sutskever, Ilya and Wu, Jeffrey , title =. Proceedings of the 41st International Con...
-
[80]
Proceedings of the AAAI Conference on Artificial Intelligence , volume =
Nayebi, Aran , title =. Proceedings of the AAAI Conference on Artificial Intelligence , volume =
-
[81]
Proceedings of the AAAI Conference on Artificial Intelligence , volume =
Recchia, Gabriel and others , title =. Proceedings of the AAAI Conference on Artificial Intelligence , volume =
-
[82]
and others , title =
McKenzie, Ian R. and others , title =. Proceedings of the AAAI Conference on Artificial Intelligence , volume =
-
[83]
Proceedings of the AAAI Conference on Artificial Intelligence , volume =
Mahbub, Taslim and others , title =. Proceedings of the AAAI Conference on Artificial Intelligence , volume =
-
[84]
Proceedings of the AAAI Conference on Artificial Intelligence , volume =
Lang, Hao and others , title =. Proceedings of the AAAI Conference on Artificial Intelligence , volume =
-
[85]
Proceedings of the AAAI Conference on Artificial Intelligence , volume =
Olson, Matthew Lyle and others , title =. Proceedings of the AAAI Conference on Artificial Intelligence , volume =
-
[86]
International Conference on Learning Representations , year =
Tan, Sijun and Zhuang, Siyuan and Montgomery, Kyle and Tang, William Yuan and Cuadron, Alejandro and Wang, Chenguang and Popa, Raluca Ada and Stoica, Ion , title =. International Conference on Learning Representations , year =
-
[87]
Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics , pages =
Xu, Austin and Bansal, Srijan and Ming, Yifei and Yavuz, Semih and Joty, Shafiq , title =. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics , pages =. 2025 , doi =
2025
-
[88]
2025 , eprint =
Li, Songze and Xu, Chuokun and Wang, Jiaying and Gong, Xueluan and Chen, Chen and Zhang, Jirui and Wang, Jun and Lam, Kwok-Yan and Ji, Shouling , title =. 2025 , eprint =
2025
-
[89]
2025 , eprint =
Zhou, Yilun and Xu, Austin and Wang, Peifeng and Xiong, Caiming and Joty, Shafiq , title =. 2025 , eprint =
2025
-
[90]
Great Models Think Alike and This Undermines
Goel, Shashwat and Str. Great Models Think Alike and This Undermines. Proceedings of the 42nd International Conference on Machine Learning , series =. 2025 , url =
2025
-
[91]
and Kantamneni, Subhash and Tegmark, Max , title =
Engels, Joshua and Baek, David D. and Kantamneni, Subhash and Tegmark, Max , title =. 2025 , eprint =
2025
-
[92]
International Conference on Learning Representations , year =
Dorner, Florian Eddie and Nastl, Vivian and Hardt, Moritz , title =. International Conference on Learning Representations , year =
-
[93]
2025 , eprint =
Kutasov, Jonathan and Sun, Yuqi and Colognese, Paul and van der Weij, Teun and Petrini, Linda and Zhang, Chen Bo Calvin and Hughes, John and Deng, Xiang and Sleight, Henry and Tracy, Tyler and Shlegeris, Buck and Benton, Joe , title =. 2025 , eprint =
2025
-
[94]
and Yoon, Seunghyun and Schuetze, Hinrich , title =
Modarressi, Ali and Deilamsalehy, Hanieh and Dernoncourt, Franck and Bui, Trung and Rossi, Ryan A. and Yoon, Seunghyun and Schuetze, Hinrich , title =. Proceedings of the 42nd International Conference on Machine Learning , series =. 2025 , url =
2025
-
[95]
Proceedings of the 2025 ACM SIGIR International Conference on the Theory of Information Retrieval , year =
Farzi, Naghmeh and Dietz, Laura , title =. Proceedings of the 2025 ACM SIGIR International Conference on the Theory of Information Retrieval , year =
2025
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.