REVIEW 2 major objections 6 minor 8 cited by
Neither Valid nor Reliable? Investigating the Use of LLMs as Judges
T0 review · 2 major / 6 minor · reviewed 2026-08-05 · deepseek-v4-flash
Pith's one-line read The paper argues that LLMs used as judges are being adopted faster than the evidence supports, and that correlation with human judgment—the field's main validation—establishes only one narrow kind of validity.
desk verdict A solid, well-cited position paper arguing that LLJ adoption has outpaced validity scrutiny — worth refereeing, but its central frequency claim rests on a review protocol too thin to carry it. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the social-science concept of construct validity, operationalized as a seven-dimension framework: face, content, convergent, discriminant, predictive, hypothesis, and consequential validity. The paper uses this framework as a checklist to show that the standard LLJ validation move—correlating machine scores with human judgments—covers only one dimension, and that reliability (often tested through position-bias and prompt-robustness studies) is a separate property. The framework does the argument's work by converting a vague worry about LLMs into a structured measurement diagnosis.
What would settle it
Conduct a preregistered systematic review of a defined corpus of LLJ papers, counting how often each of the four assumptions appears as the stated motivation and how often convergent validity is the only validation reported. If those framings are rare, or if most papers already report discriminant or predictive validity, the paper's central diagnosis fails. A second direct test: pick a benchmark where an LLJ passes a full seven-dimension validation, including robustness to adversarial prompt and response edits; a clean pass would not refute the general claim but would show the failure is conti
Extended reading notes
Core claim
The paper's central claim is that the current enthusiasm for LLMs as judges is premature because their adoption has outpaced rigorous scrutiny of their reliability and validity as evaluators. The authors use a social-science measurement framework to distinguish reliability (consistency of scores across repetitions) from validity (whether the scores actually measure the intended construct), and to decompose validity into dimensions including face, content, convergent, discriminant, predictive, hypothesis, and consequential validity. They argue that the LLJ literature has mostly relied on convergent validity—correlation with human judgments—while neglecting the other dimensions, and that even
Load-bearing premise
The load-bearing premise is that the four assumptions the paper critiques (proxy for humans, capable evaluator, scalable, cost-effective) are actually the common motivating frames in the LLJ literature; the paper's high-level qualitative review is labeled nonexhaustive and names no corpus, search protocol, or selection criteria, so if the sampled papers were chosen to fit the argument, the empirical grounding for the position collapses.
Editorial extensions
If this is right
- Correlating an LLJ's scores with human ratings should be treated as evidence of convergent validity only, not as proof that the judge is valid for the task.
- Comparisons across LLJ papers are unreliable until evaluation criteria, rating scales, and comparison modes are standardized, as the three SummEval-based studies show in miniature.
- Reliability testing—position bias, self-enhancement, adversarial prompt attacks—must be paired with validity testing; a consistent judge may still be measuring the wrong thing.
- Safety judges that rely on surface markers can misclassify harmful content as harmless, so safety decisions should not depend on a single LLJ score.
- Cost accounting for LLJs should include non-financial costs such as data-worker displacement, inference energy, and reproduced societal biases.
Reading between the lines
- If the paper is right, reported wins from LLJ-based training pipelines may partly reflect preference leakage between judge and student models rather than genuine capability gains; an explicit test would train identical models with judge rewards versus human-validated rewards and compare generalization.
- The framework suggests a new evaluation target: measuring LLJs' discriminant and predictive validity directly—for example, whether an LLJ's ranking of systems predicts downstream user satisfaction better than a human ranking does.
- Because human gold standards are themselves constructs with measurement error, some cases of low human-LLJ agreement may reflect unreliable human labels rather than an invalid judge; treating disagreement as information rather than noise could turn the paper's critique into a constructive measurement model.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This position paper argues that the widespread adoption of large language models as judges (LLJs) has outpaced rigorous validation. Drawing on measurement theory (Adcock and Collier; Jacobs and Wallach), it identifies four assumptions underlying LLJ use: proxies for human judgment, capable evaluators, scalability, and cost-effectiveness. For each assumption it surveys documented limitations—inconsistent human or LLJ judgment collection, prompt sensitivity and bias, contamination and competitive benchmarking, and overlooked economic/environmental/societal costs—and grounds the discussion in three applications: text summarization, data annotation, and safety alignment. The paper concludes with recommendations for standardizing and contextualizing LLJ evaluation. It is explicitly framed as a non-exhaustive qualitative review rather than a systematic meta-analysis.
Significance. If its central claim holds, the paper is a timely corrective: it argues that correlation with human judgments, the dominant validation strategy, is neither necessary nor sufficient for establishing the construct validity of LLJs. The measurement-theory framing is a genuine strength, as is the decomposition of the debate into four assumption types and the use of concrete cases (e.g., the SummEval instruction inconsistencies, safety-judge vulnerability). The paper is densely cited and honestly hedged—it openly says the review is not exhaustive. The main limitation is that the empirical basis for calling these assumptions 'common' and for claiming the field has 'adopted' them is a curated set of examples without a reproducible corpus or coding protocol; this representativeness problem is load-bearing for the paper's broader claim and needs to be addressed.
major comments (2)
- [§4; Tables 2–3] The paper's central premise—that the four assumptions are 'common' and that 'adoption has outpaced rigorous scrutiny'—rests on a 'high-level qualitative review of commonly cited works' (§4) with no stated corpus, search strategy, inclusion criteria, or coding scheme. The evidence shown is small and curated: approximately eight quotes in Table 2 and a three-paper SummEval comparison in Table 3. This does not establish representativeness. The 'not exhaustive' disclaimer is transparent, but transparency does not replace a defined sample. Because the frequency/representativeness claim is load-bearing for the title claim, the manuscript should either (a) specify the corpus and protocol and report inter-coder agreement, or (b) explicitly scope the argument to 'an influential strand of the literature' and avoid the stronger claim that the field as a whole adopts these assumptions.
- [§4.1; Table 3] The SummEval case study is compelling, but it is used to support a broader generalization: that the LLJ literature has 'adopted correlation with human judgment as the primary validation criterion' and that LLJ operationalizations are inconsistent. Three papers are a convenience sample; the reader cannot tell whether they were selected to fit the argument. Moreover, the inference from differing instruction texts and scales to weakened validity needs one more step: prompt variation may impair comparability across studies, but it does not by itself show that an individual LLJ lacks content or construct validity. The paper should either state how these three papers were chosen, or explicitly present the case study as illustrative rather than as evidence about the distribution of practices in the field.
minor comments (6)
- [§4.2 (Explainability)] The sentence 'none of these studies examined the faithfulness of the generated explanations' is a universal negative over a set that is not systematically enumerated. If 'these studies' refers only to the four cited works [15, 71, 48, 38], the statement is safe, but the surrounding text reads as a characterization of the literature. Also, 'faithfulness' is never defined; in the context of LLJ rationales it is ambiguous (faithful to the model's reasoning? faithful to the human-judgment construct?). Please qualify the scope and define the term.
- [Table 3] The table is hard to parse: the rows do not clearly identify which source corresponds to each instruction/scale/process combination, and the 'Scale' column mixes entries such as 'Likert Scale (1-3)', 'Binary Option or Scale 1-10', etc. Add a 'Source' column or row labels with citation keys so the reader can map each row to a specific paper.
- [§4.4] The statement that LLM-based evaluation is 'generally more economical overall' is an empirical cost claim made without a systematic cost model or a citation to one. Since the section's purpose is to complicate the cost-effectiveness assumption, this sentence should be hedged or backed by a concrete comparison, or removed in favor of 'commonly claimed.'
- [§3; Table 1] Several validity dimensions (e.g., hypothesis validity, predictive validity) are listed in Table 1 but are not explicitly operationalized in the later discussion. Mapping each challenge to a specific validity dimension would strengthen the measurement-theory contribution and avoid the impression that the framework is invoked only loosely.
- [References] In the provided version, many reference URLs are rendered as placeholder characters, making them unverifiable. Ensure the published/final manuscript contains complete, legible reference entries with functioning URLs.
- [Title] The title 'Neither Valid nor Reliable?' is broader than the evidence presented. The paper's original contribution and most of its evidence concern validity; reliability is discussed mainly through citations to existing bias/robustness work. Consider a more precise title, or add a short paragraph that explicitly delineates which aspects of reliability are or are not supported.
Circularity Check
No significant circularity: the paper's critique is anchored in external measurement theory and external empirical findings; its two self-citations are non-load-bearing supporting examples.
full rationale
This is a position paper arguing that LLM-as-a-judge adoption has outpaced validity/reliability scrutiny. The derivation chain is not a formal derivation with equations or fitted parameters; it is an argumentative synthesis. The load-bearing premises come from external sources: measurement theory [1,47], documented LLJ biases [117,51,81,104,57], human-evaluation critiques [42,120], contamination studies [53,57,34,83], and safety-judge robustness studies [24,13]. The paper's two self-citations ([10], [11]) are used only as supporting examples for side claims about safety safeguards and sociotechnical safety, not as premises on which the central claim depends. The main structural weakness is the 'high-level qualitative review of commonly cited works' in Section 4, which is disclosed as 'not exhaustive' and lacks a corpus, search protocol, or coding scheme; this makes the frequency claim that the four assumptions are 'common' unverified. Similarly, the universal negative in Section 4.2 that 'none of these studies examined the faithfulness of the generated explanations' is asserted without a systematic search. These are representativeness/selection-bias concerns, not circularity: the argument does not reduce to its own inputs by definition, and no fitted parameter is renamed as a prediction. The paper explicitly flags the limitation ('Although not exhaustive'), and the central critique would remain substantive even if some LLJ papers motivate adoption differently. Therefore no circular step meeting the quoted-evidence standard can be exhibited, and the appropriate score is 0.
Assumptions & free parameters
assumptions (6)
- domain assumption Measurement theory (Adcock and Collier's four-level framework and validity/reliability distinction, Jacobs and Wallach's construct-validity dimensions) transfers to LLM-based evaluators.
- ad hoc to paper The selected papers are representative of common LLJ motivations and framings.
- domain assumption Documented human-judgment inconsistencies (Howcroft et al., Zhou et al.) transfer to the benchmarks used to validate LLJs.
- domain assumption Cited LLM biases, prompt-attack results, and contamination findings generalize to LLJs as deployed.
- domain assumption The Superficial Alignment Hypothesis (Zhou et al. 2023) and stylistic-shift evidence (Lin et al. 2023) are accepted as sound.
- domain assumption Consequential validity (economic, environmental, societal impacts) is a legitimate criterion for judging LLJ measurement validity.
Cite this review
Pith. "Pith review of Neither Valid nor Reliable? Investigating the Use of LLMs as Judges." pith.science (2026). https://pith.science/paper/QUSVK6UF
@misc{pith2026250818076,
author = {Pith},
title = {Pith review of: Neither Valid nor Reliable? Investigating the Use of LLMs as Judges},
year = {2026},
howpublished = {\url{https://pith.science/paper/QUSVK6UF}},
note = {Machine review of arXiv:2508.18076}
}
read the original abstract
Evaluating natural language generation (NLG) systems remains a core challenge of natural language processing (NLP), further complicated by the rise of large language models (LLMs) that aims to be general-purpose. Recently, large language models as judges (LLJs) have emerged as a promising alternative to traditional metrics, but their validity remains underexplored. This position paper argues that the current enthusiasm around LLJs may be premature, as their adoption has outpaced rigorous scrutiny of their reliability and validity as evaluators. Drawing on measurement theory from the social sciences, we identify and critically assess four core assumptions underlying the use of LLJs: their ability to act as proxies for human judgment, their capabilities as evaluators, their scalability, and their cost-effectiveness. We examine how each of these assumptions may be challenged by the inherent limitations of LLMs, LLJs, or current practices in NLG evaluation. To ground our analysis, we explore three applications of LLJs: text summarization, data annotation, and safety alignment. Finally, we highlight the need for more responsible evaluation practices in LLJs evaluation, to ensure that their growing role in the field supports, rather than undermines, progress in NLG.
Forward citations
Cited by 8 Pith papers
-
Matter to Mechanism: A Benchmark for AI Co-Scientists in Materials and Battery Research
Introduces the Matter to Mechanism benchmark of 2,645 structured instances and a composite metric suite for evaluating AI co-scientists on problem-to-hypothesis reasoning in battery materials research.
-
GRASP: Deterministic argument ranking in interaction graphs
GRASP aggregates stable local LLM interaction judgments into global argument rankings via a convergent attack-defense propagation operator on interaction graphs, yielding higher reproducibility than holistic judging a...
-
Augmenting Human Evaluation with LLM Judges: How Many Human Reviews Do You Need?
The paper formulates LLM-as-judge evaluation as a two-stage missing-data problem and derives sample-size formulas via doubly robust estimators to achieve desired power while allocating more human reviews where LLM pre...
-
Evaluating AI-Generated Images of Cultural Artifacts with Community-Informed Rubrics
Community members from the UK blind community, Kerala, and Tamil Nadu helped define what counts as culturally appropriate depictions of artifacts, and the authors tested whether those definitions can be turned into re...
-
ComplexConstraints and Beyond: Expert Rubrics for RLVR
Expert atomic rubrics used as RL rewards lift a 4B model by +15.5 pp on ComplexConstraints and transfer gains to AdvancedIF, MultiChallenge, and agentic tool benchmarks.
-
ComplexConstraints and Beyond: Expert Rubrics for RLVR
Expert-curated rubrics in the new ComplexConstraints dataset improve LLM instruction following by 12-15% when used as RL training signals, with gains transferring to out-of-distribution agentic benchmarks.
-
Evaluating AI-Generated Images of Cultural Artifacts with Community-Informed Rubrics
Case studies with blind UK residents and people from Kerala and Tamil Nadu demonstrate that community input at the systematization stage produces culturally grounded definitions of appropriateness for text-to-image mo...
-
Leveraging Multimodal LLMs for Built Environment and Housing Attribute Assessment from Street-View Imagery
Fine-tuning Gemma 3 27B on modest human-labeled street-view data yields building condition scores that align with and sometimes exceed individual human raters on correlation metrics, with knowledge distillation produc...
Reference graph
Works this paper leans on
-
[1]
Measurement Validity: A Shared Standard for Qualitative and Quantitative Research
Robert Adcock and David Collier. Measurement Validity: A Shared Standard for Qualitative and Quantitative Research. American Political Science Review, 95(3):529–546, September
-
[2]
Chirag Agarwal, Sree Harsha Tanneru, and Himabindu Lakkaraju. Faithfulness vs. Plausibility: On the (Un)Reliability of Explanations from Large Language Models, March 2024. URL ������������������������������� . arXiv:2402.04614 [cs]
arXiv 2024
-
[3]
Open-Source Large Language Models Outperform Crowd Workers and Approach ChatGPT in Text-Annotation Tasks.CoRR, January 2023
Meysam Alizadeh, Maël Kubli, Zeynab Samei, Shirin Dehghani, Juan Diego Bermeo, Maria Korobeynikova, and Fabrizio Gilardi. Open-Source Large Language Models Outperform Crowd Workers and Approach ChatGPT in Text-Annotation Tasks.CoRR, January 2023. URL ������������������������������������������
2023
-
[4]
When Benchmarks are Targets: Revealing the Sensitiv- ity of Large Language Model Leaderboards
Norah Alzahrani, Hisham Alyahya, Yazeed Alnumay, Sultan AlRashed, Shaykhah Alsubaie, Yousef Almushayqih, Faisal Mirza, Nouf Alotaibi, Nora Al-Twairesh, Areeb Alowisheq, M Saiful Bari, and Haidar Khan. When Benchmarks are Targets: Revealing the Sensitiv- ity of Large Language Model Leaderboards. In Lun-Wei Ku, Andre Martins, and Vivek Srikumar, editors, Pr...
-
[5]
Yuntao Bai, Saurav Kadavath, Sandipan Kundu, Amanda Askell, Jackson Kernion, Andy Jones, Anna Chen, Anna Goldie, Azalia Mirhoseini, Cameron McKinnon, Carol Chen, Catherine Olsson, Christopher Olah, Danny Hernandez, Dawn Drain, Deep Ganguli, Dustin Li, Eli Tran-Johnson, Ethan Perez, Jamie Kerr, Jared Mueller, Jeffrey Ladish, Joshua Landau, Kamal Ndousse, K...
arXiv 2022
-
[6]
Anna Bavaresco, Raffaella Bernardi, Leonardo Bertolazzi, Desmond Elliott, Raquel Fernández, Albert Gatt, Esam Ghaleb, Mario Giulianelli, Michael Hanna, Alexander Koller, André F. T. Martins, Philipp Mondorf, Vera Neplenbroek, Sandro Pezzelle, Barbara Plank, David Schlangen, Alessandro Suglia, Aditya K. Surikuchi, Ece Takmaz, and Alberto Testoni. LLMs inst...
2024
-
[8]
Re-evaluating Evaluation in Text Summarization
Manik Bhandari, Pranav Narayan Gour, Atabak Ashfaq, Pengfei Liu, and Graham Neubig. Re-evaluating Evaluation in Text Summarization. In Bonnie Webber, Trevor Cohn, Yulan He, 11 and Yang Liu, editors, Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 9347–9359, Online, November 2020. Association for Comput...
-
[9]
Michael Buhrmester, Tracy Kwang, and Samuel D. Gosling. Amazon’s Mechanical Turk: A New Source of Inexpensive, Yet High-Quality, Data? Perspectives on Psychological Science: A Journal of the Association for Psychological Science , 6(1):3–5, January 2011. ISSN 1745-6916. doi: 10.1177/1745691610393980
Show all 132 references
-
[10]
From Representational Harms to Quality-of-Service Harms: A Case Study on Llama 2 Safety Safeguards
Khaoula Chehbouni, Megha Roshan, Emmanuel Ma, Futian Wei, Afaf Taik, Jackie Cheung, and Golnoosh Farnadi. From Representational Harms to Quality-of-Service Harms: A Case Study on Llama 2 Safety Safeguards. In Lun-Wei Ku, Andre Martins, and Vivek Srikumar, editors, Findings of ...
2024 doi
-
[11]
Beyond the Safety Bundle: Auditing the Helpful and Harmless Dataset
Khaoula Chehbouni, Jonathan Colaço Carr, Yash More, Jackie CK Cheung, and Golnoosh Farnadi. Beyond the Safety Bundle: Auditing the Helpful and Harmless Dataset. In Luis Chiruzzo, Alan Ritter, and Lu Wang, editors, Proceedings of the 2025 Conference of the Nations of the Americ...
2025
-
[12]
Humans or LLMs as the Judge? A Study on Judgement Bias
Guiming Hardy Chen, Shunian Chen, Ziche Liu, Feng Jiang, and Benyou Wang. Humans or LLMs as the Judge? A Study on Judgement Bias. In Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen, editors, Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processi...
2024 doi
-
[13]
Safer or Luckier? LLMs as Safety Evaluators Are Not Robust to Artifacts, March 2025
Hongyu Chen and Seraphina Goldfarb-Tarrant. Safer or Luckier? LLMs as Safety Evaluators Are Not Robust to Artifacts, March 2025. URL ������������������������������� . arXiv:2503.09347 [cs]
2025 arXiv
-
[14]
Cheng-Han Chiang and Hung-yi Lee. Can Large Language Models Be an Alternative to Human Evaluations? In Anna Rogers, Jordan Boyd-Graber, and Naoaki Okazaki, editors, Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers),...
2023 doi
-
[15]
A Closer Look into Using Large Language Models for Automatic Evaluation
Cheng-Han Chiang and Hung-yi Lee. A Closer Look into Using Large Language Models for Automatic Evaluation. In Houda Bouamor, Juan Pino, and Kalika Bali, editors, Findings of the Association for Computational Linguistics: EMNLP 2023, pages 8928–8942, Singapore, December 2023. A...
2023 doi
-
[16]
Angelopoulos, Tianle Li, Dacheng Li, Banghua Zhu, Hao Zhang, Michael I
Wei-Lin Chiang, Lianmin Zheng, Ying Sheng, Anastasios N. Angelopoulos, Tianle Li, Dacheng Li, Banghua Zhu, Hao Zhang, Michael I. Jordan, Joseph E. Gonzalez, and Ion Stoica. Chatbot arena: an open platform for evaluating LLMs by human preference. In Pro- ceedings of the 41st In...
2024
-
[17]
Elizabeth Clark, Tal August, Sofia Serrano, Nikita Haduong, Suchin Gururangan, and Noah A. Smith. All That‘s ‘Human’ Is Not Gold: Evaluating Human Evaluation of Generated Text. In Chengqing Zong, Fei Xia, Wenjie Li, and Roberto Navigli, editors,Proceedings of the 59th Annual M...
2021
-
[18]
Cronbach and Paul E
Lee J. Cronbach and Paul E. Meehl. Construct validity in psychological tests. Psychological Bulletin, 52(4):281–302, 1955. ISSN 1939-1455. doi: 10.1037/h0040957. Place: US Publisher: American Psychological Association
1955 doi
-
[19]
Investigating Data Contamination in Modern Benchmarks for Large Language Models
Chunyuan Deng, Yilun Zhao, Xiangru Tang, Mark Gerstein, and Arman Cohan. Investigating Data Contamination in Modern Benchmarks for Large Language Models. In Kevin Duh, Helena Gomez, and Steven Bethard, editors, Proceedings of the 2024 Conference of the North American Chapter o...
2024 doi
-
[20]
Trends in AI inference energy consumption: Beyond the performance-vs-parameter laws of deep learning
Radosvet Desislavov, Fernando Martínez-Plumed, and José Hernández-Orallo. Trends in AI inference energy consumption: Beyond the performance-vs-parameter laws of deep learning. Sustainable Computing: Informatics and Systems, 38:100857, April 2023. ISSN 2210-5379. doi: 10.1016/j...
2023
-
[21]
Opdahl, and Carl-Gustav Lindén
Laurence Dierickx, Arjen van Dalen, Andreas L. Opdahl, and Carl-Gustav Lindén. Striking the Balance in Using LLMs for Fact-Checking: A Narrative Literature Review. In Mike Preuss, Agata Leszkiewicz, Jean-Christopher Boucher, Ofer Fridman, and Lucas Stampe, editors, Disinformat...
2024 doi
-
[22]
Generaliza- tion or Memorization: Data Contamination and Trustworthy Evaluation for Large Language Models
Yihong Dong, Xue Jiang, Huanyu Liu, Zhi Jin, Bin Gu, Mengfei Yang, and Ge Li. Generaliza- tion or Memorization: Data Contamination and Trustworthy Evaluation for Large Language Models. In Lun-Wei Ku, Andre Martins, and Vivek Srikumar, editors, Findings of the Associa- tion for...
2024 doi
-
[23]
Liang, and Tatsunori B
Yann Dubois, Chen Xuechen Li, Rohan Taori, Tianyi Zhang, Ishaan Gulrajani, Jimmy Ba, Carlos Guestrin, Percy S. Liang, and Tatsunori B. Hashimoto. Al- pacaFarm: A Simulation Framework for Methods that Learn from Human Feed- back. Advances in Neural Information Processing System...
2023
-
[24]
Know Thy Judge: On the Robustness Meta-Evaluation of LLM Safety Judges, March 2025
Francisco Eiras, Eliott Zemour, Eric Lin, and Vaikkunth Mugunthan. Know Thy Judge: On the Robustness Meta-Evaluation of LLM Safety Judges, March 2025. URL ������������� ������������������ . arXiv:2503.04474 [cs]
2025 arXiv
-
[25]
Beyond correlation: The impact of human uncertainty in measuring the effectiveness of automatic evaluation and LLM-as-a-judge
Aparna Elangovan, Lei Xu, Jongwoo Ko, Mahsa Elyasi, Ling Liu, Sravan Babu Bodapati, and Dan Roth. Beyond correlation: The impact of human uncertainty in measuring the effectiveness of automatic evaluation and LLM-as-a-judge. October 2024. URL ������ ���������������������������...
2024
-
[26]
Evaluating the Carbon Impact of Large Language Models at the Inference Stage
Brad Everman, Trevor Villwock, Dayuan Chen, Noe Soto, Oliver Zhang, and Ziliang Zong. Evaluating the Carbon Impact of Large Language Models at the Inference Stage. In 2023 IEEE International Performance, Computing, and Communications Conference (IPCCC) , pages 150–157, Novembe...
2023
-
[27]
Fabbri, Wojciech Kry´sci´nski, Bryan McCann, Caiming Xiong, Richard Socher, and Dragomir Radev
Alexander R. Fabbri, Wojciech Kry´sci´nski, Bryan McCann, Caiming Xiong, Richard Socher, and Dragomir Radev. SummEval: Re-evaluating Summarization Evaluation. Transactions of the Association for Computational Linguistics, 9:391–409, April 2021. ISSN 2307-387X. doi: 10.1162/tac...
2021 doi
-
[28]
Dickerson
Benjamin Feuer, Micah Goldblum, Teresa Datta, Sanjana Nambiar, Raz Besaleli, Samuel Dooley, Max Cembalest, and John P. Dickerson. Style Outweighs Substance: Failure Modes of LLM Judges in Alignment Benchmarking. October 2024. URL ������������������� ����������������������� . 13
2024
-
[29]
When the Majority is Wrong: Modeling Annotator Disagreement for Subjective Tasks
Eve Fleisig, Rediet Abebe, and Dan Klein. When the Majority is Wrong: Modeling Annotator Disagreement for Subjective Tasks. In Houda Bouamor, Juan Pino, and Kalika Bali, editors, Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 6715...
2023 doi
-
[30]
Amazon Mechanical Turk: Gold Mine or Coal Mine ? Computational Linguistics, 37(2):413–420, April 2011
Karen Fort, Gilles Adda, and Kevin Bretonnel Cohen. Amazon Mechanical Turk: Gold Mine or Coal Mine ? Computational Linguistics, 37(2):413–420, April 2011. doi: 10.1162/COLI_ a_00057. URL �������������������������������� . Publisher: Massachusetts Institute of Technology Press ...
2011 doi
-
[31]
Gallegos, Ryan A
Isabel O. Gallegos, Ryan A. Rossi, Joe Barrow, Md Mehrab Tanjim, Sungchul Kim, Franck Dernoncourt, Tong Yu, Ruiyi Zhang, and Nesreen K. Ahmed. Bias and Fairness in Large Language Models: A Survey. Computational Linguistics, 50(3):1097–1179, September 2024. doi: 10.1162/coli_a_...
2024 doi
-
[32]
AEGIS2.0: A Diverse AI Safety Dataset and Risks Taxonomy for Alignment of LLM Guardrails
Shaona Ghosh, Prasoon Varshney, Makesh Narsimhan Sreedhar, Aishwarya Padmakumar, Traian Rebedea, Jibin Rajan Varghese, and Christopher Parisien. AEGIS2.0: A Diverse AI Safety Dataset and Risks Taxonomy for Alignment of LLM Guardrails. In Luis Chiruzzo, Alan Ritter, and Lu Wang...
2025
-
[33]
ChatGPT outperforms crowd workers for text-annotation tasks
Fabrizio Gilardi, Meysam Alizadeh, and Maël Kubli. ChatGPT outperforms crowd workers for text-annotation tasks. Proceedings of the National Academy of Sciences, 120(30):e2305016120, July 2023. doi: 10.1073/pnas.2305016120. URL �������������������������������� �����������������...
2023 doi
-
[34]
Time Travel in LLMs: Tracing Data Contamination in Large Language Models
Shahriar Golchin and Mihai Surdeanu. Time Travel in LLMs: Tracing Data Contamination in Large Language Models. October 2023. URL �������������������������������� ����������
2023
-
[35]
Kivlichan, Rachel Rosen, and Lucy Vasserman
Nitesh Goyal, Ian D. Kivlichan, Rachel Rosen, and Lucy Vasserman. Is Your Toxicity My Toxicity? Exploring the Impact of Rater Identity on Toxicity Annotation. Proc. ACM Hum.- Comput. Interact., 6(CSCW2):363:1–363:28, November 2022. doi: 10.1145/3555088. URL �������������������...
2022 doi
-
[36]
Validating LLM-as-a-Judge Systems in the Absence of Gold Labels, March 2025
Luke Guerdan, Solon Barocas, Kenneth Holstein, Hanna Wallach, Zhiwei Steven Wu, and Alexandra Chouldechova. Validating LLM-as-a-Judge Systems in the Absence of Gold Labels, March 2025. URL ����������������������������������
2025
-
[37]
WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs, December 2024
Seungju Han, Kavel Rao, Allyson Ettinger, Liwei Jiang, Bill Yuchen Lin, Nathan Lambert, Yejin Choi, and Nouha Dziri. WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs, December 2024. URL �������������������������� �����. arXiv:2406.18495 [cs]
2024 arXiv
-
[38]
AnnoLLM: Making Large Language Models to Be Better Crowdsourced Annotators
Xingwei He, Zhenghao Lin, Yeyun Gong, A-Long Jin, Hang Zhang, Chen Lin, Jian Jiao, Siu Ming Yiu, Nan Duan, and Weizhu Chen. AnnoLLM: Making Large Language Models to Be Better Crowdsourced Annotators. In Yi Yang, Aida Davani, Avi Sil, and Anoop Kumar, edi- tors, Proceedings of ...
2024
-
[39]
If in a Crowdsourced Data Annotation Pipeline, a GPT-4
Zeyu He, Chieh-Yang Huang, Chien-Kuang Cornelia Ding, Shaurya Rohatgi, and Ting- Hao Kenneth Huang. If in a Crowdsourced Data Annotation Pipeline, a GPT-4. In Pro- ceedings of the 2024 CHI Conference on Human Factors in Computing Systems , CHI ’24, pages 1–25, New York, NY , U...
2024
-
[40]
Salim, and Mark Sanderson
Danula Hettiachchi, Indigo Holcombe-James, Stephanie Livingstone, Anjalee de Silva, Matthew Lease, Flora D. Salim, and Mark Sanderson. How Crowd Worker Factors Influ- ence Subjective Annotations: A Study of Tagging Misogynistic Hate Speech in Tweets. Proceedings of the AAAI Co...
2023 doi
-
[41]
Large Language Models are Zero-Shot Rankers for Recommender Systems
Yupeng Hou, Junjie Zhang, Zihan Lin, Hongyu Lu, Ruobing Xie, Julian McAuley, and Wayne Xin Zhao. Large Language Models are Zero-Shot Rankers for Recommender Systems. In Advances in Information Retrieval: 46th European Conference on Information Retrieval, ECIR 2024, Glasgow, UK...
2024 doi
-
[42]
Howcroft, Anya Belz, Miruna-Adriana Clinciu, Dimitra Gkatzia, Sadid A
David M. Howcroft, Anya Belz, Miruna-Adriana Clinciu, Dimitra Gkatzia, Sadid A. Hasan, Saad Mahamood, Simon Mille, Emiel van Miltenburg, Sashank Santhanam, and Verena Rieser. Twenty Years of Confusion in Human Evaluation: NLG Needs Evaluation Sheets and Standardised Definition...
2020
-
[43]
Xinyu Hu, Mingqi Gao, Sen Hu, Yang Zhang, Yicheng Chen, Teng Xu, and Xiaojun Wan. Are LLM-based Evaluators Confusing NLG Quality Criteria? In Lun-Wei Ku, Andre Martins, and Vivek Srikumar, editors,Proceedings of the 62nd Annual Meeting of the Association for Computational Ling...
2024 doi
-
[44]
Is ChatGPT better than Human Annotators? Potential and Limitations of ChatGPT in Explaining Implicit Hate Speech
Fan Huang, Haewoon Kwak, and Jisun An. Is ChatGPT better than Human Annotators? Potential and Limitations of ChatGPT in Explaining Implicit Hate Speech. In Companion Proceedings of the ACM Web Conference 2023, WWW ’23 Companion, pages 294–297, New York, NY , USA, 2023. Associa...
2023
-
[45]
Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations, December 2023
Hakan Inan, Kartikeya Upasani, Jianfeng Chi, Rashi Rungta, Krithika Iyer, Yuning Mao, Michael Tontchev, Qing Hu, Brian Fuller, Davide Testuggine, and Madian Khabsa. Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations, December 2023. URL �������������������...
2023 arXiv
-
[46]
Data Janitors
Lilly Irani. Justice for “Data Janitors”, January 2015. URL ������������������������ ������������������������������
2015
-
[47]
Jacobs and Hanna Wallach
Abigail Z. Jacobs and Hanna Wallach. Measurement and Fairness. In Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency , FAccT ’21, pages 375–385, New York, NY , USA, March 2021. Association for Computing Machinery. ISBN 978-1-4503-8309-7. doi: ...
2021
-
[48]
CritiqueLLM: Towards an Informative Critique Generation Model for Evaluation of Large Language Model Generation
Pei Ke, Bosi Wen, Andrew Feng, Xiao Liu, Xuanyu Lei, Jiale Cheng, Shengyuan Wang, Aohan Zeng, Yuxiao Dong, Hongning Wang, Jie Tang, and Minlie Huang. CritiqueLLM: Towards an Informative Critique Generation Model for Evaluation of Large Language Model Generation. In Lun-Wei Ku,...
2024
-
[49]
Crowd-calibrator: Can annotator disagreement inform calibration in subjective tasks? In First Conference on Language Modeling, 2024
Urja Khurana, Eric Nalisnick, Antske Fokkens, and Swabha Swayamdipta. Crowd-calibrator: Can annotator disagreement inform calibration in subjective tasks? In First Conference on Language Modeling, 2024. URL ������������������������������������������
2024
-
[50]
Large Language Models Are State-of-the-Art Evalua- tors of Translation Quality
Tom Kocmi and Christian Federmann. Large Language Models Are State-of-the-Art Evalua- tors of Translation Quality. In Mary Nurminen, Judith Brenner, Maarit Koponen, Sirkku Latomaa, Mikhail Mikhailov, Frederike Schierl, Tharindu Ranasinghe, Eva Vanmassen- hove, Sergi Alvarez Vi...
2023
-
[51]
Benchmarking Cognitive Biases in Large Language Models as Evaluators
Ryan Koo, Minhwa Lee, Vipul Raheja, Jong Inn Park, Zae Myung Kim, and Dongyeop Kang. Benchmarking Cognitive Biases in Large Language Models as Evaluators. In Lun-Wei Ku, Andre Martins, and Vivek Srikumar, editors, Findings of the Association for Computational Linguistics: ACL ...
2024 doi
-
[52]
From Generation to Judgment: Opportunities and Challenges of LLM-as-a-judge, February 2025
Dawei Li, Bohan Jiang, Liangjie Huang, Alimohammad Beigi, Chengshuai Zhao, Zhen Tan, Amrita Bhattacharjee, Yuxuan Jiang, Canyu Chen, Tianhao Wu, Kai Shu, Lu Cheng, and Huan Liu. From Generation to Judgment: Opportunities and Challenges of LLM-as-a-judge, February 2025. URL ���...
2025
-
[53]
Preference Leakage: A Contamination Problem in LLM- as-a-judge, February 2025
Dawei Li, Renliang Sun, Yue Huang, Ming Zhong, Bohan Jiang, Jiawei Han, Xiangliang Zhang, Wei Wang, and Huan Liu. Preference Leakage: A Contamination Problem in LLM- as-a-judge, February 2025. URL ������������������������������� . arXiv:2502.01534 [cs]
2025
-
[54]
CalibraEval: Calibrating Prediction Distribution to Mitigate Selection Bias in LLMs-as-Judges, October 2024
Haitao Li, Junjie Chen, Qingyao Ai, Zhumin Chu, Yujia Zhou, Qian Dong, and Yiqun Liu. CalibraEval: Calibrating Prediction Distribution to Mitigate Selection Bias in LLMs-as-Judges, October 2024. URL ������������������������������� . arXiv:2410.15393 [cs]
2024 arXiv
-
[55]
LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods, December 2024
Haitao Li, Qian Dong, Junjie Chen, Huixue Su, Yujia Zhou, Qingyao Ai, Ziyi Ye, and Yiqun Liu. LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods, December 2024. URL ������������������������������� . arXiv:2412.05579 [cs]
2024 arXiv
-
[56]
Generative Judge for Evaluating Alignment
Junlong Li, Shichao Sun, Weizhe Yuan, Run-Ze Fan, Hai Zhao, and Pengfei Liu. Generative Judge for Evaluating Alignment. October 2023. URL ����������������������������� �������������
2023
-
[57]
Dissect- ing Human and LLM Preferences
Junlong Li, Fan Zhou, Shichao Sun, Yikai Zhang, Hai Zhao, and Pengfei Liu. Dissect- ing Human and LLM Preferences. In Lun-Wei Ku, Andre Martins, and Vivek Srikumar, editors, Proceedings of the 62nd Annual Meeting of the Association for Computational Lin- guistics (Volume 1: Lo...
2024 doi
-
[58]
SALAD-Bench: A Hierarchical and Comprehensive Safety Benchmark for Large Language Models
Lijun Li, Bowen Dong, Ruohui Wang, Xuhao Hu, Wangmeng Zuo, Dahua Lin, Yu Qiao, and Jing Shao. SALAD-Bench: A Hierarchical and Comprehensive Safety Benchmark for Large Language Models. In Lun-Wei Ku, Andre Martins, and Vivek Srikumar, editors,Findings of the Association for Com...
2024 doi
-
[59]
Islam, and Shaolei Ren
Pengfei Li, Jianyi Yang, Mohammad A. Islam, and Shaolei Ren. Making AI Less "Thirsty": Uncovering and Addressing the Secret Water Footprint of AI Models, March 2025. URL ������������������������������� . arXiv:2304.03271 [cs]. 16
2025 arXiv
-
[60]
Gonzalez, and Ion Stoica
Tianle Li, Wei-Lin Chiang, Evan Frick, Lisa Dunlap, Tianhao Wu, Banghua Zhu, Joseph E. Gonzalez, and Ion Stoica. From Crowdsourced Data to High-Quality Benchmarks: Arena- Hard and BenchBuilder Pipeline, June 2024. URL ������������������������������� . arXiv:2406.11939 [cs] version: 1
2024 arXiv
-
[61]
Split and Merge: Aligning Position Biases in LLM-based Evaluators
Zongjie Li, Chaozheng Wang, Pingchuan Ma, Daoyuan Wu, Shuai Wang, Cuiyun Gao, and Yang Liu. Split and Merge: Aligning Position Biases in LLM-based Evaluators. In Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen, editors, Proceedings of the 2024 Conference on Empirical Methods...
2024 doi
-
[62]
The Unlocking Spell on Base LLMs: Rethink- ing Alignment via In-Context Learning
Bill Yuchen Lin, Abhilasha Ravichander, Ximing Lu, Nouha Dziri, Melanie Sclar, Khyathi Chandu, Chandra Bhagavatula, and Yejin Choi. The Unlocking Spell on Base LLMs: Rethink- ing Alignment via In-Context Learning. October 2023. URL ����������������������� ���������������������...
2023
-
[63]
G- Eval: NLG Evaluation using Gpt-4 with Better Human Alignment
Yang Liu, Dan Iter, Yichong Xu, Shuohang Wang, Ruochen Xu, and Chenguang Zhu. G- Eval: NLG Evaluation using Gpt-4 with Better Human Alignment. In Houda Bouamor, Juan Pino, and Kalika Bali, editors, Proceedings of the 2023 Conference on Empirical Methods in Natural Language Pro...
2023 doi
-
[64]
LLMs as Narcissistic Evaluators: When Ego Inflates Evaluation Scores
Yiqi Liu, Nafise Moosavi, and Chenghua Lin. LLMs as Narcissistic Evaluators: When Ego Inflates Evaluation Scores. In Lun-Wei Ku, Andre Martins, and Vivek Srikumar, edi- tors, Findings of the Association for Computational Linguistics: ACL 2024 , pages 12688– 12701, Bangkok, Tha...
2024 doi
-
[65]
LLM Comparative Assessment: Zero-shot NLG Evaluation through Pairwise Comparisons using Large Language Models
Adian Liusie, Potsawee Manakul, and Mark Gales. LLM Comparative Assessment: Zero-shot NLG Evaluation through Pairwise Comparisons using Large Language Models. In Yvette Graham and Matthew Purver, editors, Proceedings of the 18th Conference of the European Chapter of the Associ...
2024
-
[66]
PrimeGuard: Safe and Helpful LLMs through Tuning-Free Routing, July 2024
Blazej Manczak, Eliott Zemour, Eric Lin, and Vaikkunth Mugunthan. PrimeGuard: Safe and Helpful LLMs through Tuning-Free Routing, July 2024. URL ��������������������� ���������� . arXiv:2407.16318 [cs]
2024 arXiv
-
[67]
Marshall, Partha S.R
Catherine C. Marshall, Partha S.R. Goguladinne, Mudit Maheshwari, Apoorva Sathe, and Frank M. Shipman. Who Broke Amazon Mechanical Turk? An Analysis of Crowd- sourcing Data Quality over Time. In Proceedings of the 15th ACM Web Science Con- ference 2023, WebSci ’23, pages 335–3...
2023
-
[68]
HarmBench: a standardized evaluation framework for automated red teaming and robust refusal
Mantas Mazeika, Long Phan, Xuwang Yin, Andy Zou, Zifan Wang, Norman Mu, Elham Sakhaee, Nathaniel Li, Steven Basart, Bo Li, David Forsyth, and Dan Hendrycks. HarmBench: a standardized evaluation framework for automated red teaming and robust refusal. In Pro- ceedings of the 41s...
2024
-
[69]
USR: An Unsupervised and Reference Free Evaluation Metric for Dialog Generation
Shikib Mehri and Maxine Eskenazi. USR: An Unsupervised and Reference Free Evaluation Metric for Dialog Generation. In Dan Jurafsky, Joyce Chai, Natalie Schluter, and Joel Tetreault, editors,Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics...
2020 doi
-
[70]
Are large language models good annota- tors? In Proceedings on, pages 38–48
Jay Mohta, Kenan Ak, Yan Xu, and Mingwei Shen. Are large language models good annota- tors? In Proceedings on, pages 38–48. PMLR, April 2023. URL �������������������� ���������������������������� . ISSN: 2640-3498
2023
-
[71]
Automated evaluation of written dis- course coherence using GPT-4
Ben Naismith, Phoebe Mulcaire, and Jill Burstein. Automated evaluation of written dis- course coherence using GPT-4. In Ekaterina Kochmar, Jill Burstein, Andrea Horbach, Ronja Laarmann-Quante, Nitin Madnani, Anaïs Tack, Victoria Yaneva, Zheng Yuan, and Torsten Zesch, editors, ...
2023
-
[72]
ChatGPT Label: Comparing the Quality of Human- Generated and LLM-Generated Annotations in Low-Resource Language NLP Tasks
Arbi Haza Nasution and Aytu˘g Onan. ChatGPT Label: Comparing the Quality of Human- Generated and LLM-Generated Annotations in Low-Resource Language NLP Tasks. IEEE Access, 12:71876–71900, 2024. ISSN 2169-3536. doi: 10.1109/ACCESS.2024.3402809. URL �����������������������������...
2024
-
[73]
GuardFormer: Guardrail Instruction Pretraining for Efficient SafeGuarding
James O’ Neill, Santhosh Subramanian, Eric Lin, Abishek Satish, and Vaikkunth Mugunthan. GuardFormer: Guardrail Instruction Pretraining for Efficient SafeGuarding. October 2024. URL ������������������������������������������
2024
-
[74]
Evaluating Content Selection in Summarization: The Pyramid Method
Ani Nenkova and Rebecca Passonneau. Evaluating Content Selection in Summarization: The Pyramid Method. In Proceedings of the Human Language Technology Conference of the North American Chapter of the Association for Computational Linguistics: HLT-NAACL 2004, pages 145–152, Bost...
2004
-
[75]
Will Orr and Edward B. Kang. AI as a Sport: On the Competitive Epistemologies of Bench- marking. In Proceedings of the 2024 ACM Conference on Fairness, Accountability, and Transparency, FAccT ’24, pages 1875–1884, New York, NY , USA, 2024. Association for Computing Machinery. ...
2024
-
[76]
Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul Christiano, Jan Leike, and...
2022 arXiv
-
[77]
Bowman, and Shi Feng
Arjun Panickssery, Samuel R. Bowman, and Shi Feng. LLM Evaluators Recognize and Favor Their Own Generations. November 2024. URL �������������������������������� ����������
2024
-
[78]
Users Favor LLM-Generated Content – Until They Know It’s AI, February 2025
Petr Parshakov, Iuliia Naidenova, Sofia Paklina, Nikita Matkin, and Cornel Nesseler. Users Favor LLM-Generated Content – Until They Know It’s AI, February 2025. URL ����� �������������������������� . arXiv:2503.16458 [cs]
2025 arXiv
-
[79]
Red Teaming Language Models with Language Models
Ethan Perez, Saffron Huang, Francis Song, Trevor Cai, Roman Ring, John Aslanides, Amelia Glaese, Nat McAleese, and Geoffrey Irving. Red Teaming Language Models with Language Models. In Yoav Goldberg, Zornitsa Kozareva, and Yue Zhang, editors,Proceedings of the 2022 Conference ...
2022 doi
-
[80]
Amazon’s Mechanical Turk a Digital Sweatshop? Transparency and Accountability in Crowdsourced Online Research
Matthew Pittman, , and Kim Sheehan. Amazon’s Mechanical Turk a Digital Sweatshop? Transparency and Accountability in Crowdsourced Online Research. Journal of Media Ethics, 31(4):260–262, October 2016. ISSN 2373-6992. doi: 10.1080/23736992.2016.1228811. URL ��������������������...
2016
-
[81]
Is LLM-as-a-Judge Robust? Investigating Universal Adversarial Attacks on Zero-shot LLM Assessment
Vyas Raina, Adian Liusie, and Mark Gales. Is LLM-as-a-Judge Robust? Investigating Universal Adversarial Attacks on Zero-shot LLM Assessment. In Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen, editors, Proceedings of the 2024 Conference on Empirical Methods in Natural Langua...
2024
-
[82]
Meta’s Llama 4 ’herd’ controversy and AI contamination, explained, April
Tiernan Ray. Meta’s Llama 4 ’herd’ controversy and AI contamination, explained, April
-
[83]
To the Cutoff
Manley Roberts, Himanshu Thakur, Christine Herlihy, Colin White, and Samuel Dooley. To the Cutoff... and Beyond? A Longitudinal Perspective on LLM Data Contamination. October
-
[84]
XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models
Paul Röttger, Hannah Kirk, Bertie Vidgen, Giuseppe Attanasio, Federico Bianchi, and Dirk Hovy. XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models. In Kevin Duh, Helena Gomez, and Steven Bethard, editors, Proceedings of the 2024 Conferen...
2024
-
[85]
From Words to Watts: Benchmarking the Energy Costs of Large Language Model Inference
Siddharth Samsi, Dan Zhao, Joseph McDonald, Baolin Li, Adam Michaleas, Michael Jones, William Bergeron, Jeremy Kepner, Devesh Tiwari, and Vijay Gadepally. From Words to Watts: Benchmarking the Energy Costs of Large Language Model Inference. In 2023 IEEE High Performance Extrem...
2023
-
[86]
doi: 10.18653/v1/2024.emnlp-main.427
Association for Computational Linguistics. doi: 10.18653/v1/2024.emnlp-main.427. URL ���������������������������������������������
2024 doi
-
[87]
Targeting the Benchmark: On Methodology in Current Natural Language Processing Research
David Schlangen. Targeting the Benchmark: On Methodology in Current Natural Language Processing Research. In Chengqing Zong, Fei Xia, Wenjie Li, and Roberto Navigli, editors, Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th I...
2021 doi
-
[88]
URL ������������������������������������������������������������� �������������������������������
-
[89]
Optimization-based Prompt Injection Attack to LLM-as-a-Judge
Jiawen Shi, Zenghui Yuan, Yinuo Liu, Yue Huang, Pan Zhou, Lichao Sun, and Neil Zhenqiang Gong. Optimization-based Prompt Injection Attack to LLM-as-a-Judge. In Proceedings of the 2024 on ACM SIGSAC Conference on Computer and Communications Security , CCS ’24, pages 660–674, Ne...
2024
-
[90]
URL ������������������������������������������
-
[91]
Silva, Rohan Khera, and Lee H
Gisele S. Silva, Rohan Khera, and Lee H. Schwamm. Reviewer Experience Detecting and Judging Human Versus Artificial Intelligence Content: The Stroke Journal Essay Contest. Stroke, 55(10):2573–2578, October 2024. ISSN 1524-4628. doi: 10.1161/STROKEAHA.124. 045012
2024 doi
-
[92]
The Leaderboard Illusion, April 2025
Shivalika Singh, Yiyang Nan, Alex Wang, Daniel D’Souza, Sayash Kapoor, Ahmet Üstün, Sanmi Koyejo, Yuntian Deng, Shayne Longpre, Noah Smith, Beyza Ermis, Marzieh Fadaee, and Sara Hooker. The Leaderboard Illusion, April 2025. URL ��������������������� ���������� . arXiv:2504.20879 [cs]
2025 arXiv
-
[93]
Privacy, Power, and Invisible Labor on Amazon Mechanical Turk
Shruti Sannon and Dan Cosley. Privacy, Power, and Invisible Labor on Amazon Mechanical Turk. In Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems, CHI ’19, pages 1–12, New York, NY , USA, 2019. Association for Computing Machinery. ISBN 978-1-4503-597...
2019
-
[94]
BERTScore is Unfair: On Social Bias in Language Model-Based Metrics for Text Generation
Tianxiang Sun, Junliang He, Xipeng Qiu, and Xuanjing Huang. BERTScore is Unfair: On Social Bias in Language Model-Based Metrics for Text Generation. In Yoav Goldberg, Zornitsa Kozareva, and Yue Zhang, editors,Proceedings of the 2022 Conference on Empirical Methods in Natural L...
2022 doi
-
[95]
Societal Biases in Language Generation: Progress and Challenges
Emily Sheng, Kai-Wei Chang, Prem Natarajan, and Nanyun Peng. Societal Biases in Language Generation: Progress and Challenges. In Chengqing Zong, Fei Xia, Wenjie Li, and Roberto Navigli, editors, Proceedings of the 59th Annual Meeting of the Associ- ation for Computational Ling...
-
[96]
Principle-Driven Self-Alignment of Language Models from Scratch with Minimal Human Supervision
Zhiqing Sun, Yikang Shen, Qinhong Zhou, Hongxin Zhang, Zhenfang Chen, David Daniel Cox, Yiming Yang, and Chuang Gan. Principle-Driven Self-Alignment of Language Models from Scratch with Minimal Human Supervision. November 2023. URL ������������������� �������������������������...
2023
-
[97]
Meta llama guard 2
Llama Team. Meta llama guard 2. ������������������������������������������ ������������������������������������ , 2024
2024
-
[98]
Large Language Models Help Humans Verify Truthfulness – Except When They Are Convincingly Wrong
Chenglei Si, Navita Goyal, Tongshuang Wu, Chen Zhao, Shi Feng, Hal Daumé Iii, and Jordan Boyd-Graber. Large Language Models Help Humans Verify Truthfulness – Except When They Are Convincingly Wrong. In Kevin Duh, Helena Gomez, and Steven Bethard, editors, 19 Proceedings of the...
2024
-
[99]
Beware the Hype: ChatGPT Didn’t Replace Human Data Annotators, April
Turkopticon. Beware the Hype: ChatGPT Didn’t Replace Human Data Annotators, April
-
[100]
Large Language Models Outperform Expert Coders and Supervised Classi- fiers at Annotating Political Social Media Messages
Petter Törnberg. Large Language Models Outperform Expert Coders and Supervised Classi- fiers at Annotating Political Social Media Messages. Sage Journals, September 2024. doi: 10.1177/08944393241286471. URL ����������������������������������������� ����������������� . Publishe...
2024 doi
-
[101]
Cheap and Fast – But is it Good? Evaluating Non-Expert Annotations for Natural Language Tasks
Rion Snow, Brendan O’Connor, Daniel Jurafsky, and Andrew Ng. Cheap and Fast – But is it Good? Evaluating Non-Expert Annotations for Natural Language Tasks. In Mirella Lapata and Hwee Tou Ng, editors, Proceedings of the 2008 Conference on Empirical Methods in Natural Language P...
2008
-
[102]
Feder Cooper, Angelina Wang, Solon Barocas, Alexandra Chouldechova, Chad Atalla, Su Lin Blodgett, Emily Corvi, P
Hanna Wallach, Meera Desai, Nicholas Pangakis, A. Feder Cooper, Angelina Wang, Solon Barocas, Alexandra Chouldechova, Chad Atalla, Su Lin Blodgett, Emily Corvi, P. Alex Dow, Jean Garcia-Gathright, Alexandra Olteanu, Stefanie Reed, Emily Sheng, Dan Vann, Jennifer Wortman Vaugha...
-
[103]
SALMON: Self-Alignment with Instructable Reward Mod- els
Zhiqing Sun, Yikang Shen, Hongxin Zhang, Qinhong Zhou, Zhenfang Chen, David Daniel Cox, Yiming Yang, and Chuang Gan. SALMON: Self-Alignment with Instructable Reward Mod- els. October 2023. URL �������������������������������������������������� ����������
2023
-
[104]
Large Language Models are not Fair Evaluators
Peiyi Wang, Lei Li, Liang Chen, Zefan Cai, Dawei Zhu, Binghuai Lin, Yunbo Cao, Lingpeng Kong, Qi Liu, Tianyu Liu, and Zhifang Sui. Large Language Models are not Fair Evaluators. In Lun-Wei Ku, Andre Martins, and Vivek Srikumar, editors,Proceedings of the 62nd Annual Meeting of...
2024 doi
-
[105]
Dai, and Quoc V
Jason Wei, Maarten Bosma, Vincent Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M. Dai, and Quoc V . Le. Finetuned Language Models are Zero-Shot Learners. October 2021. URL �������������������������������������������
2021
-
[106]
Measuring the Occupational Impact of AI: Tasks, Cognitive Abilities and AI Benchmarks
Songül Tolan, Annarosa Pesole, Fernando Martínez-Plumed, Enrique Fernández-Macías, José Hernández-Orallo, and Emilia Gómez. Measuring the Occupational Impact of AI: Tasks, Cognitive Abilities and AI Benchmarks. Journal of Artificial Intelligence Research, 71:191– 236, June 202...
2021 doi
-
[107]
Chi, Quoc V
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed H. Chi, Quoc V . Le, and Denny Zhou. Chain-of-thought prompting elicits reasoning in large language models. In Proceedings of the 36th International Conference on Neural Information Processing Sy...
2022
-
[108]
URL ���������������������������������������������������������
-
[109]
ChatGPT Can Replace the Underpaid Workers Who Train AI, Researchers Say, March 2023
Chloe Xiang. ChatGPT Can Replace the Underpaid Workers Who Train AI, Researchers Say, March 2023. URL �������������������������������������������������������� �����������������������������������������������
2023
-
[110]
Replacing Judges with Juries: Evaluating LLM Generations with a Panel of Diverse Models, May 2024
Pat Verga, Sebastian Hofstatter, Sophia Althammer, Yixuan Su, Aleksandra Piktus, Arkady Arkhangorodsky, Minjie Xu, Naomi White, and Patrick Lewis. Replacing Judges with Juries: Evaluating LLM Generations with a Panel of Diverse Models, May 2024. URL ���������������������������...
2024 arXiv
-
[111]
Vera Liao
Ziang Xiao, Susu Zhang, Vivian Lai, and Q. Vera Liao. Evaluating Evaluation Metrics: A Framework for Analyzing NLG Evaluation Metrics using Measurement Theory. In Houda 21 Bouamor, Juan Pino, and Kalika Bali, editors, Proceedings of the 2023 Conference on Em- pirical Methods i...
2023
- [112]
-
[113]
Is ChatGPT a Good NLG Evaluator? A Preliminary Study
Jiaan Wang, Yunlong Liang, Fandong Meng, Zengkui Sun, Haoxiang Shi, Zhixu Li, Jinan Xu, Jianfeng Qu, and Jie Zhou. Is ChatGPT a Good NLG Evaluator? A Preliminary Study. In Yue Dong, Wen Xiao, Lu Wang, Fei Liu, and Giuseppe Carenini, editors, Proceedings of the 4th New Frontier...
-
[114]
doi: 10.18653/v1/2023.newsum-1.1
Association for Computational Linguistics. doi: 10.18653/v1/2023.newsum-1.1. URL �����������������������������������������
2023 doi
-
[115]
Dora Zhao, Jerone T. A. Andrews, Orestis Papakyriakopoulos, and Alice Xiang. Position: measure dataset diversity, don’t just claim it. In Proceedings of the 41st International Confer- ence on Machine Learning, volume 235 of ICML’24, pages 60644–60673, Vienna, Austria,
-
[116]
Large Lan- guage Models Are Not Robust Multiple Choice Selectors
Chujie Zheng, Hao Zhou, Fandong Meng, Jie Zhou, and Minlie Huang. Large Lan- guage Models Are Not Robust Multiple Choice Selectors. October 2023. URL ������ ������������������������������������
2023
-
[117]
Chi, Tatsunori Hashimoto, Oriol Vinyals, Percy Liang, Jeff Dean, and William Fedus
Jason Wei, Yi Tay, Rishi Bommasani, Colin Raffel, Barret Zoph, Sebastian Borgeaud, Dani Yogatama, Maarten Bosma, Denny Zhou, Donald Metzler, Ed H. Chi, Tatsunori Hashimoto, Oriol Vinyals, Percy Liang, Jeff Dean, and William Fedus. Emergent Abilities of Large Language Models. T...
2022
-
[118]
Cheating Automatic LLM Benchmarks: Null Models Achieve High Win Rates
Xiaosen Zheng, Tianyu Pang, Chao Du, Qian Liu, Jing Jiang, and Min Lin. Cheating Automatic LLM Benchmarks: Null Models Achieve High Win Rates. October 2024. URL ������������������������������������������
2024
-
[119]
Taxonomy of Risks posed by Language Models
Laura Weidinger, View Profile, Jonathan Uesato, View Profile, Maribeth Rauh, View Profile, Conor Griffin, Search about this author, Po-Sen Huang, View Profile, John Mellor, View Profile, Amelia Glaese, View Profile, Myra Cheng, View Profile, Borja Balle, Search about this auth...
2022
-
[120]
Deconstructing NLG Evaluation: Evaluation Practices, Assumptions, and Their Implications
Kaitlyn Zhou, Su Lin Blodgett, Adam Trischler, Hal Daumé III, Kaheer Suleman, and Alexan- dra Olteanu. Deconstructing NLG Evaluation: Evaluation Practices, Assumptions, and Their Implications. In Marine Carpuat, Marie-Catherine de Marneffe, and Ivan Vladimir Meza Ruiz, editors...
2022
-
[121]
OpenAI Used Kenyan Workers Making $2 an Hour to Filter Traumatic Content from ChatGPT, January 2023
Chloe Xiang. OpenAI Used Kenyan Workers Making $2 an Hour to Filter Traumatic Content from ChatGPT, January 2023. URL�������������������������������������������� ������������������������������������������������������������������ �������������
2023
-
[123]
doi: 10.18653/v1/2023.emnlp-main.676
Association for Computational Linguistics. doi: 10.18653/v1/2023.emnlp-main.676. URL ���������������������������������������������
2023 doi
-
[124]
Chawla, and Xiangliang Zhang
Jiayi Ye, Yanbo Wang, Yue Huang, Dongping Chen, Qihui Zhang, Nuno Moniz, Tian Gao, Werner Geyer, Chao Huang, Pin-Yu Chen, Nitesh V . Chawla, and Xiangliang Zhang. Justice or Prejudice? Quantifying Biases in LLM-as-a-Judge. October 2024. URL �������� ����������������������������������
2024
-
[125]
ShieldGemma: Generative AI Content Moderation Based on Gemma, August 2024
Wenjun Zeng, Yuchi Liu, Ryan Mullins, Ludovic Peran, Joe Fernandez, Hamza Harkous, Karthik Narasimhan, Drew Proud, Piyush Kumar, Bhaktipriya Radharapu, Olivia Sturman, and Oscar Wahltinez. ShieldGemma: Generative AI Content Moderation Based on Gemma, August 2024. URL ���������...
2024 arXiv
-
[126]
Human favoritism, not AI aversion: People’s perceptions (and bias) toward generative AI, human experts, and human–GAI collaboration in persuasive content generation
Yunhao Zhang and Renée Gosline. Human favoritism, not AI aversion: People’s perceptions (and bias) toward generative AI, human experts, and human–GAI collaboration in persuasive content generation. Judgment and Decision Mak- ing, 18:e41, January 2023. ISSN 1930-2975. doi: 10.1...
2023 doi
-
[129]
Xing, Hao Zhang, Joseph E
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric P. Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica. Judging LLM-as-a-judge with MT-bench and Chatbot Arena. In Proceedings of the 37th International ...
2023
-
[131]
LIMA: Less Is More for Alignment
Chunting Zhou, Pengfei Liu, Puxin Xu, Srini Iyer, Jiao Sun, Yuning Mao, Xuezhe Ma, Avia Efrat, Ping Yu, Lili Yu, Susan Zhang, Gargi Ghosh, Mike Lewis, Luke Zettlemoyer, and Omer Levy. LIMA: Less Is More for Alignment. November 2023. URL ������ ������������������������������������
2023
-
[133]
Can ChatGPT Reproduce Human-Generated Labels? A Study of Social Computing Tasks, April 2023
Yiming Zhu, Peixian Zhang, Ehsan-Ul Haq, Pan Hui, and Gareth Tyson. Can ChatGPT Reproduce Human-Generated Labels? A Study of Social Computing Tasks, April 2023. URL ������������������������������� . arXiv:2304.10145 [cs]. 22
2023 arXiv
-
[2001]
doi: 10.1017/S0003055401003100
ISSN 0003-0554, 1537-5943. doi: 10.1017/S0003055401003100. URL ������ �������������������������������������������������������������������� ������������������������������������������������������������������� �����������������������������������������������������������
-
[2021]
doi: 10.18653/v1/2021.acl-long.330
Association for Computational Linguistics. doi: 10.18653/v1/2021.acl-long.330. URL �������������������������������������������
2021 doi
-
[2023]
doi: 10.18653/v1/2023.bea-1.32
Association for Computational Linguistics. doi: 10.18653/v1/2023.bea-1.32. URL ���������������������������������������
2023 doi
-
[2024]
doi: 10.18653/v1/2024.acl-long.744
Association for Computational Linguistics. doi: 10.18653/v1/2024.acl-long.744. URL �������������������������������������������
2024 doi
-
[2025]
ISBN 9798891761896
Association for Computational Linguistics. ISBN 9798891761896. URL ������ ���������������������������������������
Reviewed August 5, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.