REVIEW 3 major objections 5 minor 2 cited by
The Science of Evaluating Foundation Models
T0 review · 3 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read LLM evaluation should start from the use case, not from generic benchmarks, and the paper turns this into an ABCD checklist with pruning and documentation stages.
desk verdict A readable checklist and survey with an overstated 'no actionable guideline exists' claim that its own references contradict; fix the framing and it's worth a serious referee. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central mechanism is the ABCD framework (Algorithm, Big Data, Computation Resources, Domain Expertise) used as an organizing alphabet for evaluation, together with the three-stage workflow it feeds: a preparation checklist, applicability analysis, and documentation. The checklist maps each preparation step to the relevant ABCD letter, so choices about models, datasets, metrics, baselines, ethics and safety, and resources are made explicit before experiments begin. The applicability-analysis stage supplies the paper's main operational idea: not every evaluation dimension is needed for every task, so evaluators should assign relative weights to dimensions and prune to what is feasible, then disclose those weights in documentation. This weighting-and-disclosure mechanism is what converts the framework from a taxonomy into a decision procedure.
What would settle it
Search the literature and practitioner tooling for an existing, widely available evaluation process that already includes step-by-step instructions for defining objectives, selecting datasets and metrics, setting baselines, addressing ethics and safety, allocating resources, and documenting results; finding one that practitioners can follow end-to-end in a new domain would falsify the paper's claim that no such actionable guideline exists.
Extended reading notes
Core claim
The claim is that there is no actionable evaluation guideline incorporating a cohesive process for large language models, and that evaluation should be driven by use-case context rather than generic leaderboards. The paper formalizes the process with the ABCD framework: Algorithm covers model choices and baselines, Big Data covers selection and diversity of evaluation datasets, Computation Resources covers memory, GPU, storage, and inference constraints, and Domain Expertise covers contextually meaningful metrics and human evaluation. These four letters anchor a checklist of eight preparation steps, an applicability-analysis stage in which evaluators weight dimensions and prune unnecessary evaluations, and documentation standards that include model cards and data sheets. If the paper is right, evaluating an LLM becomes a disciplined method that can be repeated, audited, and adapted to domains such as healthcare or law.
Load-bearing premise
The load-bearing premise is the gap claim introduced in Section 1: that no prior work offers an actionable, cohesive evaluation guideline, even though Section 7 itself lists existing frameworks and tools that could be read as exactly such guidelines.
Editorial extensions
If this is right
- Following the ABCD checklist before running experiments makes model-selection decisions traceable to the stated use case.
- Resource-constrained teams can prune low-priority evaluation dimensions and still produce a defensible evaluation report.
- Disclosing dimension weights makes benchmark results comparable between teams with different priorities.
- Domain experts gain a defined role in evaluation through choosing datasets, metrics, and qualitative checks that automated benchmarks miss.
- New evaluation metrics and tools can be placed into a single process instead of being treated as competing leaderboards.
Reading between the lines
- If ABCD becomes common practice, questions like 'which model is best' would shift from aggregate leaderboard rankings to context-specific fitness-for-use statements, reducing the misleading simplicity of a single number.
- The weighting step implies a research program of eliciting and validating stakeholder weights for evaluation dimensions, since the paper acknowledges those weights are subjective.
- The framework's advice implies a testable hypothesis: teams using the checklist produce more reproducible and decision-relevant evaluations than teams relying on generic benchmarks, which a controlled comparison could check.
- The suggested multi-agent evaluation direction could operationalize ABCD by assigning each letter to an agent with distinct responsibilities, an extension the paper mentions but does not implement.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes that existing LLM evaluation literature lacks an actionable, cohesive process that integrates use-case context with ethical and operational considerations. To fill this gap, it introduces the ABCD framework (Algorithm, Big Data, Computation Resources, Domain Expertise), a nine-step evaluation checklist in Table 2, a three-stage workflow of checklist, applicability analysis, and documentation, and a targeted survey of evaluation dimensions (performance, robustness, fairness, explainability, safety) and metrics. The paper explicitly disclaims offering a new evaluation method or an exhaustive survey, and instead emphasizes formalizing the process and providing practical tools.
Significance. The paper's strength is its compact organization of a large body of evaluation literature into usable categories, and the ABCD checklist is a coherent, readable starting point for practitioners. Several existing frameworks (HELM, LalaEval, fmeval, Chang et al., Peng et al.) are cited, which makes the paper a useful pointer into the space. However, the central claim of a missing actionable guideline is contradicted by those same citations, and the paper provides no worked example or empirical demonstration that the ABCD checklist improves evaluation decisions. With a reframed contribution and a concrete application, the material could serve as a useful tutorial or position piece, but as written it does not substantiate the claimed novelty.
major comments (3)
- [Section 1 / Section 7] The paper's load-bearing claim in Section 1 that 'there exists no actionable evaluation guideline incorporating a cohesive process' is internally inconsistent with Section 7. Section 7.2 describes LalaEval as 'a holistic human evaluation framework for domain-specific LLMs, encompassing domain specification, criteria establishment, benchmark dataset creation, evaluation rubric construction, and thorough analysis of evaluation outcomes,' which is precisely an actionable, domain-aware process. Section 7.1 describes Peng et al.'s two-stage framework from core abilities to agent applications and Chang et al.'s categorization of evaluation methods, and Section 7.2 describes fmeval as an open-source library covering both performance and responsible-AI dimensions. Since these existing frameworks and tools provide structured, context-aware evaluation processes, the gap claim as stated is contradicted by the paper's own survey. The contribution should be reframed as a synthesis or operational checklist that consolidates existing guidelines, with an explicit paragraph stating what ABCD adds beyond terminology.
- [Section 2.3 / Table 1] The memory-requirement guidance is internally inconsistent. Section 2.3 first states that a 7B-parameter model 'requires approximately 28 GB of memory, assuming 4 bytes per parameter,' then immediately gives a rule of thumb of approximately 2 × X GB for X billion parameters in bfloat16/float16. Table 1 lists the 7B row as 14 GB and all rows follow the 2 bytes-per-parameter scaling. If the 4 bytes-per-parameter figure refers to FP32 weights and the table refers to BF16 weights, this distinction must be stated; if the table is meant to include runtime activation memory, the relationship is mislabeled. Because Table 1 is presented as planning guidance and the checklist includes 'Allocate Resources (C),' an inconsistent resource model weakens the paper's practical utility.
- [Section 5 / Table 2] The paper claims to formalize the evaluation process, but Table 2's checklist consists of generic project-management steps (define objectives, prioritize dimensions, select datasets, identify metrics, establish baselines, address ethics, allocate resources, document, iterate) with no illustration of how the ABCD decomposition changes a concrete evaluation decision. Section 5.2 offers informal examples of selectively weighting dimensions, but there is no end-to-end use case, case study, or comparison with an existing framework such as LalaEval or HELM. Adding at least one worked example (e.g., evaluating a model for healthcare question answering or code generation) and, if feasible, a comparison with a baseline evaluation practice would substantiate the claim that the framework is actionable and useful.
minor comments (5)
- [Section 1] The phrase 'how to systemically approach LLM evaluation' should be 'systematically,' and the sentence structure in 'current research [10, 58] lacks a comprehensive...' should be rephrased to make clear that the references do not themselves lack comprehensiveness.
- [Section 2.3] The claim that models 'exceeding 100 billion parameters demand exponentially more memory' is inaccurate relative to Table 1, which shows linear growth at 2 bytes per parameter; replace 'exponentially' with 'proportionally' or specify which overheads become nonlinear.
- [Figure 1] The workflow diagram shows five unlabeled boxes and no arrow labels or stage names; annotate it to match Sections 5.1–5.3 so that the relationship between the checklist, applicability analysis, and documentation is clear.
- [Section 5.3] The bullet beginning 'Employing standardized documentation tools...' is a sentence continuation rather than a parallel bullet item; merge it into the previous line or rewrite it as a proper bullet.
- [Section 3.1] The heading 'Entity/Word Extraction are tasks' should read 'Entity/Word Extraction is a task category,' and the subsequent sentence beginning 'This category encompasses...' should be adjusted for number agreement.
Circularity Check
No significant circularity: the paper is a survey and framework proposal with no fitted quantities, no predictions derived from inputs, and no load-bearing self-citation chain.
full rationale
This is a review/framework paper rather than a derivation: it proposes the ABCD evaluation framework, surveys existing evaluation dimensions, metrics, and tools, and provides a checklist. There are no equations, fitted parameters, or empirical predictions whose output is equivalent to an input by construction. The paper's self-citations ([87], [92]) are background references to the authors' prior work on LLM-as-evaluator benchmarks and healthcare data augmentation; they are not used to justify the central framework or to forbid alternatives. The main contestable claim is the Section 1 assertion that 'there exists no actionable evaluation guideline incorporating a cohesive process,' which is a novelty/gap claim and, if anything, a correctness concern given the frameworks cited in Section 7 (HELM, Chang et al., Peng et al., fmeval, LalaEval). But an overstated gap claim is not circularity: the paper does not define its contribution in terms of that gap, nor does it fit a parameter and then rename the fit as a prediction. The ABCD checklist is a generic organizational device, not a renamed empirical result presented as unification. Accordingly, no circular step can be quoted and exhibited under the specified patterns, and the appropriate finding is no significant circularity.
Assumptions & free parameters
assumptions (3)
- ad hoc to paper The ABCD decomposition (Algorithm, Big Data, Computation Resources, Domain Expertise) is a sufficient organizing basis for LLM evaluation.
- domain assumption A subset of evaluation dimensions can be pruned for a task, and the relative importance and weights are inherently subjective.
- domain assumption Existing literature provides no actionable evaluation guideline incorporating a cohesive process.
Cite this review
Pith. "Pith review of The Science of Evaluating Foundation Models." pith.science (2026). https://pith.science/paper/3KACSY7B
@misc{pith2026250209670,
author = {Pith},
title = {Pith review of: The Science of Evaluating Foundation Models},
year = {2026},
howpublished = {\url{https://pith.science/paper/3KACSY7B}},
note = {Machine review of arXiv:2502.09670}
}
read the original abstract
The emergent phenomena of large foundation models have revolutionized natural language processing. However, evaluating these models presents significant challenges due to their size, capabilities, and deployment across diverse applications. Existing literature often focuses on individual aspects, such as benchmark performance or specific tasks, but fails to provide a cohesive process that integrates the nuances of diverse use cases with broader ethical and operational considerations. This work focuses on three key aspects: (1) Formalizing the Evaluation Process by providing a structured framework tailored to specific use-case contexts, (2) Offering Actionable Tools and Frameworks such as checklists and templates to ensure thorough, reproducible, and practical evaluations, and (3) Surveying Recent Work with a targeted review of advancements in LLM evaluation, emphasizing real-world applications.
Figures
Forward citations
Cited by 2 Pith papers
-
DF3DV-1K: A Large-Scale Dataset and Benchmark for Distractor-Free Novel View Synthesis
DF3DV-1K supplies 1,048 scenes with clean and cluttered image pairs plus a challenging 41-scene subset to benchmark and improve distractor-free radiance field methods.
-
Deterministic Inference across Tensor Parallel Sizes That Eliminates Training-Inference Mismatch
Tree-Based Invariant Kernels fix the floating-point reduction order across GPUs, making LLM logits and sampled tokens bitwise identical for tensor-parallel sizes 1/2/4/8 and exactly matching vLLM (TP=4) with FSDP (TP=1).
Reference graph
Works this paper leans on
-
[1]
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Al- tenschmidt, Sam Altman, Shyamal Anadkat, et al. 2023. Gpt-4 tech- nical report. arXiv preprint arXiv:2303.08774 (2023)
arXiv 2023
-
[2]
Hizkiel Mitiku Alemayehu, Hamada M Zahera, and Axel- Cyrille Ngonga Ngomo. 2024. Error Analysis of Multilingual Lan- guage Models in Machine Translation: A Case Study of English- Amharic Translation. In Proceedings of the 2024 Conference on Empir- ical Methods in Natural Language Processing . 19758–19768
2024
-
[3]
Zico Kolter, Matt Fredrikson, et al
Maksym Andriushchenko, Alexandra Souly, Mateusz Dziemian, Derek Duenas, Maxwell Lin, Justin Wang, Dan Hendrycks, Andy Zou, Yuan et al. Zico Kolter, Matt Fredrikson, et al. 2024. Agentharm: A benchmark for measuring harmfulness of llm agents. arXiv preprint arXiv:2410.09024 (2024)
arXiv 2024
-
[4]
Jinze Bai, Shuai Bai, Yunfei Chu, Zeyu Cui, Kai Dang, Xiaodong Deng, Yang Fan, Wenbin Ge, Yu Han, Fei Huang, et al. 2023. Qwen technical report. arXiv preprint arXiv:2309.16609 (2023)
arXiv 2023
-
[5]
Payal Bajaj, Daniel Campos, Nick Craswell, Li Deng, Jianfeng Gao, Xiaodong Liu, Rangan Majumder, Andrew McNamara, Bhaskar Mitra, Tri Nguyen, Mir Rosenberg, Xia Song, Alina Stoica, Saurabh Tiwary, and Tong Wang. 2018. MS MARCO: A Human Generated MAchine Reading COmprehension Dataset. arXiv:1611.09268 [cs.CL] https: //arxiv.org/abs/1611.09268
arXiv 2018
-
[6]
Satanjeev Banerjee and Alon Lavie. 2005. METEOR: An Automatic Metric for MT Evaluation with Improved Correlation with Human Judgments. In Proceedings of the ACL Workshop on Intrinsic and Extrin- sic Evaluation Measures for Machine Translation and/or Summariza- tion, Jade Goldstein, Alon Lavie, Chin-Yew Lin, and Clare Voss (Eds.). Association for Computati...
2005
-
[7]
Barry Becker and Ronny Kohavi. 1996. Adult. UCI Machine Learning Repository. DOI: https://doi.org/10.24432/C5XW20
doi:10.24432/c5xw20 1996
-
[8]
Samuel R Bowman, Gabor Angeli, Christopher Potts, and Christo- pher D Manning. 2015. A large annotated corpus for learning natural language inference. arXiv preprint arXiv:1508.05326 (2015)
arXiv 2015
Show all 109 references
-
[9]
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Ka- plan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al . 2020. Language Models are Few- Shot Learners. In Advances in Neural Information Processing Systems , Vol. 33. 1877–1901
2020
-
[10]
Yupeng Chang, Xu Wang, Jindong Wang, Yuan Wu, Linyi Yang, Kaijie Zhu, Hao Chen, Xiaoyuan Yi, Cunxiang Wang, Yidong Wang, et al. 2024. A survey on evaluation of large language models. ACM Transactions on Intelligent Systems and Technology 15, 3 (2024), 1–45
2024
-
[11]
Yanda Chen, Ruiqi Zhong, Narutatsu Ri, Chen Zhao, He He, Jacob Steinhardt, Zhou Yu, and Kathleen McKeown. 2023. Do models explain themselves? counterfactual simulatability of natural language explanations. arXiv preprint arXiv:2307.08678 (2023)
2023 arXiv
-
[12]
Wei-Lin Chiang, Lianmin Zheng, Ying Sheng, Anastasios Nikolas An- gelopoulos, Tianle Li, Dacheng Li, Hao Zhang, Banghua Zhu, Michael Jordan, Joseph E Gonzalez, et al . 2024. Chatbot arena: An open platform for evaluating llms by human preference. arXiv preprint arXiv:2403.04132 (2024)
2024 arXiv
-
[13]
George Chrysostomou and Nikolaos Aletras. 2021. Improving the faithfulness of attention-based explanations with task-specific infor- mation for text classification. arXiv preprint arXiv:2105.02657 (2021)
2021 arXiv
-
[14]
Nick Craswell. 2009. Mean Reciprocal Rank . Springer US, Boston, MA, 1703–1703. https://doi.org/10.1007/978-0-387-39940-9_488
2009 doi
-
[15]
Francesco Croce, Maksym Andriushchenko, Vikash Sehwag, Edoardo Debenedetti, Nicolas Flammarion, Mung Chiang, Prateek Mittal, and Matthias Hein. 2021. RobustBench: a standardized adversarial ro- bustness benchmark. arXiv:2010.09670 [cs.LG] https://arxiv.org/abs/ 2010.09670
2021 arXiv
-
[16]
Tianyu Cui, Yanling Wang, Chuanpu Fu, Yong Xiao, Sijia Li, Xinhao Deng, Yunpeng Liu, Qinglin Zhang, Ziyi Qiu, Peiyang Li, et al. 2024. Risk taxonomy, mitigation, and assessment benchmarks of large language model systems. arXiv preprint arXiv:2401.05778 (2024)
2024 arXiv
-
[17]
Maria De-Arteaga, Alexey Romanov, Hanna Wallach, Jennifer Chayes, Christian Borgs, Alexandra Chouldechova, Sahin Geyik, Krishnaram Kenthapadi, and Adam Tauman Kalai. 2019. Bias in bios: A case study of semantic representation bias in a high-stakes setting. In proceedings of th...
2019
-
[18]
Jay DeYoung, Sarthak Jain, Nazneen Fatema Rajani, Eric Lehman, Caiming Xiong, Richard Socher, and Byron C Wallace. 2019. ERASER: A benchmark to evaluate rationalized NLP models. arXiv preprint arXiv:1911.03429 (2019)
2019 arXiv
-
[19]
William Dieterich, Christina Mendoza, and Tim Brennan. 2016. COM- PAS risk scales: Demonstrating accuracy equity and predictive parity. Northpointe Inc 7, 4 (2016), 1–36
2016
-
[20]
Esin Durmus, He He, and Mona Diab. 2020. FEQA: A Question Answering Evaluation Framework for Faithfulness Assessment in Abstractive Summarization. In Proceedings of the 58th Annual Meet- ing of the Association for Computational Linguistics , Dan Jurafsky, Joyce Chai, Natalie S...
2020 doi
-
[21]
Javid Ebrahimi, Anyi Rao, Daniel Lowd, and Dejing Dou. 2018. Hot- Flip: White-Box Adversarial Examples for Text Classification. In Proceedings of the 56th Annual Meeting of the Association for Com- putational Linguistics (Volume 2: Short Papers) , Iryna Gurevych and Yusuke Miy...
2018 doi
-
[22]
James Foulds, Rashidul Islam, Kamrun Naher Keya, and Shimei Pan. 2019. An Intersectional Definition of Fairness. arXiv:1807.08362 [cs.LG] https://arxiv.org/abs/1807.08362
2019 arXiv
-
[23]
Yonatan Geifman and Ran El-Yaniv. 2017. Selective Classification for Deep Neural Networks. arXiv:1705.08500 [cs.LG] https://arxiv.org/ abs/1705.08500
2017 arXiv
-
[24]
Tanya Goyal and Greg Durrett. 2020. Evaluating Factuality in Gen- eration with Dependency-level Entailment. In Findings of the As- sociation for Computational Linguistics: EMNLP 2020 , Trevor Cohn, Yulan He, and Yang Liu (Eds.). Association for Computational Linguis- tics, Onl...
2020 doi
-
[25]
Weinberger
Chuan Guo, Geoff Pleiss, Yu Sun, and Kilian Q. Weinberger. 2017. On Calibration of Modern Neural Networks. arXiv:1706.04599 [cs.LG] https://arxiv.org/abs/1706.04599
2017 arXiv
-
[26]
Taicheng Guo, Xiuying Chen, Yaqi Wang, Ruidi Chang, Shichao Pei, Nitesh V Chawla, Olaf Wiest, and Xiangliang Zhang. 2024. Large lan- guage model based multi-agents: A survey of progress and challenges. arXiv preprint arXiv:2402.01680 (2024)
2024 arXiv
-
[27]
Zishan Guo, Renren Jin, Chuang Liu, Yufei Huang, Dan Shi, Linhao Yu, Yan Liu, Jiaxuan Li, Bojian Xiong, Deyi Xiong, et al. 2023. Evaluating large language models: A comprehensive survey. arXiv preprint arXiv:2310.19736 (2023)
2023 arXiv
- [28]
-
[29]
Hsin-Yi Hsieh, Shih-Cheng Huang, and Richard Tsai. 2024. TWBias: A Benchmark for Assessing Social Bias in Traditional Chinese Large Language Models through a Taiwan Cultural Lens. In Findings of the Association for Computational Linguistics: EMNLP 2024 . 8688–8704
2024
-
[30]
Simon Hughes, Minseok Bae, and Miaoran Li. 2023. Vectara Hal- lucination Leaderboard. https://github.com/vectara/hallucination- leaderboard
2023
-
[31]
Albert Q Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bam- ford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, et al. 2023. Mis- tral 7B. arXiv preprint arXiv:2310.06825 (2023)
2023 arXiv
-
[32]
Di Jin, Zhijing Jin, Joey Tianyi Zhou, and Peter Szolovits. 2020. Is BERT Really Robust? A Strong Baseline for Natural Language Attack The Science of Evaluating Foundation Models on Text Classification and Entailment. arXiv:1907.11932 [cs.CL] https://arxiv.org/abs/1907.11932
2020 arXiv
-
[33]
Mandar Joshi, Eunsol Choi, Daniel Weld, and Luke Zettlemoyer
-
[34]
Tom Kocmi, Eleftherios Avramidis, Rachel Bawden, Ondřej Bojar, An- ton Dvorkovich, Christian Federmann, Mark Fishel, Markus Freitag, Thamme Gowda, Roman Grundkiewicz, Barry Haddow, Marzena Karpinska, Philipp Koehn, Benjamin Marie, Christof Monz, Kenton Murray, Masaaki Nagata, ...
2024
-
[35]
Earnshaw, Imran S
Pang Wei Koh, Shiori Sagawa, Henrik Marklund, Sang Michael Xie, Marvin Zhang, Akshay Balsubramani, Weihua Hu, Michihiro Yasunaga, Richard Lanas Phillips, Irena Gao, Tony Lee, Etienne David, Ian Stavness, Wei Guo, Berton A. Earnshaw, Imran S. Haque, Sara Beery, Jure Leskovec, A...
2021 arXiv
-
[36]
Wojciech Kryscinski, Bryan McCann, Caiming Xiong, and Richard Socher. 2020. Evaluating the Factual Consistency of Abstractive Text Summarization. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP) , Bonnie Webber, Trevor Cohn, Yul...
2020 doi
-
[37]
Dai, Jakob Uszkor- eit, Quoc Le, and Slav Petrov
Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polo- sukhin, Jacob Devlin, Kenton Lee, Kristina Toutanova, Llion Jones, Matthew Kelcey, Ming-Wei Chang, Andrew M. Dai, Jakob Uszkor- eit, Quoc Le, and ...
2019 doi
-
[38]
Tamera Lanham, Anna Chen, Ansh Radhakrishnan, Benoit Steiner, Carson Denison, Danny Hernandez, Dustin Li, Esin Durmus, Evan Hubinger, Jackson Kernion, et al. 2023. Measuring faithfulness in chain-of-thought reasoning. arXiv preprint arXiv:2307.13702 (2023)
2023 arXiv
-
[39]
Md Tahmid Rahman Laskar, Sawsan Alqahtani, M Saiful Bari, Miza- nur Rahman, Mohammad Abdullah Matin Khan, Haidar Khan, Israt Jahan, Amran Bhuiyan, Chee Wei Tan, Md Rizwan Parvez, et al. 2024. A systematic survey and critical review on evaluating large language models: Challeng...
2024
-
[40]
Mina Lee, Megha Srivastava, Amelia Hardy, John Thickstun, Esin Durmus, Ashwin Paranjape, Ines Gerard-Ursin, Xiang Lisa Li, Faisal Ladhak, Frieda Rong, et al. 2022. Evaluating human-language model interaction. arXiv preprint arXiv:2212.09746 (2022)
2022 arXiv
-
[41]
Noah Lee, Na Min An, and James Thorne. 2023. Can Large Language Models Capture Dissenting Human Voices? arXiv:2305.13788 [cs.CL] https://arxiv.org/abs/2305.13788
2023 arXiv
-
[42]
Junyi Li, Xiaoxue Cheng, Wayne Xin Zhao, Jian-Yun Nie, and Ji- Rong Wen. 2023. Halueval: A large-scale hallucination evaluation benchmark for large language models.arXiv preprint arXiv:2305.11747 (2023)
2023 arXiv
-
[43]
Jiwei Li, Michel Galley, Chris Brockett, Jianfeng Gao, and Bill Dolan
-
[45]
Manning, Christo- pher Ré, Diana Acosta-Navas, Drew A
Percy Liang, Rishi Bommasani, Tony Lee, Dimitris Tsipras, Dilara Soylu, Michihiro Yasunaga, Yian Zhang, Deepak Narayanan, Yuhuai Wu, Ananya Kumar, Benjamin Newman, Binhang Yuan, Bobby Yan, Ce Zhang, Christian Cosgrove, Christopher D. Manning, Christo- pher Ré, Diana Acosta-Nav...
2023 arXiv
-
[46]
Yuanzhi Liang, Linchao Zhu, and Yi Yang. 2024. AntEval: Quanti- tatively Evaluating Informativeness and Expressiveness of Agent Social Interactions. arXiv preprint arXiv:2401.06509 (2024)
2024 arXiv
-
[47]
Chin-Yew Lin. 2004. ROUGE: A Package for Automatic Evalua- tion of Summaries. In Text Summarization Branches Out . Associ- ation for Computational Linguistics, Barcelona, Spain, 74–81. https: //aclanthology.org/W04-1013/
2004
-
[48]
Qingyu Lu, Baopu Qiu, Liang Ding, Kanjian Zhang, Tom Kocmi, and Dacheng Tao. 2023. Error Analysis Prompting Enables Human-Like Translation Evaluation in Large Language Models. arXiv preprint arXiv:2303.13809 (2023)
2023 arXiv
-
[49]
Andrew Maas, Raymond E Daly, Peter T Pham, Dan Huang, Andrew Y Ng, and Christopher Potts. 2011. Learning word vectors for sentiment analysis. In Proceedings of the 49th annual meeting of the association for computational linguistics: Human language technologies . 142–150
2011
-
[50]
Marta Marchiori Manerba, Karolina Stańczak, Riccardo Guidotti, and Isabelle Augenstein. 2023. Social bias probing: Fairness benchmarking for language models. arXiv preprint arXiv:2311.09090 (2023)
2023 arXiv
-
[51]
Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher
-
[52]
Moin Nadeem, Anna Bethke, and Siva Reddy. 2020. StereoSet: Measur- ing stereotypical bias in pretrained language models. arXiv preprint arXiv:2004.09456 (2020)
2020 arXiv
-
[53]
Ramesh Nallapati, Bowen Zhou, Caglar Gulcehre, Bing Xiang, et al
-
[54]
Pointer sentinel mixture models.arXiv preprint arXiv:1609.07843 (2016)
2016 arXiv
-
[55]
Cohen, and Mirella Lapata
Shashi Narayan, Shay B. Cohen, and Mirella Lapata. 2018. Don’t Give Me the Details, Just the Summary! Topic-Aware Convolutional Neural Networks for Extreme Summarization. In Proceedings of the Yuan et al. 2018 Conference on Empirical Methods in Natural Language Processing , El...
2018 doi
-
[56]
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002. BLEU: a method for automatic evaluation of machine translation. In Proceedings of the 40th Annual Meeting of the Association for Compu- tational Linguistics (ACL). 311–318
2002
-
[57]
arXiv preprint arXiv:1602.06023 (2016)
Abstractive text summarization using sequence-to-sequence rnns and beyond. arXiv preprint arXiv:1602.06023 (2016)
2016 arXiv
-
[58]
Nikita Nangia, Clara Vania, Rasika Bhalerao, and Samuel R Bowman
-
[59]
Sameer Pradhan, Alessandro Moschitti, Nianwen Xue, Hwee Tou Ng, Anders Björkelund, Olga Uryupina, Yuchen Zhang, and Zhi Zhong. 2013. Towards robust linguistic analysis using ontonotes. In Proceedings of the Seventeenth Conference on Computational Natural Language Learning. 143–152
2013
-
[60]
Chengwei Qin, Aston Zhang, Zhuosheng Zhang, Jiaao Chen, Michi- hiro Yasunaga, and Diyi Yang. 2023. Is ChatGPT a General-Purpose Natural Language Processing Task Solver? arXiv:2302.06476 [cs.CL] https://arxiv.org/abs/2302.06476
2023 arXiv
-
[61]
Yiwei Qin, Kaiqiang Song, Yebowen Hu, Wenlin Yao, Sangwoo Cho, Xiaoyang Wang, Xuansheng Wu, Fei Liu, Pengfei Liu, and Dong Yu
-
[62]
Alicia Parrish, Angelica Chen, Nikita Nangia, Vishakh Padmakumar, Jason Phang, Jana Thompson, Phu Mon Htut, and Samuel R Bowman
-
[63]
Ansh Radhakrishnan, Karina Nguyen, Anna Chen, Carol Chen, Car- son Denison, Danny Hernandez, Esin Durmus, Evan Hubinger, Jack- son Kernion, Kamil ˙e Lukoši ¯ut˙e, et al . 2023. Question decomposi- tion improves the faithfulness of model-generated reasoning. arXiv preprint arXi...
2023 arXiv
-
[64]
Ji-Lun Peng, Sijia Cheng, Egil Diau, Yung-Yu Shih, Po-Heng Chen, Yen-Ting Lin, and Yun-Nung Chen. 2024. A Survey of Useful LLM Evaluation. arXiv preprint arXiv:2406.00936 (2024)
2024 arXiv
-
[65]
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang
-
[66]
Abhilasha Ravichander, Siddharth Dalmia, Maria Ryskina, Florian Metze, Eduard Hovy, and Alan W Black. 2021. NoiseQA: Challenge Set Evaluation for User-Centric Question Answering. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational...
2021 doi
-
[67]
Rush, Sumit Chopra, and Jason Weston
Alexander M. Rush, Sumit Chopra, and Jason Weston. 2015. A Neural Attention Model for Abstractive Sentence Summarization. Proceed- ings of the 2015 Conference on Empirical Methods in Natural Language Processing (2015). https://doi.org/10.18653/v1/d15-1044
2015 doi
-
[68]
Erik F Sang and Fien De Meulder. 2003. Introduction to the CoNLL- 2003 shared task: Language-independent named entity recognition. arXiv preprint cs/0306050 (2003)
2003 arXiv
-
[69]
Han Qiu, Jiaxing Huang, Peng Gao, Qin Qi, Xiaoqin Zhang, Ling Shao, and Shijian Lu. 2024. LongHalQA: Long-Context Hallucination Evaluation for MultiModal Large Language Models. arXiv preprint arXiv:2410.09962 (2024)
2024 arXiv
-
[70]
Yaozong Shen, Lijie Wang, Ying Chen, Xinyan Xiao, Jing Liu, and Hua Wu. 2022. An Interpretability Evaluation Benchmark for Pre-trained Language Models. arXiv preprint arXiv:2207.13948 (2022)
2022 arXiv
-
[71]
P Rajpurkar. 2016. Squad: 100,000+ questions for machine compre- hension of text. arXiv preprint arXiv:1606.05250 (2016)
2016 arXiv
-
[72]
Till Speicher, Hoda Heidari, Nina Grgic-Hlaca, Krishna P Gummadi, Adish Singla, Adrian Weller, and Muhammad Bilal Zafar. 2018. A unified approach to quantifying algorithmic unfairness: Measuring individual &group unfairness via inequality indices. In Proceedings of the 24th AC...
2018
-
[73]
In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing , Jian Su, Kevin Duh, and Xavier Car- reras (Eds.)
SQuAD: 100,000+ Questions for Machine Comprehension of Text. In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing , Jian Su, Kevin Duh, and Xavier Car- reras (Eds.). Association for Computational Linguistics, Austin, Texas, 2383–2392. https...
2016 doi
-
[74]
Annalisa Szymanski, Simret Araya Gebreegziabher, Oghenemaro Anuyah, Ronald A Metoyer, and Toby Jia-Jun Li. 2024. Comparing Criteria Development Across Domain Experts, Lay Users, and Models in Large Language Model Evaluation. arXiv preprint arXiv:2410.02054 (2024)
2024
-
[75]
Thomas Yu Chow Tam, Sonish Sivarajkumar, Sumit Kapoor, Alisa V Stolyar, Katelyn Polanska, Karleigh R McCarthy, Hunter Osterhoudt, Xizhi Wu, Shyam Visweswaran, Sunyang Fu, et al. 2024. A framework for human evaluation of large language models in healthcare derived from literatu...
2024
-
[76]
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al . 2023. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971 (2023)
2023 arXiv
-
[77]
Pola Schwöbel, Luca Franceschi, Muhammad Bilal Zafar, Keerthan Vasist, Aman Malhotra, Tomer Shenhar, Pinal Tailor, Pinar Yilmaz, Michael Diamond, and Michele Donini. 2024. Evaluating Large Lan- guage Models with fmeval. arXiv preprint arXiv:2407.12872 (2024)
2024 arXiv
-
[78]
Alex Wang. 2018. Glue: A multi-task benchmark and analy- sis platform for natural language understanding. arXiv preprint arXiv:1804.07461 (2018)
2018 arXiv
-
[79]
Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christo- pher D Manning, Andrew Y Ng, and Christopher Potts. 2013. Re- cursive deep models for semantic compositionality over a sentiment treebank. In Proceedings of the 2013 conference on empirical methods in natural lang...
2013
-
[80]
Truong, Simran Arora, Mantas Mazeika, Dan Hendrycks, Zinan Lin, Yu Cheng, Sanmi Koyejo, Dawn Song, and Bo Li
Boxin Wang, Weixin Chen, Hengzhi Pei, Chulin Xie, Mintong Kang, Chenhui Zhang, Chejian Xu, Zidi Xiong, Ritik Dutta, Rylan Schaeffer, Sang T. Truong, Simran Arora, Mantas Mazeika, Dan Hendrycks, Zinan Lin, Yu Cheng, Sanmi Koyejo, Dawn Song, and Bo Li. 2024. DecodingTrust: A Com...
2024 arXiv
-
[81]
Chongyan Sun, Ken Lin, Shiwei Wang, Hulong Wu, Chengfei Fu, and Zhen Wang. 2024. LalaEval: A Holistic Human Evaluation Frame- work for Domain-Specific Large Language Models. arXiv preprint arXiv:2408.13338 (2024)
2024 arXiv
-
[82]
Junyang Wang, Yuhang Wang, Guohai Xu, Jing Zhang, Yukai Gu, Haitao Jia, Ming Yan, Ji Zhang, and Jitao Sang. 2023. An llm-free The Science of Evaluating Foundation Models multi-dimensional benchmark for mllms hallucination evaluation. arXiv preprint arXiv:2311.07397 (2023)
2023 arXiv
-
[83]
Longyue Wang, Chenyang Lyu, Tianbo Ji, Zhirui Zhang, Dian Yu, Shuming Shi, and Zhaopeng Tu. 2023. Document-Level Machine Translation with Large Language Models. arXiv:2304.02210 [cs.CL] https://arxiv.org/abs/2304.02210
2023 arXiv
-
[84]
Smith, and Teruko Mitamura
Mengqiu Wang, Noah A. Smith, and Teruko Mitamura. 2007. What is the Jeopardy Model? A Quasi-Synchronous Grammar for QA. In Proceedings of the 2007 Joint Conference on Empirical Methods in Natural Language Processing and Computational Natural Lan- guage Learning (EMNLP-CoNLL), ...
2007
-
[85]
Miles Turpin, Julian Michael, Ethan Perez, and Samuel Bowman
-
[86]
Advances in Neural Information Processing Systems 36 (2024)
Language models don’t always say what they think: unfaithful explanations in chain-of-thought prompting. Advances in Neural Information Processing Systems 36 (2024)
2024
-
[87]
Yicheng Wang, Jiayi Yuan, Yu-Neng Chuang, Zhuoer Wang, Yingchi Liu, Mark Cusick, Param Kulkarni, Zhengping Ji, Yasser Ibrahim, and Xia Hu. 2024. DHP Benchmark: Are LLMs Good NLG Evaluators? arXiv preprint arXiv:2408.13704 (2024)
2024 arXiv
-
[88]
Alex Wang, Kyunghyun Cho, and Mike Lewis. 2020. Asking and Answering Questions to Evaluate the Factual Consistency of Sum- maries. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, Dan Jurafsky, Joyce Chai, Natalie Schluter, and Joel T...
2020 doi
-
[89]
Adina Williams, Nikita Nangia, and Samuel R Bowman. 2017. A broad-coverage challenge corpus for sentence understanding through inference. arXiv preprint arXiv:1704.05426 (2017)
2017 arXiv
-
[90]
Boxin Wang, Chejian Xu, Shuohang Wang, Zhe Gan, Yu Cheng, Jian- feng Gao, Ahmed Hassan Awadallah, and Bo Li. 2022. Adversarial GLUE: A Multi-Task Benchmark for Robustness Evaluation of Lan- guage Models. arXiv:2111.02840 [cs.CL] https://arxiv.org/abs/2111. 02840
2022 arXiv
-
[91]
Wangsong Yin, Mengwei Xu, Yuanchun Li, and Xuanzhe Liu
-
[92]
Jiayi Yuan, Ruixiang Tang, Xiaoqian Jiang, and Xia Hu. 2024. Large language models for healthcare data augmentation: An example on patient-trial matching. In AMIA Annual Symposium Proceedings, Vol. 2023. 1324
2024
-
[93]
Tongxin Yuan, Zhiwei He, Lingzhong Dong, Yiming Wang, Ruijie Zhao, Tian Xia, Lizhen Xu, Binglin Zhou, Fangqi Li, Zhuosheng Zhang, et al. 2024. R-judge: Benchmarking safety risk awareness for llm agents. arXiv preprint arXiv:2401.10019 (2024)
2024 arXiv
-
[94]
Xiao Wang, Qin Liu, Tao Gui, Qi Zhang, Yicheng Zou, Xin Zhou, Jiacheng Ye, Yongxin Zhang, Rui Zheng, Zexiong Pang, Qinzhuo Wu, Zhengyan Li, Chong Zhang, Ruotian Ma, Zichu Fei, Ruijian Cai, Jun Zhao, Xingwu Hu, Zhiheng Yan, Yiding Tan, Yuan Hu, Qiyuan Bian, Zhihua Liu, Shan Qin...
2021
-
[95]
Yining Wang, Liwei Wang, Yuanzhi Li, Di He, Tie-Yan Liu, and Wei Chen. 2013. A Theoretical Analysis of NDCG Type Ranking Measures. arXiv:1304.6480 [cs.LG] https://arxiv.org/abs/1304.6480
2013 arXiv
-
[96]
Xiang Zhang, Junbo Zhao, and Yann LeCun. 2015. Character-level convolutional networks for text classification. Advances in neural information processing systems 28 (2015)
2015
-
[97]
Yufei Wang, Wanjun Zhong, Liangyou Li, Fei Mi, Xingshan Zeng, Wenyong Huang, Lifeng Shang, Xin Jiang, and Qun Liu. 2023. Align- ing large language models with human: A survey. arXiv preprint arXiv:2307.12966 (2023)
2023 arXiv
-
[98]
Kaijie Zhu, Jindong Wang, Jiaheng Zhou, Zichen Wang, Hao Chen, Yidong Wang, Linyi Yang, Wei Ye, Yue Zhang, Neil Zhenqiang Gong, and Xing Xie. 2024. PromptRobust: Towards Evaluating the Robustness of Large Language Models on Adversarial Prompts. arXiv:2306.04528 [cs.CL] https:/...
2024 arXiv
-
[99]
Lin Xu, Zhiyuan Hu, Daquan Zhou, Hongyu Ren, Zhen Dong, Kurt Keutzer, See Kiong Ng, and Jiashi Feng. 2024. Magic: Investigation of large language model powered multi-agent in cognition, adaptability, rationality and collaboration. In Proceedings of the 2024 Conference on Empir...
2024
-
[100]
Yukun Zhu, Ryan Kiros, Richard Zemel, Ruslan Salakhutdinov, Raquel Urtasun, Antonio Torralba, and Sanja Fidler. 2015. Aligning Books and Movies: Towards Story-like Visual Explanations by Watching Movies and Reading Books. arXiv:1506.06724 [cs.CV] https://arxiv. org/abs/1506.06724
2015 arXiv
-
[101]
arXiv preprint arXiv:2403.11805 (2024)
Llm as a system service on mobile devices. arXiv preprint arXiv:2403.11805 (2024)
2024 arXiv
-
[104]
Xiaohan Yuan, Jinfeng Li, Dongxia Wang, Yuefeng Chen, Xiaofeng Mao, Longtao Huang, Hui Xue, Wenhai Wang, Kui Ren, and Jingyi Wang. 2024. S-Eval: Automatic and Adaptive Test Generation for Benchmarking Safety Evaluation of Large Language Models. arXiv preprint arXiv:2405.14191 (2024)
2024 arXiv
-
[105]
Weinberger, and Yoav Artzi
Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger, and Yoav Artzi. 2020. BERTScore: Evaluating Text Generation with BERT. arXiv:1904.09675 [cs.CL] https://arxiv.org/abs/1904.09675
2020 arXiv
-
[107]
Haiyan Zhao, Hanjie Chen, Fan Yang, Ninghao Liu, Huiqi Deng, Hengyi Cai, Shuaiqiang Wang, Dawei Yin, and Mengnan Du. 2024. Explainability for large language models: A survey.ACM Transactions on Intelligent Systems and Technology 15, 2 (2024), 1–38
2024
-
[109]
Kaijie Zhu, Qinlin Zhao, Hao Chen, Jindong Wang, and Xing Xie. 2024. PromptBench: A Unified Library for Evaluation of Large Language Models. arXiv:2312.07910 [cs.AI] https://arxiv.org/abs/2312.07910
2024 arXiv
-
[2016]
A Diversity-Promoting Objective Function for Neural Conver- sation Models. InProceedings of the 2016 Conference of the North Amer- ican Chapter of the Association for Computational Linguistics: Human Language Technologies, Kevin Knight, Ani Nenkova, and Owen Ram- bow (Eds.). A...
2016 doi
-
[2017]
In Proceedings of the 55th An- nual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , Regina Barzilay and Min-Yen Kan (Eds.)
TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehension. In Proceedings of the 55th An- nual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , Regina Barzilay and Min-Yen Kan (Eds.). Associa- tion for Computatio...
-
[2020]
arXiv preprint arXiv:2010.00133 (2020)
CrowS-pairs: A challenge dataset for measuring social biases in masked language models. arXiv preprint arXiv:2010.00133 (2020)
2020 arXiv
-
[2021]
arXiv preprint arXiv:2110.08193 (2021)
BBQ: A hand-built bias benchmark for question answering. arXiv preprint arXiv:2110.08193 (2021)
2021 arXiv
-
[2024]
arXiv preprint arXiv:2401.03601 (2024)
Infobench: Evaluating instruction following ability in large language models. arXiv preprint arXiv:2401.03601 (2024)
2024 arXiv
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.