REVIEW 4 major objections 6 minor 27 references
MEQA: A Meta-Evaluation Framework for Question & Answer LLM Benchmarks
T0 review · 4 major / 6 minor · reviewed 2026-08-16 · deepseek-v4-flash
Pith's one-line read MEQA proposes a standardized scorecard for judging question-answering benchmarks and applies it to nine cybersecurity benchmarks.
desk verdict Useful rubric, unvalidated scores: the MEQA framework is a solid contribution to benchmark meta-evaluation, but its demonstration on cybersecurity benchmarks doesn't yet support the comparative claims. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the MEQA scorecard: eight named criteria—memorization robustness, prompt robustness, evaluation design, evaluator design, reproducibility, comparability, validity, and reliability—each expanded into sub-criteria rubrics, for a total of 44. Each sub-criterion carries a 1-5 description of what a low, medium, or high score looks like, plus an N/A option for criteria that do not apply to a benchmark. The rubric is what carries the argument: it converts qualitative judgments about a benchmark into numbers that can be averaged, compared, and used for gap analysis. A second mechanism is the evaluator protocol, in which human raters and, optionally, an LLM are given the sub-criterion definition and few-shot examples before scoring.
What would settle it
Take the same 44 sub-criteria and the same nine benchmarks to an independent panel of cybersecurity and evaluation experts who have not seen the paper's scores, and compute inter-rater agreement (e.g., Fleiss' kappa or ICC) plus the resulting benchmark rankings. If kappa or ICC falls below conventional thresholds or the rankings change materially, the MEQA scores are an artifact of its own raters rather than a stable property of the benchmarks.
Extended reading notes
Core claim
The central claim is that benchmark quality can be meaningfully decomposed, quantified, and compared: MEQA's eight criteria and 44 sub-criteria are meant to cover the main failure modes of QA benchmarks, and the 1-5 scores are meant to produce comparable numbers rather than vague impressions. The demonstration on nine cybersecurity benchmarks reports that HarmBench-Cyber and WMDP-Cyber score highest (3.6 and 3.5), SecQA and SECURE lowest (2.7), and that per-benchmark variability is large (standard deviations around 1.0-1.6). It also reports that few-shot LLM scoring tended to match the human evaluators, especially at extreme scores, which the paper takes as evidence that automated meta-evaluation can scale. The paper concludes that most cybersecurity benchmarks already do well on reproducibility and comparability but need work on prompt robustness and reliability.
Load-bearing premise
The scorecard is only as trustworthy as the evaluators who fill it in, and the paper's evidence for that trust comes from three human raters who are also the authors, with no chance-corrected agreement measure and no outside validation of their scores.
Editorial extensions
If this is right
- Benchmark developers can run MEQA before releasing a new QA benchmark and get a concrete list of weak sub-criteria to fix, instead of relying on intuition.
- Automated meta-evaluation becomes practical: the paper reports that few-shot LLM scores tend to match human scores, so the 44-item rubric can be applied to many benchmarks at low cost.
- The published scores give a baseline for cybersecurity benchmarks: HarmBench-Cyber at 3.6 and WMDP-Cyber at 3.5 are the current top of the set, and SecQA and SECURE at 2.7 sit at the bottom.
- Large within-benchmark standard deviations imply that a single mean score hides uneven quality, so consumers of benchmark results should look at sub-criteria profiles rather than only the headline number.
Reading between the lines
- If the pattern of low prompt robustness and reliability generalizes beyond these nine benchmarks, then many current QA leaderboards may be prompt-format artifacts; a quick check would re-run the same models on rephrased prompts and see whether rankings move.
- The paper reports over 80% exact agreement among its three human raters but no chance-corrected statistic; computing Fleiss' kappa or ICC from the sub-criteria scores would separate the rubric's clarity from the raters' shared leniency.
- Because MEQA averages sub-criteria into a mean, correlated sub-criteria could inflate or dilute the signal; a profile of per-criterion scores or a weighted aggregation would be a more conservative reading of the same data.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes MEQA, a meta-evaluation framework for question-answering (QA) benchmarks, organized around eight criteria (memorization robustness, prompt robustness, evaluation design, evaluator design, reproducibility, comparability, validity, reliability) and 44 sub-criteria, each scored on a 1–5 scale. The framework is demonstrated on nine cybersecurity QA benchmarks using three human evaluators (the authors) and GPT-4o as an LLM evaluator. The paper reports overall mean scores per benchmark (Table 2), per-criterion scores (Figure 1), and per-sub-criterion scores (Appendix B), and concludes that most benchmarks are strong in reproducibility and comparability but weak in prompt robustness and reliability. The central claim is that MEQA provides standardized, quantifiable, and comparable assessments that enable meaningful intra-benchmark comparisons and gap analysis for benchmark developers.
Significance. If MEQA were shown to be reliable and valid, it would address a genuine gap in the literature: prior meta-evaluation work has largely focused on isolated aspects such as safety-washing, reproducibility, or prompt sensitivity. The framework's 44 sub-criteria operationalize a broad set of considerations from prior work, and the paper is transparent about its proof-of-concept status in Appendix D. The framework has face validity as a structured checklist for benchmark auditing. However, the demonstration in this manuscript does not yet establish that MEQA produces stable or meaningful measurements: human-evaluator reliability is unreported, the LLM evaluation is calibrated on the same human scores, and the aggregation procedure is unspecified. These gaps directly weaken the central claim of 'meaningful intra-benchmark comparisons'; the current results are best interpreted as a pilot illustration rather than a validated instrument.
major comments (4)
- [Section 2.1, Table 2, Appendix D] The reliability of the human evaluator scores is not established. The paper reports only that the three human evaluators 'agreed exactly on the majority (over 80%) of the sub-criteria' with disagreements at most 1 point, but it provides no chance-corrected agreement measure (e.g., Cohen's or Fleiss' kappa, or ICC) and no per-sub-criterion score distributions. On a 5-point scale, high exact-agreement rates can coexist with strong central tendency and little discriminative power between benchmarks, so the reported agreement does not rule out that the mean differences in Table 2 (e.g., 3.6 vs 2.7) reflect rater noise rather than genuine differences in benchmark quality. Since the central claim of meaningful intra-benchmark comparisons rests on these scores, this is a load-bearing omission.
- [Appendix C, Section 2.1] The statement that GPT-4o 'tended to match human evaluators' is not evidence of independent agreement. Appendix C explains that the reported LLM results were obtained with few-shot prompting that included scores for benchmarks already evaluated by humans, so the LLM's agreement is partly a consequence of the calibration examples. The paper should report the zero-shot results separately, specify how many and which few-shot examples were used, and compute agreement statistics (e.g., Cohen's kappa or ICC) on benchmarks that were excluded from the few-shot set. Without this, the LLM-evaluator results cannot be interpreted as validating the framework or as evidence that automated meta-evaluation is scalable.
- [Section 2.1, Appendix B, Table 2] The aggregation of sub-criterion scores into criterion scores and overall means is unspecified. The text states that sub-criterion scores are 'aggregated to create Figure 1, Table 2 and Table 11' (Appendix B), but no formula is given: it is unclear whether criteria are equally weighted, whether N/A sub-criteria are excluded or imputed, and whether the overall mean is the mean of criteria or of sub-criteria. This makes the reported scores non-reproducible and undercuts the claims of 'quantifiable scores' and 'standardized assessments.' The aggregation rule should be stated explicitly and, ideally, implemented in the released code so that the mapping from raw sub-criterion ratings to the reported tables can be verified.
- [Appendix D, Abstract, Conclusions] The validation is self-referential. The human evaluators are the paper's authors, and the LLM evaluator is calibrated against those same evaluators. The manuscript provides no external anchor for the scores—no independent expert panel, no comparison with known benchmark properties, and no correlation with downstream performance or independently measured benchmark quality. The limitations in Appendix D acknowledge the proof-of-concept nature, but the abstract and conclusions nevertheless state that MEQA 'highlights the benchmarks' strengths and weaknesses' and 'enable[s] meaningful intra-benchmark comparisons.' These claims should be tempered until the instrument is validated against an external criterion, or the paper should be reframed as a proposal with pilot results.
minor comments (6)
- [Section 3] The word 'bechmarks' is a typo; it should read 'benchmarks.'
- [Table 3] The row 'Memorization Detection' contains a stray LaTeX command 'textit' in the text; this should be removed or rendered properly.
- [Table 8] The row 'Cross-Model Consistency' contains the typo 'appraches'; it should read 'approaches.'
- [Table 9] The row 'Representative data' contains the typo 'vulnerabilty'; it should read 'vulnerability.'
- [Table 6] The row 'Annotator Expertise' contains the typo 'Prefereable'; it should read 'Preferable.'
- [Appendix F] The code is described as 'available anonymously,' but the manuscript gives no repository identifier, URL, or anonymized link; please provide a way for readers to access the code.
Circularity Check
LLM-evaluator agreement is partly self-referential because few-shot prompts contained the human scores it is compared against; the core MEQA scoring framework itself is not circular.
-
fitted input called prediction
[Section 2.1 and Appendix C]
"The automated scoring tended to match human evaluators, particularly when producing extreme scores (1 or 5). ... We attempted both zero-shot prompting and few-shot prompting as preliminary tests. In the latter case, we included scores for benchmarks that had been already evaluated by humans, this improved the LLM’s accuracy. Our results in this paper are those obtained using few-shot prompting."
The paper presents GPT-4o's agreement with human evaluators as evidence that the LLM evaluator works, but the reported results were produced with few-shot prompts that included human scores for benchmarks already evaluated by humans. The model therefore had the target values in its context, so the observed 'match' is expected and is not independent confirmation. This is calibrating on the labels and then presenting the resulting alignment as a finding. The circularity is partial: it affects the LLM-evaluator validation claim, but not the human-score benchmark rankings, which are independent of the few-shot prompt contents.
full rationale
The central MEQA framework is not circular by construction: the eight criteria and 44 sub-criteria are explicitly synthesized from previously published evaluation literature, and the per-benchmark scores are the evaluators' 1-5 judgments using the rubric, not quantities derived from an equation that presupposes the result. The paper does not invoke any load-bearing self-citation, uniqueness theorem, or ansatz smuggled in from the authors' prior work; the references are external. The genuine circular step is confined to the LLM-evaluator component: Appendix C discloses that the reported GPT-4o results came from few-shot prompting that included human scores for already-evaluated benchmarks, and Section 2.1 then cites GPT-4o's tendency to match human evaluators as support. That agreement is contaminated by construction because the human scores were placed in the prompt as exemplars. The human-evaluator-based rankings and the framework's utility as a gap-analysis instrument are not circular, though their reliability is under-supported by the absence of kappa/ICC statistics and external validation; that is a validity concern rather than a circularity concern under the rules of this review. Overall, the paper merits a moderate circularity score because a supporting 'prediction-like' claim reduces to its calibration inputs, while the core derivation remains self-contained.
Assumptions & free parameters
free parameters (1)
- Sub-criterion aggregation weights
assumptions (4)
- domain assumption The eight criteria are the relevant and sufficient dimensions of QA benchmark quality.
- domain assumption Human expert evaluation is the ground truth for benchmark quality.
- domain assumption GPT-4o scores are a valid proxy for human evaluation.
- domain assumption The 1-5 rating scale yields meaningful, comparable scores across sub-criteria.
Cite this review
Pith. "Pith review of MEQA: A Meta-Evaluation Framework for Question & Answer LLM Benchmarks." pith.science (2026). https://pith.science/paper/JEDYDDSZ
@misc{pith2026250414039,
author = {Pith},
title = {Pith review of: MEQA: A Meta-Evaluation Framework for Question & Answer LLM Benchmarks},
year = {2026},
howpublished = {\url{https://pith.science/paper/JEDYDDSZ}},
note = {Machine review of arXiv:2504.14039}
}
read the original abstract
As Large Language Models (LLMs) advance, their potential for widespread societal impact grows simultaneously. Hence, rigorous LLM evaluations are both a technical necessity and social imperative. While numerous evaluation benchmarks have been developed, there remains a critical gap in meta-evaluation: effectively assessing benchmarks' quality. We propose MEQA, a framework for the meta-evaluation of question and answer (QA) benchmarks, to provide standardized assessments, quantifiable scores, and enable meaningful intra-benchmark comparisons. We demonstrate this approach on cybersecurity benchmarks, using human and LLM evaluators, highlighting the benchmarks' strengths and weaknesses. We motivate our choice of test domain by AI models' dual nature as powerful defensive tools and security threats.
Figures
Figures from the paper (6 more)
Reference graph
Works this paper leans on
-
[1]
Bach, S. H., Sanh, V ., Yong, Z. X., Webson, A., Raffel, C., Nayak, N. V ., Sharma, A., Kim, T., Bari, M. S., F´evry, T., Alyafeai, Z., Dey, M., Santilli, A., Sun, Z., Ben-David, S., Xu, C., Chhablani, G., Wang, H., Fries, J. A., AlShaibani, M. S., Sharma, S., Thakker, U., Almubarak, K., Tang, X., Jiang, M. T., and Rush, A. M. Promptsource: An integrated ...
arXiv 2022
-
[2]
Bhatt, M., Chennabasappa, S., Nikolaidis, C., Wan, S., Ev- timov, I., Gabi, D., Song, D., Ahmad, F., Aschermann, C., Fontana, L., Frolov, S., Giri, R. P., Kapil, D., Kozyrakis, Y ., LeBlanc, D., Milazzo, J., Straumann, A., Synnaeve, G., V on- timitta, V ., Whitman, S., and Saxe, J. Purple llama cybersece- val: A secure coding benchmark for language models...
arXiv 2023
-
[3]
Cyberseceval 2: A wide-ranging cybersecurity evaluation suite for large language models, 2024
Bhatt, M., Chennabasappa, S., Li, Y ., Nikolaidis, C., Song, D., Wan, S., Ahmad, F., Aschermann, C., Chen, Y ., Kapil, D., Molnar, D., Whitman, S., and Saxe, J. Cyberseceval 2: A wide-ranging cybersecurity evaluation suite for large language models, 2024. URL https://arxiv.org/ abs/2404.13161
arXiv 2024
-
[4]
T., Nguyen, L., Mahara, A., Lightcap, Z., Frazier, R., Fieblinger, R., Torales, G
Bhusal, D., Alam, M. T., Nguyen, L., Mahara, A., Lightcap, Z., Frazier, R., Fieblinger, R., Torales, G. L., and Rastogi, N. Secure: Benchmarking large language models for cybersecu- rity advisory, 2024. URL https://arxiv.org/abs/ 2405.20441
arXiv 2024
-
[5]
Biderman, S., Schoelkopf, H., Sutawika, L., Gao, L., Tow, J., Abbasi, B., Aji, A. F., Ammanamanchi, P. S., Black, S., Clive, J., DiPofi, A., Etxaniz, J., Fattori, B., Forde, J. Z., Foster, C., Hsu, J., Jaiswal, M., Lee, W. Y ., Li, H., Lovering, C., 2 MEQA Figure 1. Scores of cybersecurity benchmarks per criterion. N/A indicates inapplicable criteria (e.g...
arXiv 2024
-
[6]
Clark, E., August, T., Serrano, S., Haduong, N., Gururangan, S., and Smith, N. A. All that’s ‘human’ is not gold: Evaluat- ing human evaluation of generated text. In Zong, C., Xia, F., Li, W., and Navigli, R. (eds.),Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natura...
2021
-
[7]
Evaluating superhuman models with consistency checks, 2023
Fluri, L., Paleka, D., and Tram`er, F. Evaluating superhuman models with consistency checks, 2023. URL https:// arxiv.org/abs/2306.09983
arXiv 2023
-
[8]
A framework for few-shot language model evaluation, 12 2023
Gao, L., Tow, J., Abbasi, B., Biderman, S., Black, S., DiPofi, A., Foster, C., Golding, L., Hsu, J., Le Noac’h, A., Li, H., Mc- Donell, K., Muennighoff, N., Ociepa, C., Phang, J., Reynolds, L., Schoelkopf, H., Skowron, A., Sutawika, L., Tang, E., Thite, A., Wang, B., Wang, K., and Zou, A. A framework for few-shot language model evaluation, 12 2023. URL ht...
arXiv 2023
Show all 27 references
-
[9]
Jacobs, A. Z. and Wallach, H. M. Measurement and fair- ness. CoRR, abs/1912.05511, 2019. URL http://arxiv. org/abs/1912.05511
1912 arXiv
-
[10]
Sevenllm: Benchmarking, eliciting, and enhancing abilities of large language models in cyber threat intelligence, 2024
Ji, H., Yang, J., Chai, L., Wei, C., Yang, L., Duan, Y ., Wang, Y ., Sun, T., Guo, H., Li, T., Ren, C., and Li, Z. Sevenllm: Benchmarking, eliciting, and enhancing abilities of large language models in cyber threat intelligence, 2024. URL https://arxiv.org/abs/2405.03446
2024 arXiv
-
[11]
Seceval: A comprehensive benchmark for eval- uating cybersecurity knowledge of foundation models
Li, G., Li, Y ., Guannan, W., Yang, H., and Yu, Y . Seceval: A comprehensive benchmark for eval- uating cybersecurity knowledge of foundation models. https://github.com/XuanwuAI/SecEval, 2023
2023
-
[12]
D., Dombrowski, A.-K., Goel, S., Phan, L., Mukobi, G., Helm-Burger, N., Lababidi, R., Justen, L., Liu, A
Li, N., Pan, A., Gopal, A., Yue, S., Berrios, D., Gatti, A., Li, J. D., Dombrowski, A.-K., Goel, S., Phan, L., Mukobi, G., Helm-Burger, N., Lababidi, R., Justen, L., Liu, A. B., Chen, M., Barrass, I., Zhang, O., Zhu, X., Tamirisa, R., Bharathi, B., Khoja, A., Zhao, Z., Herbert...
2024 arXiv
-
[13]
ROUGE: A package for automatic evaluation of summaries
Lin, C.-Y . ROUGE: A package for automatic evaluation of summaries. In Text Summarization Branches Out , pp. 74–81, Barcelona, Spain, July 2004. Association for Com- putational Linguistics. URL https://aclanthology. org/W04-1013
2004
-
[14]
Secqa: A concise question-answering dataset for evaluating large language models in computer security, 2023
Liu, Z. Secqa: A concise question-answering dataset for evaluating large language models in computer security, 2023. URL https://arxiv.org/abs/2312.15838
2023 arXiv
-
[15]
Harmbench: A standardized evaluation frame- work for automated red teaming and robust refusal, 2024
Mazeika, M., Phan, L., Yin, X., Zou, A., Wang, Z., Mu, N., Sakhaee, E., Li, N., Basart, S., Li, B., Forsyth, D., and Hendrycks, D. Harmbench: A standardized evaluation frame- work for automated red teaming and robust refusal, 2024. URL https://arxiv.org/abs/2402.04249
2024 arXiv
-
[16]
State of what art? a call for multi-prompt llm evaluation, 2024
Mizrahi, M., Kaplan, G., Malkin, D., Dror, R., Shahaf, D., and Stanovsky, G. State of what art? a call for multi-prompt llm evaluation, 2024. URL https://arxiv.org/abs/ 2401.00595
2024 arXiv
-
[17]
H., Fitz, S., and Hendrycks, D
Ren, R., Basart, S., Khoja, A., Gatti, A., Phan, L., Yin, X., Mazeika, M., Pan, A., Mukobi, G., Kim, R. H., Fitz, S., and Hendrycks, D. Safetywashing: Do ai safety benchmarks actually measure safety progress?, 2024. URL https:// arxiv.org/abs/2407.21792
2024 arXiv
-
[18]
Jinja2 Documentation, 2024
Ronacher, A. Jinja2 Documentation, 2024. URL https: //jinja.palletsprojects.com/. Accessed: 2024- 09-17
2024
-
[19]
First tragedy, then parse: History repeats itself in the new era of large language models, 2024
Saphra, N., Fleisig, E., Cho, K., and Lopez, A. First tragedy, then parse: History repeats itself in the new era of large language models, 2024. URL https://arxiv.org/ abs/2311.05020
2024 arXiv
-
[20]
Quantifying language models’ sensitivity to spurious features in prompt design or: How i learned to start worrying about prompt for- matting, 2024
Sclar, M., Choi, Y ., Tsvetkov, Y ., and Suhr, A. Quantifying language models’ sensitivity to spurious features in prompt design or: How i learned to start worrying about prompt for- matting, 2024. URL https://arxiv.org/abs/2310. 11324
2024
-
[21]
Large language models are inconsistent and biased evaluators, 2024
Stureborg, R., Alikaniotis, D., and Suhara, Y . Large language models are inconsistent and biased evaluators, 2024. URL https://arxiv.org/abs/2405.01724
2024 arXiv
-
[22]
Subramonian, A., Yuan, X., au2, H. D. I., and Blodgett, S. L. It takes two to tango: Navigating conceptualizations of nlp tasks and measurements of performance, 2023. URL https://arxiv.org/abs/2305.09022
2023 arXiv
-
[23]
A., Jain, R., Bisztray, T., and Debbah, M
Tihanyi, N., Ferrag, M. A., Jain, R., Bisztray, T., and Debbah, M. Cybermetric: A benchmark dataset based on retrieval- augmented generation for evaluating llms in cybersecurity knowledge, 2024. URL https://arxiv.org/abs/ 2402.07688
2024 arXiv
-
[24]
Xiao, Z., Zhang, S., Lai, V ., and Liao, Q. V . Evaluating evaluation metrics: A framework for analyzing nlg evaluation metrics using measurement theory, 2023. URL https: //arxiv.org/abs/2305.14889
2023 arXiv
-
[25]
Skill-mix: a flexible and expandable family of evaluations for ai models, 2023
Yu, D., Kaur, S., Gupta, A., Brown-Cohen, J., Goyal, A., and Arora, S. Skill-mix: a flexible and expandable family of evaluations for ai models, 2023. URL https://arxiv. org/abs/2310.17567
2023 arXiv
-
[26]
P., Zhang, H., Gonzalez, J
Zheng, L., Chiang, W.-L., Sheng, Y ., Zhuang, S., Wu, Z., Zhuang, Y ., Lin, Z., Li, Z., Li, D., Xing, E. P., Zhang, H., Gonzalez, J. E., and Stoica, I. Judging llm-as-a-judge with mt-bench and chatbot arena, 2023. URL https://arxiv. org/abs/2306.05685
2023 arXiv
-
[27]
canary strings
Zhu, K., Chen, J., Wang, J., Gong, N. Z., Yang, D., and Xie, X. Dyval: Dynamic evaluation of large language models for reasoning tasks, 2024. URL https://arxiv.org/ abs/2309.17167. 4 MEQA Appendices A. Criteria and sub-criteria We outline our meta-evaluation criteria in Table ...
2024 arXiv
Reviewed August 16, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.