Pith. sign in

REVIEW 4 major objections 6 minor 27 references

MEQA: A Meta-Evaluation Framework for Question & Answer LLM Benchmarks

T0 review · 4 major / 6 minor · reviewed 2026-08-16 · deepseek-v4-flash

Pith's one-line read MEQA proposes a standardized scorecard for judging question-answering benchmarks and applies it to nine cybersecurity benchmarks.

desk verdict Useful rubric, unvalidated scores: the MEQA framework is a solid contribution to benchmark meta-evaluation, but its demonstration on cybersecurity benchmarks doesn't yet support the comparative claims. read the letter →

arxiv 2504.14039 v1 pith:JEDYDDSZ submitted 2025-04-18 cs.CL cs.AI

classification cs.CLcs.AI
keywords meta-evaluationquestion-answeringbenchmarksLLMevaluationbenchmarkqualitycybersecurityLLM-as-judgerubricpromptrobustness
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper sets out to fill a gap: there are many benchmarks for question-answering LLMs, but no standard way to judge the quality of the benchmarks themselves. It proposes MEQA, a meta-evaluation framework that turns that judgment into a scorecard: eight criteria, 44 concrete sub-criteria, each scored from 1 to 5 by human or LLM evaluators. Applied to nine cybersecurity QA benchmarks, the scorecard yields mean scores from 2.7 to 3.6 and exposes a pattern—strong reproducibility and comparability, weak prompt robustness and reliability. The paper offers MEQA as a reusable blueprint for benchmark developers to audit and improve their benchmarks.

What carries the argument

The load-bearing object is the MEQA scorecard: eight named criteria—memorization robustness, prompt robustness, evaluation design, evaluator design, reproducibility, comparability, validity, and reliability—each expanded into sub-criteria rubrics, for a total of 44. Each sub-criterion carries a 1-5 description of what a low, medium, or high score looks like, plus an N/A option for criteria that do not apply to a benchmark. The rubric is what carries the argument: it converts qualitative judgments about a benchmark into numbers that can be averaged, compared, and used for gap analysis. A second mechanism is the evaluator protocol, in which human raters and, optionally, an LLM are given the sub-criterion definition and few-shot examples before scoring.

What would settle it

Take the same 44 sub-criteria and the same nine benchmarks to an independent panel of cybersecurity and evaluation experts who have not seen the paper's scores, and compute inter-rater agreement (e.g., Fleiss' kappa or ICC) plus the resulting benchmark rankings. If kappa or ICC falls below conventional thresholds or the rankings change materially, the MEQA scores are an artifact of its own raters rather than a stable property of the benchmarks.

Watch

Extended reading notes

Core claim

The central claim is that benchmark quality can be meaningfully decomposed, quantified, and compared: MEQA's eight criteria and 44 sub-criteria are meant to cover the main failure modes of QA benchmarks, and the 1-5 scores are meant to produce comparable numbers rather than vague impressions. The demonstration on nine cybersecurity benchmarks reports that HarmBench-Cyber and WMDP-Cyber score highest (3.6 and 3.5), SecQA and SECURE lowest (2.7), and that per-benchmark variability is large (standard deviations around 1.0-1.6). It also reports that few-shot LLM scoring tended to match the human evaluators, especially at extreme scores, which the paper takes as evidence that automated meta-evaluation can scale. The paper concludes that most cybersecurity benchmarks already do well on reproducibility and comparability but need work on prompt robustness and reliability.

Load-bearing premise

The scorecard is only as trustworthy as the evaluators who fill it in, and the paper's evidence for that trust comes from three human raters who are also the authors, with no chance-corrected agreement measure and no outside validation of their scores.

Editorial extensions

If this is right

  • Benchmark developers can run MEQA before releasing a new QA benchmark and get a concrete list of weak sub-criteria to fix, instead of relying on intuition.
  • Automated meta-evaluation becomes practical: the paper reports that few-shot LLM scores tend to match human scores, so the 44-item rubric can be applied to many benchmarks at low cost.
  • The published scores give a baseline for cybersecurity benchmarks: HarmBench-Cyber at 3.6 and WMDP-Cyber at 3.5 are the current top of the set, and SecQA and SECURE at 2.7 sit at the bottom.
  • Large within-benchmark standard deviations imply that a single mean score hides uneven quality, so consumers of benchmark results should look at sub-criteria profiles rather than only the headline number.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the pattern of low prompt robustness and reliability generalizes beyond these nine benchmarks, then many current QA leaderboards may be prompt-format artifacts; a quick check would re-run the same models on rephrased prompts and see whether rankings move.
  • The paper reports over 80% exact agreement among its three human raters but no chance-corrected statistic; computing Fleiss' kappa or ICC from the sub-criteria scores would separate the rubric's clarity from the raters' shared leniency.
  • Because MEQA averages sub-criteria into a mean, correlated sub-criteria could inflate or dilute the signal; a profile of per-criterion scores or a weighted aggregation would be a more conservative reading of the same data.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. The paper proposes MEQA, a meta-evaluation framework for question-answering (QA) benchmarks, organized around eight criteria (memorization robustness, prompt robustness, evaluation design, evaluator design, reproducibility, comparability, validity, reliability) and 44 sub-criteria, each scored on a 1–5 scale. The framework is demonstrated on nine cybersecurity QA benchmarks using three human evaluators (the authors) and GPT-4o as an LLM evaluator. The paper reports overall mean scores per benchmark (Table 2), per-criterion scores (Figure 1), and per-sub-criterion scores (Appendix B), and concludes that most benchmarks are strong in reproducibility and comparability but weak in prompt robustness and reliability. The central claim is that MEQA provides standardized, quantifiable, and comparable assessments that enable meaningful intra-benchmark comparisons and gap analysis for benchmark developers.

Significance. If MEQA were shown to be reliable and valid, it would address a genuine gap in the literature: prior meta-evaluation work has largely focused on isolated aspects such as safety-washing, reproducibility, or prompt sensitivity. The framework's 44 sub-criteria operationalize a broad set of considerations from prior work, and the paper is transparent about its proof-of-concept status in Appendix D. The framework has face validity as a structured checklist for benchmark auditing. However, the demonstration in this manuscript does not yet establish that MEQA produces stable or meaningful measurements: human-evaluator reliability is unreported, the LLM evaluation is calibrated on the same human scores, and the aggregation procedure is unspecified. These gaps directly weaken the central claim of 'meaningful intra-benchmark comparisons'; the current results are best interpreted as a pilot illustration rather than a validated instrument.

major comments (4)
  1. [Section 2.1, Table 2, Appendix D] The reliability of the human evaluator scores is not established. The paper reports only that the three human evaluators 'agreed exactly on the majority (over 80%) of the sub-criteria' with disagreements at most 1 point, but it provides no chance-corrected agreement measure (e.g., Cohen's or Fleiss' kappa, or ICC) and no per-sub-criterion score distributions. On a 5-point scale, high exact-agreement rates can coexist with strong central tendency and little discriminative power between benchmarks, so the reported agreement does not rule out that the mean differences in Table 2 (e.g., 3.6 vs 2.7) reflect rater noise rather than genuine differences in benchmark quality. Since the central claim of meaningful intra-benchmark comparisons rests on these scores, this is a load-bearing omission.
  2. [Appendix C, Section 2.1] The statement that GPT-4o 'tended to match human evaluators' is not evidence of independent agreement. Appendix C explains that the reported LLM results were obtained with few-shot prompting that included scores for benchmarks already evaluated by humans, so the LLM's agreement is partly a consequence of the calibration examples. The paper should report the zero-shot results separately, specify how many and which few-shot examples were used, and compute agreement statistics (e.g., Cohen's kappa or ICC) on benchmarks that were excluded from the few-shot set. Without this, the LLM-evaluator results cannot be interpreted as validating the framework or as evidence that automated meta-evaluation is scalable.
  3. [Section 2.1, Appendix B, Table 2] The aggregation of sub-criterion scores into criterion scores and overall means is unspecified. The text states that sub-criterion scores are 'aggregated to create Figure 1, Table 2 and Table 11' (Appendix B), but no formula is given: it is unclear whether criteria are equally weighted, whether N/A sub-criteria are excluded or imputed, and whether the overall mean is the mean of criteria or of sub-criteria. This makes the reported scores non-reproducible and undercuts the claims of 'quantifiable scores' and 'standardized assessments.' The aggregation rule should be stated explicitly and, ideally, implemented in the released code so that the mapping from raw sub-criterion ratings to the reported tables can be verified.
  4. [Appendix D, Abstract, Conclusions] The validation is self-referential. The human evaluators are the paper's authors, and the LLM evaluator is calibrated against those same evaluators. The manuscript provides no external anchor for the scores—no independent expert panel, no comparison with known benchmark properties, and no correlation with downstream performance or independently measured benchmark quality. The limitations in Appendix D acknowledge the proof-of-concept nature, but the abstract and conclusions nevertheless state that MEQA 'highlights the benchmarks' strengths and weaknesses' and 'enable[s] meaningful intra-benchmark comparisons.' These claims should be tempered until the instrument is validated against an external criterion, or the paper should be reframed as a proposal with pilot results.
minor comments (6)
  1. [Section 3] The word 'bechmarks' is a typo; it should read 'benchmarks.'
  2. [Table 3] The row 'Memorization Detection' contains a stray LaTeX command 'textit' in the text; this should be removed or rendered properly.
  3. [Table 8] The row 'Cross-Model Consistency' contains the typo 'appraches'; it should read 'approaches.'
  4. [Table 9] The row 'Representative data' contains the typo 'vulnerabilty'; it should read 'vulnerability.'
  5. [Table 6] The row 'Annotator Expertise' contains the typo 'Prefereable'; it should read 'Preferable.'
  6. [Appendix F] The code is described as 'available anonymously,' but the manuscript gives no repository identifier, URL, or anonymized link; please provide a way for readers to access the code.

Circularity Check

1 steps flagged · score 4.0 of 10

LLM-evaluator agreement is partly self-referential because few-shot prompts contained the human scores it is compared against; the core MEQA scoring framework itself is not circular.

  1. fitted input called prediction [Section 2.1 and Appendix C]
    "The automated scoring tended to match human evaluators, particularly when producing extreme scores (1 or 5). ... We attempted both zero-shot prompting and few-shot prompting as preliminary tests. In the latter case, we included scores for benchmarks that had been already evaluated by humans, this improved the LLM’s accuracy. Our results in this paper are those obtained using few-shot prompting."

    The paper presents GPT-4o's agreement with human evaluators as evidence that the LLM evaluator works, but the reported results were produced with few-shot prompts that included human scores for benchmarks already evaluated by humans. The model therefore had the target values in its context, so the observed 'match' is expected and is not independent confirmation. This is calibrating on the labels and then presenting the resulting alignment as a finding. The circularity is partial: it affects the LLM-evaluator validation claim, but not the human-score benchmark rankings, which are independent of the few-shot prompt contents.

full rationale

The central MEQA framework is not circular by construction: the eight criteria and 44 sub-criteria are explicitly synthesized from previously published evaluation literature, and the per-benchmark scores are the evaluators' 1-5 judgments using the rubric, not quantities derived from an equation that presupposes the result. The paper does not invoke any load-bearing self-citation, uniqueness theorem, or ansatz smuggled in from the authors' prior work; the references are external. The genuine circular step is confined to the LLM-evaluator component: Appendix C discloses that the reported GPT-4o results came from few-shot prompting that included human scores for already-evaluated benchmarks, and Section 2.1 then cites GPT-4o's tendency to match human evaluators as support. That agreement is contaminated by construction because the human scores were placed in the prompt as exemplars. The human-evaluator-based rankings and the framework's utility as a gap-analysis instrument are not circular, though their reliability is under-supported by the absence of kappa/ICC statistics and external validation; that is a validity concern rather than a circularity concern under the rules of this review. Overall, the paper merits a moderate circularity score because a supporting 'prediction-like' claim reduces to its calibration inputs, while the core derivation remains self-contained.

Assumptions & free parameters 1 free parameters · 4 assumptions · 0 invented entities

MEQA is a scoring rubric and not an invented physical or mathematical entity; no new objects are postulated. The framework's assumptions are mainly about the validity of evaluation judgments and the relevance of the chosen criteria, all resting on prior literature and the authors' domain assumptions.

free parameters (1)
  • Sub-criterion aggregation weights
    Overall criterion scores appear to be unweighted means of sub-criterion scores, an implicit choice that affects all reported benchmark rankings.
assumptions (4)
  • domain assumption The eight criteria are the relevant and sufficient dimensions of QA benchmark quality.
    Synthesized from prior meta-evaluation work; no empirical demonstration of exhaustiveness or independence is provided.
  • domain assumption Human expert evaluation is the ground truth for benchmark quality.
    Human scores serve as the reference for LLM calibration; no independent objective measure of benchmark quality is used.
  • domain assumption GPT-4o scores are a valid proxy for human evaluation.
    Only qualitative evidence of agreement is provided; no quantitative validation (e.g., correlation or kappa) is reported.
  • domain assumption The 1-5 rating scale yields meaningful, comparable scores across sub-criteria.
    No psychometric validation of the scale or rubric is presented; the granularity is asserted rather than demonstrated.

how reviews work

0 comments
Cite this review

Pith. "Pith review of MEQA: A Meta-Evaluation Framework for Question & Answer LLM Benchmarks." pith.science (2026). https://pith.science/paper/JEDYDDSZ

@misc{pith2026250414039,
  author       = {Pith},
  title        = {Pith review of: MEQA: A Meta-Evaluation Framework for Question & Answer LLM Benchmarks},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/JEDYDDSZ}},
  note         = {Machine review of arXiv:2504.14039}
}
read the original abstract

As Large Language Models (LLMs) advance, their potential for widespread societal impact grows simultaneously. Hence, rigorous LLM evaluations are both a technical necessity and social imperative. While numerous evaluation benchmarks have been developed, there remains a critical gap in meta-evaluation: effectively assessing benchmarks' quality. We propose MEQA, a framework for the meta-evaluation of question and answer (QA) benchmarks, to provide standardized assessments, quantifiable scores, and enable meaningful intra-benchmark comparisons. We demonstrate this approach on cybersecurity benchmarks, using human and LLM evaluators, highlighting the benchmarks' strengths and weaknesses. We motivate our choice of test domain by AI models' dual nature as powerful defensive tools and security threats.

Figures

Figures reproduced from arXiv: 2504.14039 by the authors.

Figure 1
Figure 1. Scores of cybersecurity benchmarks per criterion. N/A indicates inapplicable criteria (e.g. SECURE uses pre-defined correct answers; evaluator design does not apply). 3 [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. Scores of cybersecurity benchmarks across memorization robustness sub-criteria. 12 [PITH_FULL_IMAGE:figures/full_fig_p012_2.png] view at source ↗
Figure 3
Figure 3. Scores of cybersecurity benchmarks across prompt robustness sub-criteria. 13 [PITH_FULL_IMAGE:figures/full_fig_p013_3.png] view at source ↗
Figures from the paper (6 more)
Figure 4
Figure 4. Figure 4: Scores of cybersecurity benchmarks across evaluation design sub-criteria. 14 [PITH_FULL_IMAGE:figures/full_fig_p014_4.png]
Figure 5
Figure 5. Figure 5: Scores of cybersecurity benchmarks across evaluator design sub-criteria. 15 [PITH_FULL_IMAGE:figures/full_fig_p015_5.png]
Figure 6
Figure 6. Figure 6: Scores of cybersecurity benchmarks across reproducibility sub-criteria. 16 [PITH_FULL_IMAGE:figures/full_fig_p016_6.png]
Figure 7
Figure 7. Figure 7: Scores of cybersecurity benchmarks across comparability sub-criteria. 17 [PITH_FULL_IMAGE:figures/full_fig_p017_7.png]
Figure 8
Figure 8. Figure 8: Scores of cybersecurity benchmarks across validity sub-criteria. 18 [PITH_FULL_IMAGE:figures/full_fig_p018_8.png]
Figure 9
Figure 9. Figure 9: Scores of cybersecurity benchmarks across reliability sub-criteria. 19 [PITH_FULL_IMAGE:figures/full_fig_p019_9.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

27 extracted references · 4 canonical work pages

  1. [1]

    H., Sanh, V ., Yong, Z

    Bach, S. H., Sanh, V ., Yong, Z. X., Webson, A., Raffel, C., Nayak, N. V ., Sharma, A., Kim, T., Bari, M. S., F´evry, T., Alyafeai, Z., Dey, M., Santilli, A., Sun, Z., Ben-David, S., Xu, C., Chhablani, G., Wang, H., Fries, J. A., AlShaibani, M. S., Sharma, S., Thakker, U., Almubarak, K., Tang, X., Jiang, M. T., and Rush, A. M. Promptsource: An integrated ...

  2. [2]

    P., Kapil, D., Kozyrakis, Y ., LeBlanc, D., Milazzo, J., Straumann, A., Synnaeve, G., V on- timitta, V ., Whitman, S., and Saxe, J

    Bhatt, M., Chennabasappa, S., Nikolaidis, C., Wan, S., Ev- timov, I., Gabi, D., Song, D., Ahmad, F., Aschermann, C., Fontana, L., Frolov, S., Giri, R. P., Kapil, D., Kozyrakis, Y ., LeBlanc, D., Milazzo, J., Straumann, A., Synnaeve, G., V on- timitta, V ., Whitman, S., and Saxe, J. Purple llama cybersece- val: A secure coding benchmark for language models...

  3. [3]

    Cyberseceval 2: A wide-ranging cybersecurity evaluation suite for large language models, 2024

    Bhatt, M., Chennabasappa, S., Li, Y ., Nikolaidis, C., Song, D., Wan, S., Ahmad, F., Aschermann, C., Chen, Y ., Kapil, D., Molnar, D., Whitman, S., and Saxe, J. Cyberseceval 2: A wide-ranging cybersecurity evaluation suite for large language models, 2024. URL https://arxiv.org/ abs/2404.13161

  4. [4]

    T., Nguyen, L., Mahara, A., Lightcap, Z., Frazier, R., Fieblinger, R., Torales, G

    Bhusal, D., Alam, M. T., Nguyen, L., Mahara, A., Lightcap, Z., Frazier, R., Fieblinger, R., Torales, G. L., and Rastogi, N. Secure: Benchmarking large language models for cybersecu- rity advisory, 2024. URL https://arxiv.org/abs/ 2405.20441

  5. [5]

    F., Ammanamanchi, P

    Biderman, S., Schoelkopf, H., Sutawika, L., Gao, L., Tow, J., Abbasi, B., Aji, A. F., Ammanamanchi, P. S., Black, S., Clive, J., DiPofi, A., Etxaniz, J., Fattori, B., Forde, J. Z., Foster, C., Hsu, J., Jaiswal, M., Lee, W. Y ., Li, H., Lovering, C., 2 MEQA Figure 1. Scores of cybersecurity benchmarks per criterion. N/A indicates inapplicable criteria (e.g...

  6. [6]

    Clark, E., August, T., Serrano, S., Haduong, N., Gururangan, S., and Smith, N. A. All that’s ‘human’ is not gold: Evaluat- ing human evaluation of generated text. In Zong, C., Xia, F., Li, W., and Navigli, R. (eds.),Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natura...

  7. [7]

    Evaluating superhuman models with consistency checks, 2023

    Fluri, L., Paleka, D., and Tram`er, F. Evaluating superhuman models with consistency checks, 2023. URL https:// arxiv.org/abs/2306.09983

  8. [8]

    A framework for few-shot language model evaluation, 12 2023

    Gao, L., Tow, J., Abbasi, B., Biderman, S., Black, S., DiPofi, A., Foster, C., Golding, L., Hsu, J., Le Noac’h, A., Li, H., Mc- Donell, K., Muennighoff, N., Ociepa, C., Phang, J., Reynolds, L., Schoelkopf, H., Skowron, A., Sutawika, L., Tang, E., Thite, A., Wang, B., Wang, K., and Zou, A. A framework for few-shot language model evaluation, 12 2023. URL ht...

Show all 27 references
  1. [9]

    Jacobs, A. Z. and Wallach, H. M. Measurement and fair- ness. CoRR, abs/1912.05511, 2019. URL http://arxiv. org/abs/1912.05511

  2. [10]

    Sevenllm: Benchmarking, eliciting, and enhancing abilities of large language models in cyber threat intelligence, 2024

    Ji, H., Yang, J., Chai, L., Wei, C., Yang, L., Duan, Y ., Wang, Y ., Sun, T., Guo, H., Li, T., Ren, C., and Li, Z. Sevenllm: Benchmarking, eliciting, and enhancing abilities of large language models in cyber threat intelligence, 2024. URL https://arxiv.org/abs/2405.03446

  3. [11]

    Seceval: A comprehensive benchmark for eval- uating cybersecurity knowledge of foundation models

    Li, G., Li, Y ., Guannan, W., Yang, H., and Yu, Y . Seceval: A comprehensive benchmark for eval- uating cybersecurity knowledge of foundation models. https://github.com/XuanwuAI/SecEval, 2023

  4. [12]

    D., Dombrowski, A.-K., Goel, S., Phan, L., Mukobi, G., Helm-Burger, N., Lababidi, R., Justen, L., Liu, A

    Li, N., Pan, A., Gopal, A., Yue, S., Berrios, D., Gatti, A., Li, J. D., Dombrowski, A.-K., Goel, S., Phan, L., Mukobi, G., Helm-Burger, N., Lababidi, R., Justen, L., Liu, A. B., Chen, M., Barrass, I., Zhang, O., Zhu, X., Tamirisa, R., Bharathi, B., Khoja, A., Zhao, Z., Herbert...

  5. [13]

    ROUGE: A package for automatic evaluation of summaries

    Lin, C.-Y . ROUGE: A package for automatic evaluation of summaries. In Text Summarization Branches Out , pp. 74–81, Barcelona, Spain, July 2004. Association for Com- putational Linguistics. URL https://aclanthology. org/W04-1013

  6. [14]

    Secqa: A concise question-answering dataset for evaluating large language models in computer security, 2023

    Liu, Z. Secqa: A concise question-answering dataset for evaluating large language models in computer security, 2023. URL https://arxiv.org/abs/2312.15838

  7. [15]

    Harmbench: A standardized evaluation frame- work for automated red teaming and robust refusal, 2024

    Mazeika, M., Phan, L., Yin, X., Zou, A., Wang, Z., Mu, N., Sakhaee, E., Li, N., Basart, S., Li, B., Forsyth, D., and Hendrycks, D. Harmbench: A standardized evaluation frame- work for automated red teaming and robust refusal, 2024. URL https://arxiv.org/abs/2402.04249

  8. [16]

    State of what art? a call for multi-prompt llm evaluation, 2024

    Mizrahi, M., Kaplan, G., Malkin, D., Dror, R., Shahaf, D., and Stanovsky, G. State of what art? a call for multi-prompt llm evaluation, 2024. URL https://arxiv.org/abs/ 2401.00595

  9. [17]

    H., Fitz, S., and Hendrycks, D

    Ren, R., Basart, S., Khoja, A., Gatti, A., Phan, L., Yin, X., Mazeika, M., Pan, A., Mukobi, G., Kim, R. H., Fitz, S., and Hendrycks, D. Safetywashing: Do ai safety benchmarks actually measure safety progress?, 2024. URL https:// arxiv.org/abs/2407.21792

  10. [18]

    Jinja2 Documentation, 2024

    Ronacher, A. Jinja2 Documentation, 2024. URL https: //jinja.palletsprojects.com/. Accessed: 2024- 09-17

  11. [19]

    First tragedy, then parse: History repeats itself in the new era of large language models, 2024

    Saphra, N., Fleisig, E., Cho, K., and Lopez, A. First tragedy, then parse: History repeats itself in the new era of large language models, 2024. URL https://arxiv.org/ abs/2311.05020

  12. [20]

    Quantifying language models’ sensitivity to spurious features in prompt design or: How i learned to start worrying about prompt for- matting, 2024

    Sclar, M., Choi, Y ., Tsvetkov, Y ., and Suhr, A. Quantifying language models’ sensitivity to spurious features in prompt design or: How i learned to start worrying about prompt for- matting, 2024. URL https://arxiv.org/abs/2310. 11324

  13. [21]

    Large language models are inconsistent and biased evaluators, 2024

    Stureborg, R., Alikaniotis, D., and Suhara, Y . Large language models are inconsistent and biased evaluators, 2024. URL https://arxiv.org/abs/2405.01724

  14. [22]

    Subramonian, A., Yuan, X., au2, H. D. I., and Blodgett, S. L. It takes two to tango: Navigating conceptualizations of nlp tasks and measurements of performance, 2023. URL https://arxiv.org/abs/2305.09022

  15. [23]

    A., Jain, R., Bisztray, T., and Debbah, M

    Tihanyi, N., Ferrag, M. A., Jain, R., Bisztray, T., and Debbah, M. Cybermetric: A benchmark dataset based on retrieval- augmented generation for evaluating llms in cybersecurity knowledge, 2024. URL https://arxiv.org/abs/ 2402.07688

  16. [24]

    Xiao, Z., Zhang, S., Lai, V ., and Liao, Q. V . Evaluating evaluation metrics: A framework for analyzing nlg evaluation metrics using measurement theory, 2023. URL https: //arxiv.org/abs/2305.14889

  17. [25]

    Skill-mix: a flexible and expandable family of evaluations for ai models, 2023

    Yu, D., Kaur, S., Gupta, A., Brown-Cohen, J., Goyal, A., and Arora, S. Skill-mix: a flexible and expandable family of evaluations for ai models, 2023. URL https://arxiv. org/abs/2310.17567

  18. [26]

    P., Zhang, H., Gonzalez, J

    Zheng, L., Chiang, W.-L., Sheng, Y ., Zhuang, S., Wu, Z., Zhuang, Y ., Lin, Z., Li, Z., Li, D., Xing, E. P., Zhang, H., Gonzalez, J. E., and Stoica, I. Judging llm-as-a-judge with mt-bench and chatbot arena, 2023. URL https://arxiv. org/abs/2306.05685

  19. [27]

    canary strings

    Zhu, K., Chen, J., Wang, J., Gong, N. Z., Yang, D., and Xie, X. Dyval: Dynamic evaluation of large language models for reasoning tasks, 2024. URL https://arxiv.org/ abs/2309.17167. 4 MEQA Appendices A. Criteria and sub-criteria We outline our meta-evaluation criteria in Table ...

Pith tools

Reviewed August 16, 2026 · model on record in the stance chip above.