Pith. sign in

REVIEW 4 major objections 6 minor 1 cited by

WGSR-Bench: Wargame-based Game-theoretic Strategic Reasoning Benchmark for Large Language Models

T0 review · 4 major / 6 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read The paper argues that current LLMs lag expert human strategists by large margins on a new 1,208-question wargame benchmark built from real adversarial replays.

desk verdict A promising wargame-based strategic reasoning benchmark, but the headline human-AI gaps rest on an unvalidated LLM judge and unreleased data, so the numbers are not yet established. read the letter →

arxiv 2506.10264 v1 pith:7GURODNM submitted 2025-06-12 cs.AI

classification cs.AI
keywords strategicreasoninglargelanguagemodelswargamebenchmarkgame-theoreticsituationalawarenessopponentmodelingpolicygeneration
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

WGSR-Bench sets out to make strategic reasoning measurable for LLMs by turning wargame play into 1,208 questions about reading a battlefield, reading an opponent, and choosing a policy. The paper's central claim is that this three-part structure, the S-POE architecture, covers the main sub-capacities of strategic reasoning, and that current large models are far behind trained humans on all three. On situational awareness the best model scores 58.2 versus 79.9 for humans; on opponent modeling the best model scores 49.1 versus 85.7 for professionals; on policy generation GPT-4.1 scores 60.0 versus 92.3 for elite strategists. If the benchmark is a fair test, then state-of-the-art LLMs have genuine but shallow strategic ability: they approach humans on simple spatial judgments and beat regular-level players, while collapsing on high-risk prediction, multi-agent coordination, and long-horizon planning.

What carries the argument

The carrying structure is the S-POE cognitive framework, which splits strategic reasoning into situational awareness, opponent risk assessment, and policy generation, and the wargame environment that supplies the questions. Wargame is defined as a high-complexity adversarial simulation with incomplete information, multi-agent dynamics, and no single best move; the paper uses real replays from such games as the source of all Q&A pairs. For PGG-Bench the evaluation is carried by a six-dimension rubric adapted from Bloom's taxonomy, with Qwen-2.5-VL-72B as the automated judge producing structured JSON scores.

What would settle it

Take a random sample of PGG-Bench responses from elite humans and from GPT-4.1, have several independent human experts score them with the same six-dimension rubric, and compare the expert rankings with the Qwen-2.5-VL-72B rankings; if the experts do not clearly rank elite humans above GPT-4.1, the reported 32.3-point gap is an artifact of the scorer.

Watch

Extended reading notes

Core claim

On its own terms, the paper establishes that LLM strategic reasoning can be decomposed and evaluated through wargame-derived questions, and that current models underperform trained humans at every level of that decomposition. The benchmark samples 1,208 Q&A pairs from a large real replay database and organizes them into MM-SA-Bench (environmental situational awareness, 424 pairs), PsyR-OM-Bench (opponent risk and reward modeling, 420 pairs), and PGG-Bench (policy generation across game-theoretic paradigms, 364 pairs). The reported results show a consistent ordering: human experts highest, professional and general humans next, then the best models, with the largest model deficits in complex situation analysis, high-risk opponent prediction, coalition coordination, and strategies beyond roughly 150 steps. The paper also integrates the three tasks into an LLM wargame agent, with preliminary win rates that it reads as evidence that strategic LLM agents are still at an early stage.

Load-bearing premise

The whole comparison rests on the benchmark's answer keys and automated scores actually being measures of good strategy; the paper does not report expert validation or human-judge agreement for either.

Editorial extensions

If this is right

  • If WGSR-Bench measures what it claims, deploying current LLMs as autonomous strategic decision-makers in high-stakes wargame-like settings is not yet supported by their measured performance.
  • The benchmark gives separable targets for improvement: closing the 48.4-point gap on complex situation analysis, the roughly 33-point deficit in coalition coordination, and the sharp decline in planning beyond about 150 steps are distinct engineering problems.
  • Because the three sub-benchmarks are modular, a model can be strong in one component and weak in another, so progress can be tracked per sub-capacity instead of through a single end-to-end win rate.
  • Because the questions come from a growing real replay database, the benchmark can be extended beyond 1,208 pairs without rebuilding the task suite.
  • The reported wargame-agent win rates, ranging roughly from 0.05 to 0.65 across scenarios, indicate that end-to-end agent behavior lags the component-level abilities that the Q&A tasks measure separately.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Editorial inference: the consistent model weakness on short-horizon and high-risk choices, alongside relative strength on long-term reward, hints that LLMs carry a cautious-planning prior from pretraining; the paper notes the asymmetry but does not directly test this explanation.
  • Editorial inference: because no human-judge agreement is reported for the automated scorer, the precise numeric gaps in PGG-Bench should be read as provisional even if the qualitative ordering, experts above current models, is plausible.
  • Editorial inference: the S-POE decomposition could be lifted out of the military domain and applied to negotiation, cybersecurity, and market competition; the wargame data is the testbed, not the definition of strategic reasoning itself.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. The paper introduces WGSR-Bench, a wargame-based benchmark for evaluating strategic reasoning in large language models, containing 1,208 Q&A pairs across three sub-benchmarks (MM-SA-Bench for situational awareness, PsyR-OM-Bench for opponent modeling, and PGG-Bench for policy generation). All items are claimed to be sourced from a real adversarial replay database of over 400,000 battle replays. The authors evaluate a range of LLMs and human participants stratified into three expertise levels, reporting large human-AI gaps (e.g., PGG-Bench: GPT-4.1 60.0 vs elite humans 92.3; MM-SA: best model 58.2 vs humans 79.9; PsyR: best LLM 49.1 vs 77.9 for general-level humans). The paper also sketches an LLM-based wargame agent built on the OODA loop. The central methodological issue is that the validity of the ground-truth labels and of the LLM-based judge is not established, so the reported quantitative gaps must be treated as provisional.

Significance. If the validity of the ground truth and scoring can be established, WGSR-Bench would be a useful contribution: it draws on a large external replay database (400k replays, 225 scenarios), spans three sub-capabilities under a coherent S-POE/OODA framing, evaluates a broad set of LLMs against stratified human groups, and the PGG-Bench scoring design uses a structured rubric grounded in assessment theory. The paper also honestly admits in Section 6 that integration of the three modules 'remains under investigation.' However, the current manuscript does not demonstrate label validity or judge reliability, and it contains internal count inconsistencies; until those are addressed, the headline quantitative claims should be viewed as unverified.

major comments (4)
  1. [§5.3.5] The central PGG-Bench scores are produced by Qwen-2.5-VL-72B with a six-dimensional rubric, but the paper reports no agreement between this judge and human expert ratings, no calibration, and no consistency analysis. Because the headline gaps (e.g., GPT-4.1 60.0 vs elite humans 92.3 in Section 5.5.1) are computed from these judge scores, the validity of all PGG-Bench conclusions rests on an unvalidated LLM judge. Please provide a human-judge correlation study on a held-out sample (e.g., Cohen's kappa or Spearman correlation), dimension-level agreement, and confidence intervals for the reported scores.
  2. [§2 and §5.2/5.3.1] The manuscript states in Section 2 that PGG-Bench is 'subdivided into 29 subtasks,' but Sections 5.2 and 5.3.1 state there are '28 distinct decision types' across 364 Q&A pairs. If the benchmark structure is uncertain at this level, the aggregated numbers cannot be interpreted. Please reconcile the counts and specify the exact mapping from game types to subtasks to individual items.
  3. [§3.3 and §4.3] For MM-SA-Bench and PsyR-OM-Bench, the correct answers are said to be 'exclusively sourced from a real adversarial database,' but sourcing from replays does not determine a unique correct answer for the multiple-choice questions. The paper reports no annotation protocol, no expert adjudication, and no inter-annotator agreement for these labels. Without evidence that the ground truth is valid, the reported human-model gaps (e.g., 79.9 vs 58.2 in MM-SA) are not yet established. Please provide the annotation guidelines, the number of annotators, and agreement statistics.
  4. [§5.5.2] The text refers to 'seven critical evaluation dimensions' in the radar chart, while Section 5.3.5 defines a six-dimensional rubric (Factual Correctness, Logical Consistency, Game Principles Adherence, Outcome Prediction, Clarity and Completeness, Innovation Bonus); additionally, Figure 15 appears to include a 'Similarity' dimension that is not part of the rubric. This inconsistency makes the multi-dimensional capability analysis ambiguous and should be corrected and tied to the defined rubric.
minor comments (6)
  1. [Throughout] There are numerous typos and grammatical errors (e.g., 'evalution', 'planing', 'Minecrarft', 'Breakdownn', 'polices', 'T he'); a careful proofread is needed.
  2. [§7] The conclusion states that WGSR-Bench features 'a 4-layer structure, 9 object categories, 39 action types,' none of which are defined or used elsewhere in the paper; either define these quantities or remove them.
  3. [Figures] The figures are referenced by number but their content is not included in the manuscript rendering; please ensure the figures are embedded and self-contained, and that all ranking and radar-chart data are also reported in text or a supplement.
  4. [Availability] No release URL, data license, or code repository is provided, so the benchmark cannot be reproduced or extended by other groups; please include an availability statement.
  5. [§3.4-§5.5] All quantitative comparisons are reported as point estimates without confidence intervals, significance tests, or multiple-run variance; given the random sampling of items and stochastic LLM inference, basic uncertainty quantification should be added.
  6. [§1] The related work discussion does not quantitatively compare WGSR-Bench to existing benchmarks such as GTBench, AvalonBench, or SmartPlay on any shared metric, which weakens the novelty claim; at minimum, discuss how difficulty and validity compare with these benchmarks.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity; the benchmark is grounded in an external replay database, and the LLM-as-judge concern is a validity gap rather than a circular derivation.

full rationale

WGSR-Bench is an evaluation artifact rather than a derivation, and its central quantities are not defined in terms of the claims they support. MM-SA-Bench and PsyR-OM-Bench use multiple-choice questions whose correct answers are grounded in the external MiaoSuan replay database ('exclusively sourced from a real adversarial database, which contains over 400,000 battle replays'), so the labels are not constructed from the models being scored. PGG-Bench's open-ended responses are scored by Qwen-2.5-VL-72B under a six-dimension rubric; this is an LLM-as-judge measurement choice, and while the paper reports no human-judge correlation, the judge is a fixed external model, not one of the evaluated systems, and the same judge is applied to human and AI responses. No fitted parameter is later renamed as a prediction, no self-citation supplies the load-bearing validity of the benchmark, and no uniqueness theorem is imported from the authors' prior work. The only self-citation ([21]) defines wargame as a high-complexity scenario and is not load-bearing. The 29-vs-28 subtask inconsistency is an internal consistency issue, not circularity. The absence of expert adjudication and inter-annotator agreement is a validity risk, but it does not make the benchmark's claims equivalent to their inputs by construction.

Assumptions & free parameters 2 free parameters · 4 assumptions · 1 invented entities

The central evaluation rests on assumptions about data validity and scoring validity rather than on mathematical derivation. There are no fitted parameters in the usual sense, but the automated scoring weights are undisclosed, and both the replay-derived labels and the S-POE decomposition are taken as given. These assumptions are not independently established in the paper.

free parameters (2)
  • PGG-Bench scoring dimension weights = not disclosed
    The final 0-100 score is a weighted aggregation with weights 'adjusted based on task characteristics', but the weights are never specified, so any PGG-Bench score is not uniquely reproducible.
  • Human expertise-level assignment criteria = not disclosed
    Participants are split into elite, professional, and general levels without quantitative thresholds or counts, so human baseline comparisons cannot be reconstructed.
assumptions (4)
  • domain assumption The MiaoSuan replay database is a valid source of strategic ground truth.
    Section 2 states Q&A pairs are 'exclusively sourced from a real adversarial database' with over 400,000 replays, but no validation of answer correctness or expert review is described.
  • domain assumption The multiple-choice questions have unambiguous correct answers.
    Section 3.3 says all tasks use a unified multiple-choice format, but no item statistics, distractor analysis, or human agreement data are provided.
  • ad hoc to paper Qwen-2.5-VL-72B produces valid strategic reasoning scores.
    Section 5.3.5 assigns this LLM as the primary scoring model without reporting correlation with human ratings or calibration against known strategic outcomes.
  • domain assumption The S-POE decomposition (situation awareness, opponent modeling, policy generation) is a sufficient decomposition of strategic reasoning.
    Section 1 introduces S-POE as the cognitive framework, but no empirical or theoretical justification is given that these three components cover strategic reasoning.
invented entities (1)
  • S-POE structured cognitive framework
    purpose: Organizes WGSR-Bench into three sub-benchmarks for policy generation, opponent risk assessment, and environmental situation awareness.
    The framework is introduced as the paper's own design; no independent validation shows these three components are the correct or complete decomposition of strategic reasoning.

how reviews work

0 comments
Cite this review

Pith. "Pith review of WGSR-Bench: Wargame-based Game-theoretic Strategic Reasoning Benchmark for Large Language Models." pith.science (2026). https://pith.science/paper/7GURODNM

@misc{pith2026250610264,
  author       = {Pith},
  title        = {Pith review of: WGSR-Bench: Wargame-based Game-theoretic Strategic Reasoning Benchmark for Large Language Models},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/7GURODNM}},
  note         = {Machine review of arXiv:2506.10264}
}
read the original abstract

Recent breakthroughs in Large Language Models (LLMs) have led to a qualitative leap in artificial intelligence' s performance on reasoning tasks, particularly demonstrating remarkable capabilities in mathematical, symbolic, and commonsense reasoning. However, as a critical component of advanced human cognition, strategic reasoning, i.e., the ability to assess multi-agent behaviors in dynamic environments, formulate action plans, and adapt strategies, has yet to be systematically evaluated or modeled. To address this gap, this paper introduces WGSR-Bench, the first strategy reasoning benchmark for LLMs using wargame as its evaluation environment. Wargame, a quintessential high-complexity strategic scenario, integrates environmental uncertainty, adversarial dynamics, and non-unique strategic choices, making it an effective testbed for assessing LLMs' capabilities in multi-agent decision-making, intent inference, and counterfactual reasoning. WGSR-Bench designs test samples around three core tasks, i.e., Environmental situation awareness, Opponent risk modeling and Policy generation, which serve as the core S-POE architecture, to systematically assess main abilities of strategic reasoning. Finally, an LLM-based wargame agent is designed to integrate these parts for a comprehensive strategy reasoning assessment. With WGSR-Bench, we hope to assess the strengths and limitations of state-of-the-art LLMs in game-theoretic strategic reasoning and to advance research in large model-driven strategic intelligence.

Figures

Figures reproduced from arXiv: 2506.10264 by the authors.

Figure 1
Figure 1. Overall framework of WGSR-Bench. introduced MAgIC [16], which consists of two social de￾duction games and three game-theory scenarios, to evaluate LLMs’ reasoning, planing and other social abilities. Wu et al. developed SmartPlay [17], a games collection, to evaluate LLMs as agents, where reasoning, planning, and learning from history abilities are required. In addition to the aforementioned two categories of strate… view at source ↗
Figure 2
Figure 2. Data used for WGSR-Bench. 2 FRAMEWORK AND DATA DESCRIPTION The overall framework of WGSR-Bench is displayed in Fig￾ure 1. MM-SA-Bench focuses on three tasks: object recogni￾tion, spatial relationship recognition, and situational reason￾ing analysis. It is further divided into 7 subtasks, including environmental identification, friend-foe positional relation￾ships, and advantage/disadvantage assessment, comprising a … view at source ↗
Figure 3
Figure 3. Overall ranking comparison showing detailed performance break [PITH_FULL_IMAGE:figures/full_fig_p004_3.png] view at source ↗
Figures from the paper (11 more)
Figure 4
Figure 4. Figure 4: Human-AI overall performance comparison in MM-SA-Bench [PITH_FULL_IMAGE:figures/full_fig_p004_4.png]
Figure 5
Figure 5. Figure 5: Six core capability scores comparison across humans and AI [PITH_FULL_IMAGE:figures/full_fig_p005_5.png]
Figure 6
Figure 6. Figure 6: Capability radar chart showing multi-dimensional performance [PITH_FULL_IMAGE:figures/full_fig_p005_6.png]
Figure 8
Figure 8. Figure 8: Human-AI overall performance comparison in PsyR-OM-Bench [PITH_FULL_IMAGE:figures/full_fig_p007_8.png]
Figure 10
Figure 10. Figure 10: Performance across different subtasks showing task-specific [PITH_FULL_IMAGE:figures/full_fig_p008_10.png]
Figure 11
Figure 11. Figure 11: Unified evaluation process for PGG-Bench: identical scenarios presented to both humans and AI, with responses fed into a single LLM [PITH_FULL_IMAGE:figures/full_fig_p011_11.png]
Figure 12
Figure 12. Figure 12: Six-dimensional automatic scoring pipeline with structured JSON evaluation output [PITH_FULL_IMAGE:figures/full_fig_p012_12.png]
Figure 13
Figure 13. Figure 13: Overall ranking comparison showing detailed performance [PITH_FULL_IMAGE:figures/full_fig_p012_13.png]
Figure 15
Figure 15. Figure 15: Multi-dimensional capability radar chart showing performance [PITH_FULL_IMAGE:figures/full_fig_p013_15.png]
Figure 16
Figure 16. Figure 16: Performance across four game types showing task-specific [PITH_FULL_IMAGE:figures/full_fig_p013_16.png]
Figure 17
Figure 17. Figure 17: LLM-based wargame agent. complex negotiations. As models approach professional￾level performance, PGG-Bench will serve as a validation framework for deploying strategic AI systems in high￾stakes applications, providing confidence metrics for system reliability and cap…

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Multi-level Value Alignment in Agentic AI Systems: Survey and Perspectives

    cs.AI 2025-06 conditional novelty 4.0 of 10

    A survey proposes a macro-meso-micro value framework for agentic AI alignment and maps applications, methods, and benchmarks onto it.

Reference graph

Works this paper leans on

36 extracted references · 19 canonical work pages · cited by 1 Pith paper

  1. [1]

    A survey of large language models,

    W. X. Zhao, K. Zhou, J. Liet al., “A survey of large language models,”arXiv:2303.18223v16, 2025

  2. [2]

    Gpt-4 technical report,

    OpenAI, “Gpt-4 technical report,”arXiv:2303.08774v6, 2024

  3. [3]

    Deepseek-v3 technical report,

    DeepSeek-AI, “Deepseek-v3 technical report,”arXiv:2412.19437v2, 2024

  4. [4]

    From system 1 to system 2: A survey of reasoning large language models,

    Z.-Z. Li, D. Zhang, M.-L. Zhanget al., “From system 1 to system 2: A survey of reasoning large language models,” arXiv:2502.17419v4, 2025

  5. [5]

    Towards large reasoning models: A survey of reinforced reasoning with large language models,

    F. Xu, Q. Hao, Z. Zonget al., “Towards large reasoning models: A survey of reinforced reasoning with large language models,” arXiv:2501.09686v3, 2025

  6. [6]

    Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning,

    DeepSeek-AI, “Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning,”arXiv:2501.12948v1, 2025

  7. [7]

    Large language mod- els for mathematical reasoning: Progresses and challenges,

    R. Ahn, Janice Verma and R. Lou, “Large language mod- els for mathematical reasoning: Progresses and challenges,” arXiv:2402.00157v4, 2024

  8. [8]

    Symbol-llm: Leverage language models for symbolic system in visual human activity reasoning,

    X. Wu, Y.-L. Li, J. Sun, and C. Lu, “Symbol-llm: Leverage language models for symbolic system in visual human activity reasoning,” inAdvances in Neural Information Processing Systems, 2023

Show all 36 references
  1. [9]

    Large language models as com- monsense knowledge for large-scale task planning,

    Z. Zhao, W. S. Lee, and D. Hsu, “Large language models as com- monsense knowledge for large-scale task planning,” inAdvances in Neural Information Processing Systems, 2023

  2. [10]

    Llm as a mastermind: A survey of strategic reasoning with large language models,

    Y. Zhang, S. Mao, T. Geet al., “Llm as a mastermind: A survey of strategic reasoning with large language models,” arXiv:2404.01230v1, 2024

  3. [11]

    Towards reasoning era: A survey of long chain-of-thought for reasoning large language models,

    Q. Chen, L. Qin, J. Liuet al., “Towards reasoning era: A survey of long chain-of-thought for reasoning large language models,” arXiv:2503.09567v3, 2025

  4. [12]

    Avalonbench: Evaluating llms playing the game of avalon,

    J. Light, M. Cai, S. Shen, and Z. Hu, “Avalonbench: Evaluating llms playing the game of avalon,”arXiv:2310.05036v3, 2023

  5. [13]

    Large language models play starcraft ii: Benchmarks and a chain of summarization approach,

    Z. Li, C. Lu, X. Xuet al., “Large language models play starcraft ii: Benchmarks and a chain of summarization approach,” inAdvances in Neural Information Processing Systems, 2024

  6. [14]

    Hierarchical expert prompt for large- language-model: An approach defeat elite ai in textstarcraft ii for the first time,

    Z. Li, C. Lu, and X. Xu, “Hierarchical expert prompt for large- language-model: An approach defeat elite ai in textstarcraft ii for the first time,”arXiv:2502.11122v1, 2025

  7. [15]

    Gtbench: Uncovering the strategic reasoning limitations of llms via game-theoretic evalua- tions,

    J. Duan, R. Zhang, J. Diffenderferet al., “Gtbench: Uncovering the strategic reasoning limitations of llms via game-theoretic evalua- tions,”arXiv:2402.12348v2, 2024

  8. [16]

    Magic: Investigation of large language model powered multi-agent in cognition, adaptability, rationality and collaboration,

    L. Xu, Z. Hu, D. Zhouet al., “Magic: Investigation of large language model powered multi-agent in cognition, adaptability, rationality and collaboration,”arXiv:2311.08562v3, 2023

  9. [17]

    Smartplay: A benchmark for llms as intelligent agents,

    Y. Wu, X. Tang, T. M. Mitchell, and Y. Li, “Smartplay: A benchmark for llms as intelligent agents,”arXiv:2310.01557v5, 2024

  10. [18]

    Mineplanner: A bench- mark for long-horizon planning in large minecraft worlds,

    W. Hill, I. Liu, I. A. D. M. Kochet al., “Mineplanner: A bench- mark for long-horizon planning in large minecraft worlds,” arXiv:2312.12891v2, 2024

  11. [19]

    Benchmarking agentic workflow generation,

    S. Qiao, R. Fang, Z. Qiuet al., “Benchmarking agentic workflow generation,”arXiv:2410.07869v3, 2025

  12. [20]

    Opendeception: Bench- marking and investigating ai deceptive behaviors via open-ended interaction simulation,

    Y. Wu, X. Pan, G. Hong, and M. Yang, “Opendeception: Bench- marking and investigating ai deceptive behaviors via open-ended interaction simulation,”arXiv:2504.13707v1, 2025

  13. [21]

    Intelligent decision making technol- ogy and challenge of wargame,

    Q. Yin, M. Zhao, W. Niet al., “Intelligent decision making technol- ogy and challenge of wargame,”Acta Automatica Sinica, vol. 49, pp. 913–928, 2023

  14. [22]

    L. W. Anderson and D. R. Krathwohl,A Taxonomy for Learning, Teaching, and Assessing: A Revision of Bloom’s Taxonomy of Educa- tional Objectives. Boston, MA: Allyn & Bacon, 2001

  15. [23]

    A five- dimensional framework for authentic assessment,

    J. T. M. Gulikers, T. J. Bastiaens, and P . A. Kirschner, “A five- dimensional framework for authentic assessment,”Educational Technology Research and Development, vol. 52, no. 3, pp. 67–86, 2004

  16. [24]

    Appropriate criteria: Key to effective rubrics,

    S. M. Brookhart, “Appropriate criteria: Key to effective rubrics,” Frontiers in Education, vol. 3, pp. 1–12, 2018, article 22

  17. [25]

    W. J. Popham,Modern Educational Measurement: Practical Guidelines for Educational Leaders, 3rd ed. Boston, MA: Allyn & Bacon, 2000

  18. [26]

    A review of rubric use in higher education,

    Y. M. Reddy and H. Andrade, “A review of rubric use in higher education,”Assessment & Evaluation in Higher Education, vol. 35, no. 4, pp. 435–448, 2010

  19. [27]

    Wiggins,Educative Assessment: Designing Assessments to Inform and Improve Student Performance

    G. Wiggins,Educative Assessment: Designing Assessments to Inform and Improve Student Performance. San Francisco, CA: Jossey-Bass, 1998

  20. [28]

    A revision of bloom’s taxonomy: An overview,

    D. R. Krathwohl, “A revision of bloom’s taxonomy: An overview,” Theory Into Practice, vol. 41, no. 4, pp. 212–218, 2002

  21. [29]

    B. S. Bloom, Ed.,Taxonomy of Educational Objectives: The Classifica- tion of Educational Goals. Handbook I: Cognitive Domain. New York: Longmans, Green, 1956

  22. [30]

    Improving instruction and assessment via bloom’s taxonomy and descriptive rubrics,

    K. R. Gosselin and N. Okamoto, “Improving instruction and assessment via bloom’s taxonomy and descriptive rubrics,” in Proceedings of the 2018 ASEE Annual Conference & Exposition, Salt Lake City, UT, 2018, paper 10.18260/1-2–30630

  23. [31]

    Llm-rubric: A multidimensional, calibrated approach to auto- mated evaluation of natural language texts,

    H. Hashemi, J. Eisner, C. Rosset, B. Van Durme, and C. Kedzie, “Llm-rubric: A multidimensional, calibrated approach to auto- mated evaluation of natural language texts,” inProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long P...

  24. [32]

    Gptscore: Evaluate as you desire,

    J. Fu, S.-K. Ng, Z. Jiang, and P . Liu, “Gptscore: Evaluate as you desire,” inProceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Lan- guage Technologies (Volume 1: Long Papers), 2024

  25. [33]

    G-eval: Nlg evaluation using gpt-4 with better human alignment,

    Y. Liu, A. R. Iter, Y. Xu, S. Wang, R. Xu, and L. Zhu, “G-eval: Nlg evaluation using gpt-4 with better human alignment,” in Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 2023

  26. [34]

    Large language models are not fair evaluators,

    P . Wang, L. Li, L. Chen, D. Zhu, B. Lin, Y. Cao, Q. Liu, T. Liu, and Z. Sui, “Large language models are not fair evaluators,” inProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2024

  27. [35]

    Realtoxicityprompts: Evaluating neural toxic degeneration in language models,

    S. Gehman, S. Gururangan, M. Sap, Y. Choi, and N. A. Smith, “Realtoxicityprompts: Evaluating neural toxic degeneration in language models,” inFindings of the Association for Computational Linguistics: EMNLP 2020, 2020, pp. 3356–3369

  28. [36]

    Why we need new evaluation metrics for nlg,

    J. Novikova, O. Du ˇsek, A. C. Curry, and V . Rieser, “Why we need new evaluation metrics for nlg,” inProceedings of the 2017 Conference on Empirical Methods in Natural Language Processing, Copenhagen, Denmark, 2017, pp. 2241–2252

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.