Pith. sign in

REVIEW 2 cited by

Are LLM-based Evaluators Confusing NLG Quality Criteria?

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.12055 v2 pith:SE26JVK2 submitted 2024-02-19 cs.CL

classification cs.CL
keywords criteriadifferentevaluationllmsclassificationfurtherissuesllm-based
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Some prior work has shown that LLMs perform well in NLG evaluation for different tasks. However, we discover that LLMs seem to confuse different evaluation criteria, which reduces their reliability. For further verification, we first consider avoiding issues of inconsistent conceptualization and vague expression in existing NLG quality criteria themselves. So we summarize a clear hierarchical classification system for 11 common aspects with corresponding different criteria from previous studies involved. Inspired by behavioral testing, we elaborately design 18 types of aspect-targeted perturbation attacks for fine-grained analysis of the evaluation behaviors of different LLMs. We also conduct human annotations beyond the guidance of the classification system to validate the impact of the perturbations. Our experimental results reveal confusion issues inherent in LLMs, as well as other noteworthy phenomena, and necessitate further research and improvements for LLM-based evaluation.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Transparent Reference-free Automated Evaluation of Open-Ended User Survey Responses

    cs.CL 2025-10 conditional novelty 5.0 of 10

    An LLM-judge framework with a gibberish filter scores human survey responses on effort, relevance, and completeness, matching expert ratings (Spearman up to 0.86 English) better than length, embedding, and DeepEval baselines.

  2. Decision Information Meets Large Language Models: The Future of Explainable Operations Research

    cs.AI 2025-02 conditional novelty 4.0 of 10

    An LLM framework that couples what-if analysis with graph edit distance on linear programs can generate more accurate and more detailed explanations for operations research queries than existing LLM baselines.

Pith tools