REVIEW 9 cited by
Evaluation of Text Generation: A Survey
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
The paper surveys evaluation methods of natural language generation (NLG) systems that have been developed in the last few years. We group NLG evaluation methods into three categories: (1) human-centric evaluation metrics, (2) automatic metrics that require no training, and (3) machine-learned metrics. For each category, we discuss the progress that has been made and the challenges still being faced, with a focus on the evaluation of recently proposed NLG tasks and neural NLG models. We then present two examples for task-specific NLG evaluations for automatic text summarization and long text generation, and conclude the paper by proposing future research directions.
Forward citations
Cited by 9 Pith papers
-
The Reader is the Metric: How Textual Features and Reader Profiles Explain Conflicting Evaluations of AI Creative Writing
Reader evaluations of AI versus human stories split into two measurable preference profiles, surface-focused and holistic, that track reader expertise and explain why prior studies disagree.
-
LLM-as-a-Judge for Evaluating System Responses in Conversational Music Recommendation
LLM judges agree moderately with human experts when scoring conversational music recommendation responses, outperform reference-based metrics, but are not reliable enough to replace human evaluation.
-
grapheme-kit: Grapheme-Level Metrics and Text Processing for Multilingual NLP
A new open-source library computes distance, similarity, and evaluation metrics on grapheme clusters rather than Unicode code points, and corrects ZWJ/ZWNJ segmentation for Tamil and Sinhala.
-
Preconditioned Test-Time Adaptation for Out-of-Distribution Debiasing in Narrative Generation
CAP-TTA triggers context-aware preconditioned LoRA updates on high bias-risk OOD prompts to reduce toxicity in LLM narrative generation while preserving fluency and avoiding catastrophic forgetting.
-
What Are We Measuring in NLG? A Meta-Analysis of Evaluation Trends 2020-2025
Across 3,334 NLG papers from 2020-2025, legacy n-gram metrics persist, LLM-as-a-judge usage outpaced human validation, and reported LLM-human correlations collapse on criteria like fluency.
-
COGENT: A Curriculum-oriented Framework for Generating Grade-appropriate Educational Content
Adding curriculum decomposition, readability constraints, and wonder-based topics to LLM prompts yields science reading passages with higher curriculum alignment and human-comparable comprehensibility.
-
Fidelity-Diversity Metrics for Text
Optimal-transport fidelity and diversity scores on text embeddings disentangle support mismatch from mass coverage and predict GSM8K finetuning accuracy on synthetic math data.
-
From Multimodal Perception to Strategic Reasoning: A Survey on AI-Generated Game Commentary
A survey that organizes AI-generated game commentary research into a taxonomy of three commentator capabilities and three commentary types, with a review of methods, datasets, and metrics.
-
FairJudge: An Adaptive, Debiased, and Consistent LLM-as-a-Judge
A three-stage SFT-DPO-GRPO curriculum on a JudgeLM-derived dataset improves LLM judge agreement/F1, but debiasing and consistency claims are evaluated on a benchmark built from the same pipeline.
Discussion (0). Continue with ORCID to comment.