Pith. sign in

REVIEW 9 cited by

Evaluation of Text Generation: A Survey

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2006.14799 v2 pith:5IEM3IDD submitted 2020-06-26 cs.CL cs.LG

classification cs.CLcs.LG
keywords evaluationgenerationmetricstextautomaticbeenmethodscategories
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The paper surveys evaluation methods of natural language generation (NLG) systems that have been developed in the last few years. We group NLG evaluation methods into three categories: (1) human-centric evaluation metrics, (2) automatic metrics that require no training, and (3) machine-learned metrics. For each category, we discuss the progress that has been made and the challenges still being faced, with a focus on the evaluation of recently proposed NLG tasks and neural NLG models. We then present two examples for task-specific NLG evaluations for automatic text summarization and long text generation, and conclude the paper by proposing future research directions.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 9 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 198 citations worldwide. Full citation record

  1. The Reader is the Metric: How Textual Features and Reader Profiles Explain Conflicting Evaluations of AI Creative Writing

    cs.CL 2025-06 conditional novelty 7.0 of 10

    Reader evaluations of AI versus human stories split into two measurable preference profiles, surface-focused and holistic, that track reader expertise and explain why prior studies disagree.

  2. LLM-as-a-Judge for Evaluating System Responses in Conversational Music Recommendation

    cs.IR 2026-07 conditional novelty 6.0 of 10

    LLM judges agree moderately with human experts when scoring conversational music recommendation responses, outperform reference-based metrics, but are not reliable enough to replace human evaluation.

  3. grapheme-kit: Grapheme-Level Metrics and Text Processing for Multilingual NLP

    cs.CL 2026-07 conditional novelty 6.0 of 10

    A new open-source library computes distance, similarity, and evaluation metrics on grapheme clusters rather than Unicode code points, and corrects ZWJ/ZWNJ segmentation for Tamil and Sinhala.

  4. Preconditioned Test-Time Adaptation for Out-of-Distribution Debiasing in Narrative Generation

    cs.CL 2026-03 unverdicted novelty 6.0 of 10

    CAP-TTA triggers context-aware preconditioned LoRA updates on high bias-risk OOD prompts to reduce toxicity in LLM narrative generation while preserving fluency and avoiding catastrophic forgetting.

  5. What Are We Measuring in NLG? A Meta-Analysis of Evaluation Trends 2020-2025

    cs.CL 2026-01 conditional novelty 6.0 of 10

    Across 3,334 NLG papers from 2020-2025, legacy n-gram metrics persist, LLM-as-a-judge usage outpaced human validation, and reported LLM-human correlations collapse on criteria like fluency.

  6. COGENT: A Curriculum-oriented Framework for Generating Grade-appropriate Educational Content

    cs.CL 2025-06 conditional novelty 6.0 of 10

    Adding curriculum decomposition, readability constraints, and wonder-based topics to LLM prompts yields science reading passages with higher curriculum alignment and human-comparable comprehensibility.

  7. Fidelity-Diversity Metrics for Text

    cs.CL 2026-07 conditional novelty 5.5 of 10

    Optimal-transport fidelity and diversity scores on text embeddings disentangle support mismatch from mass coverage and predict GSM8K finetuning accuracy on synthetic math data.

  8. From Multimodal Perception to Strategic Reasoning: A Survey on AI-Generated Game Commentary

    cs.CL 2025-06 conditional novelty 5.0 of 10

    A survey that organizes AI-generated game commentary research into a taxonomy of three commentator capabilities and three commentary types, with a review of methods, datasets, and metrics.

  9. FairJudge: An Adaptive, Debiased, and Consistent LLM-as-a-Judge

    cs.CL 2026-02 reject novelty 4.0 of 10

    A three-stage SFT-DPO-GRPO curriculum on a JudgeLM-derived dataset improves LLM judge agreement/F1, but debiasing and consistency claims are evaluated on a benchmark built from the same pipeline.

Pith tools