Pith. sign in

REVIEW 3 cited by

SummEval: Re-evaluating Summarization Evaluation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2007.12626 v4 pith:ZG6U2HT6 submitted 2020-07-24 cs.CL

classification cs.CL
keywords evaluationsummarizationmetricsautomatichumanmodelssharealong
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The scarcity of comprehensive up-to-date studies on evaluation metrics for text summarization and the lack of consensus regarding evaluation protocols continue to inhibit progress. We address the existing shortcomings of summarization evaluation methods along five dimensions: 1) we re-evaluate 14 automatic evaluation metrics in a comprehensive and consistent fashion using neural summarization model outputs along with expert and crowd-sourced human annotations, 2) we consistently benchmark 23 recent summarization models using the aforementioned automatic evaluation metrics, 3) we assemble the largest collection of summaries generated by models trained on the CNN/DailyMail news dataset and share it in a unified format, 4) we implement and share a toolkit that provides an extensible and unified API for evaluating summarization models across a broad range of automatic metrics, 5) we assemble and share the largest and most diverse, in terms of model types, collection of human judgments of model-generated summaries on the CNN/Daily Mail dataset annotated by both expert judges and crowd-source workers. We hope that this work will help promote a more complete evaluation protocol for text summarization as well as advance research in developing evaluation metrics that better correlate with human judgments.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Revisiting Metric Reliability for Fine-grained Evaluation of Machine Translation and Summarization in Indian Languages

    cs.CL 2025-10 conditional novelty 6.0 of 10

    LLM-as-judge metrics (DeepSeek-V3 most of all) correlate best with human ratings across six Indian languages, though all segment-level correlations are low and many differences lack confidence intervals.

  2. IDEAlign: Comparing Large Language Models to Human Experts in Open-ended Interpretive Annotations

    cs.CL 2025-09 conditional novelty 5.0 of 10

    IDEAlign uses a pick-the-odd-one-out triplet task to measure idea-level similarity between LLMs and expert human annotations, and shows LLM judges using this protocol outperform lexical and vector-based baselines.

  3. LLMs as Architects and Critics for Multi-Source Opinion Summarization

    cs.CL 2025-07 conditional novelty 5.0 of 10

    A new benchmark and prompt framework for generating and automatically evaluating product summaries that blend customer reviews with product metadata, with the best evaluator reaching 0.74 average Spearman correlation ...

Pith tools