Pith. sign in

REVIEW 2 cited by

A Comparative Study of Quality Evaluation Methods for Text Summarization

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.00747 v1 pith:YN4EIJBW submitted 2024-06-30 cs.CL cs.AI

classification cs.CLcs.AI
keywords evaluationsummarizationtextautomaticevaluatinghumanmetricscomparative
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Evaluating text summarization has been a challenging task in natural language processing (NLP). Automatic metrics which heavily rely on reference summaries are not suitable in many situations, while human evaluation is time-consuming and labor-intensive. To bridge this gap, this paper proposes a novel method based on large language models (LLMs) for evaluating text summarization. We also conducts a comparative study on eight automatic metrics, human evaluation, and our proposed LLM-based method. Seven different types of state-of-the-art (SOTA) summarization models were evaluated. We perform extensive experiments and analysis on datasets with patent documents. Our results show that LLMs evaluation aligns closely with human evaluation, while widely-used automatic metrics such as ROUGE-2, BERTScore, and SummaC do not and also lack consistency. Based on the empirical comparison, we propose a LLM-powered framework for automatically evaluating and improving text summarization, which is beneficial and could attract wide attention among the community.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. GigaChat Audio: Time-aware Large Audio Language Model

    eess.AS 2026-07 conditional novelty 6.0 of 10

    Interleaving periodic time markers with continuous audio tokens, plus duration-mixture synthetic training, yields stable temporal grounding for an audio LLM on inputs up to 120 minutes.

  2. MIDAS: Multi-LLM Iterative Data-Adaptive Summarization

    cs.CL 2026-08 conditional novelty 5.0 of 10

    A multi-LLM prompt optimizer that induces domain-specific formatting policies from reference summaries outperforms generic critique-driven prompt optimizers on enterprise ticket summarization.

Pith tools