Pith. sign in

REVIEW 8 cited by

Summarization is (Almost) Dead

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2309.09558 v1 pith:3LIV44L2 submitted 2023-09-18 cs.CL

classification cs.CL
keywords summariesllmssummarizationdatasetsevaluationhumanllm-generatedmodels
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

How well can large language models (LLMs) generate summaries? We develop new datasets and conduct human evaluation experiments to evaluate the zero-shot generation capability of LLMs across five distinct summarization tasks. Our findings indicate a clear preference among human evaluators for LLM-generated summaries over human-written summaries and summaries generated by fine-tuned models. Specifically, LLM-generated summaries exhibit better factual consistency and fewer instances of extrinsic hallucinations. Due to the satisfactory performance of LLMs in summarization tasks (even surpassing the benchmark of reference summaries), we believe that most conventional works in the field of text summarization are no longer necessary in the era of LLMs. However, we recognize that there are still some directions worth exploring, such as the creation of novel datasets with higher quality and more reliable evaluation methods.

Discussion (0). Sign in to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. VA-Blueprint: Uncovering Building Blocks for Visual Analytics System Design

    cs.HC 2025-08 conditional novelty 6.0 of 10

    A semi-automated methodology and public knowledge base catalog the building blocks of 101 urban visual analytics systems, using GPT-4 for extraction with human-in-the-loop correction and expert validation.

  2. From Alerts to Intelligence: A Novel LLM-Aided Framework for Host-based Intrusion Detection

    cs.CR 2025-07 conditional novelty 6.0 of 10

    SHIELD, an LLM-aided pipeline combining a masked autoencoder, deterministic data augmentation, and multi-level prompting, detects host-based attacks with high precision on three public datasets.

  3. Fair Document Valuation in LLM Summaries via Shapley Values

    cs.CL 2025-05 reject novelty 6.0 of 10

    Cluster Shapley groups semantically similar documents via embeddings and computes cluster-level Shapley values, claiming better efficiency-accuracy trade-offs than Monte Carlo and Kernel SHAP on Amazon review summarization.

  4. An Empirical Study of Many-Shot In-Context Learning for Machine Translation of Low-Resource Languages

    cs.CL 2026-04 unverdicted novelty 5.0 of 10

    BM25-retrieved many-shot examples match much larger random sets for translating English into ten truly low-resource languages, and ICL still helps after fine-tuning.

  5. Rethinking Hate Speech Detection on Social Media: Can LLMs Replace Traditional Models?

    cs.CL 2025-06 conditional novelty 5.0 of 10

    On three hate speech datasets, including a new code-mixed IndoHateMix benchmark, fine-tuned LLMs such as LLaMA-3.1 beat multilingual BERT models, with the largest gains on code-mixed Indian text.

  6. A Framework for Generating Conversational Recommendation Datasets from Behavioral Interactions

    cs.IR 2025-06 reject novelty 5.0 of 10

    ConvRecStudio generates roughly 38K synthetic multi-turn recommendation dialogs across three domains from historical interactions, and a cross-attention transformer fusing history with dialog beats dialog-only and his...

  7. Summarization for Generative Relation Extraction in the Microbiome Domain

    cs.CL 2025-06 conditional novelty 5.0 of 10

    Using LLM-generated summaries as input improves instruction-tuned generative relation extraction in the low-resource microbiome domain, though fine-tuned BERT models remain more accurate.

  8. (Towards) Scalable Reliable Automated Evaluation with Large Language Models

    cs.CL 2026-07 conditional novelty 4.0 of 10

    Multi-LLM pairwise Elo ranking with adjustable consensus thresholds produces rankings of competency profiles that average Spearman ρ≈0.83 with expert judgments.

Pith tools