Pith. sign in

REVIEW 14 cited by

News Summarization and Evaluation in the Era of GPT-3

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2209.12356 v2 pith:SM5S6LWJ submitted 2022-09-26 cs.CL

classification cs.CL
keywords summarizationgpt-3modelssummariesevaluateevaluationfine-tunedkeyword-based
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The recent success of prompting large language models like GPT-3 has led to a paradigm shift in NLP research. In this paper, we study its impact on text summarization, focusing on the classic benchmark domain of news summarization. First, we investigate how GPT-3 compares against fine-tuned models trained on large summarization datasets. We show that not only do humans overwhelmingly prefer GPT-3 summaries, prompted using only a task description, but these also do not suffer from common dataset-specific issues such as poor factuality. Next, we study what this means for evaluation, particularly the role of gold standard test sets. Our experiments show that both reference-based and reference-free automatic metrics cannot reliably evaluate GPT-3 summaries. Finally, we evaluate models on a setting beyond generic summarization, specifically keyword-based summarization, and show how dominant fine-tuning approaches compare to prompting. To support further research, we release: (a) a corpus of 10K generated summaries from fine-tuned and prompt-based models across 4 standard summarization benchmarks, (b) 1K human preference judgments comparing different systems for generic- and keyword-based summarization.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 14 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 182 citations worldwide. Full citation record

  1. ExAnte: A Benchmark for Ex-Ante Inference in Large Language Models

    cs.LG 2025-05 conditional novelty 7.0 of 10

    Models leak future knowledge despite explicit temporal cutoffs, as quantified by the ExAnte benchmark across four tasks.

  2. Am I More Pointwise or Pairwise? Revealing Position Bias in Rubric-Based LLM-as-a-Judge

    cs.CL 2026-02 conditional novelty 6.0 of 10

    Score-option position in rubrics causes consistent, model-specific bias in LLM-as-a-judge evaluations, and averaging over balanced rubric permutations partially corrects it.

  3. Topic-Guided Reinforcement Learning with LLMs for Enhancing Multi-Document Summarization

    cs.CL 2025-09 conditional novelty 6.0 of 10

    A topic-F1 reward measuring alignment between summary and source-document topics, combined with GRPO training, improves multi-document summarization over several baselines.

  4. Evaluating the Evaluators: Are readability metrics good measures of readability?

    cs.CL 2025-08 conditional novelty 6.0 of 10

    Classic readability formulas correlate weakly with human readability judgments for plain-language summaries (FKGL r=0.16), the best LLM judge reaches r=0.56, and the two evaluator families rank datasets nearly opposite.

  5. DiscoSum: Discourse-aware News Summarization

    cs.CL 2025-06 conditional novelty 6.0 of 10

    DiscoSum pairs news articles with cross-platform human summaries and shows that beam search guided by a discourse labeler produces summaries that better match a target sentence structure.

  6. M$^3$FinMeeting: A Multilingual, Multi-Sector, and Multi-Task Financial Meeting Understanding Evaluation Dataset

    cs.CL 2025-06 conditional novelty 6.0 of 10

    M3FinMeeting is a new 600-meeting, trilingual, multi-sector benchmark with three financial meeting understanding tasks, on which current LLMs achieve only moderate judged quality scores.

  7. Scaling Multi-Document Event Summarization: Evaluating Compression vs. Full-Text Approaches

    cs.CL 2025-02 conditional novelty 6.0 of 10

    Full-context and retrieval-based methods outperform hierarchical and incremental compression for large-scale multi-document summarization, though compression methods show strong intermediate information retention.

  8. "Pragmatic Tools or Empowering Friends?" Discovering and Co-Designing Personality-Aligned AI Writing Companions

    cs.HC 2025-09 conditional novelty 5.0 of 10

    Writers grouped into four MBTI-based profiles showed divergent preferences for AI writing companion features, demonstrated by two contrasting prototypes in a small proof-of-concept study.

  9. Mobile Application Review Summarization using Chain of Density Prompting

    cs.SE 2025-06 conditional novelty 5.0 of 10

    Adapting the Chain of Density prompt to define entities as app features yields denser and more readable summaries of mobile app reviews than the original prompt, vanilla prompting, or extractive baselines.

  10. Summarization for Generative Relation Extraction in the Microbiome Domain

    cs.CL 2025-06 conditional novelty 5.0 of 10

    Using LLM-generated summaries as input improves instruction-tuned generative relation extraction in the low-resource microbiome domain, though fine-tuned BERT models remain more accurate.

  11. Feeling Guilty Being a c(ai)borg: Navigating the Tensions Between Guilt and Empowerment in AI Use

    cs.CY 2025-05 conditional novelty 5.0 of 10

    Using three authors' year-long autoethnographic reflections, the paper argues that guilt about AI use can be reframed as growth through literacy, transparency, and critical engagement.

  12. Principled Content Selection to Generate Diverse and Personalized Multi-Document Summaries

    cs.CL 2025-05 conditional novelty 5.0 of 10

    Selecting LLM-extracted key points with a diversity-aware determinantal point process before rewriting improves source coverage in multi-document news summarization.

  13. UiS-IAI@LiveRAG: Retrieval-Augmented Information Nugget-Based Generation of Responses

    cs.IR 2025-06 conditional novelty 4.0 of 10

    A nugget-based RAG pipeline with query rewriting and cluster-based summarization is applied to the LiveRAG challenge, where few rewrites plus the original query improve recall and larger document cutoffs hit diminishi...

  14. A Multi-Model Metric-based Selection Framework for Abstractive Text summarization

    cs.CL 2026-06 unverdicted novelty 3.0 of 10

    Selecting among T5, PEGASUS and LED summaries by averaging ROUGE-L + BLEU + BERTScore yields 88.63 % BERTScore on CNN/DailyMail, beating the individual models and several reported LLMs.

Pith tools