REVIEW 14 cited by
News Summarization and Evaluation in the Era of GPT-3
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
The recent success of prompting large language models like GPT-3 has led to a paradigm shift in NLP research. In this paper, we study its impact on text summarization, focusing on the classic benchmark domain of news summarization. First, we investigate how GPT-3 compares against fine-tuned models trained on large summarization datasets. We show that not only do humans overwhelmingly prefer GPT-3 summaries, prompted using only a task description, but these also do not suffer from common dataset-specific issues such as poor factuality. Next, we study what this means for evaluation, particularly the role of gold standard test sets. Our experiments show that both reference-based and reference-free automatic metrics cannot reliably evaluate GPT-3 summaries. Finally, we evaluate models on a setting beyond generic summarization, specifically keyword-based summarization, and show how dominant fine-tuning approaches compare to prompting. To support further research, we release: (a) a corpus of 10K generated summaries from fine-tuned and prompt-based models across 4 standard summarization benchmarks, (b) 1K human preference judgments comparing different systems for generic- and keyword-based summarization.
Forward citations
Cited by 14 Pith papers
-
ExAnte: A Benchmark for Ex-Ante Inference in Large Language Models
Models leak future knowledge despite explicit temporal cutoffs, as quantified by the ExAnte benchmark across four tasks.
-
Am I More Pointwise or Pairwise? Revealing Position Bias in Rubric-Based LLM-as-a-Judge
Score-option position in rubrics causes consistent, model-specific bias in LLM-as-a-judge evaluations, and averaging over balanced rubric permutations partially corrects it.
-
Topic-Guided Reinforcement Learning with LLMs for Enhancing Multi-Document Summarization
A topic-F1 reward measuring alignment between summary and source-document topics, combined with GRPO training, improves multi-document summarization over several baselines.
-
Evaluating the Evaluators: Are readability metrics good measures of readability?
Classic readability formulas correlate weakly with human readability judgments for plain-language summaries (FKGL r=0.16), the best LLM judge reaches r=0.56, and the two evaluator families rank datasets nearly opposite.
-
DiscoSum: Discourse-aware News Summarization
DiscoSum pairs news articles with cross-platform human summaries and shows that beam search guided by a discourse labeler produces summaries that better match a target sentence structure.
-
M$^3$FinMeeting: A Multilingual, Multi-Sector, and Multi-Task Financial Meeting Understanding Evaluation Dataset
M3FinMeeting is a new 600-meeting, trilingual, multi-sector benchmark with three financial meeting understanding tasks, on which current LLMs achieve only moderate judged quality scores.
-
Scaling Multi-Document Event Summarization: Evaluating Compression vs. Full-Text Approaches
Full-context and retrieval-based methods outperform hierarchical and incremental compression for large-scale multi-document summarization, though compression methods show strong intermediate information retention.
-
"Pragmatic Tools or Empowering Friends?" Discovering and Co-Designing Personality-Aligned AI Writing Companions
Writers grouped into four MBTI-based profiles showed divergent preferences for AI writing companion features, demonstrated by two contrasting prototypes in a small proof-of-concept study.
-
Mobile Application Review Summarization using Chain of Density Prompting
Adapting the Chain of Density prompt to define entities as app features yields denser and more readable summaries of mobile app reviews than the original prompt, vanilla prompting, or extractive baselines.
-
Summarization for Generative Relation Extraction in the Microbiome Domain
Using LLM-generated summaries as input improves instruction-tuned generative relation extraction in the low-resource microbiome domain, though fine-tuned BERT models remain more accurate.
-
Feeling Guilty Being a c(ai)borg: Navigating the Tensions Between Guilt and Empowerment in AI Use
Using three authors' year-long autoethnographic reflections, the paper argues that guilt about AI use can be reframed as growth through literacy, transparency, and critical engagement.
-
Principled Content Selection to Generate Diverse and Personalized Multi-Document Summaries
Selecting LLM-extracted key points with a diversity-aware determinantal point process before rewriting improves source coverage in multi-document news summarization.
-
UiS-IAI@LiveRAG: Retrieval-Augmented Information Nugget-Based Generation of Responses
A nugget-based RAG pipeline with query rewriting and cluster-based summarization is applied to the LiveRAG challenge, where few rewrites plus the original query improve recall and larger document cutoffs hit diminishi...
-
A Multi-Model Metric-based Selection Framework for Abstractive Text summarization
Selecting among T5, PEGASUS and LED summaries by averaging ROUGE-L + BLEU + BERTScore yields 88.63 % BERTScore on CNN/DailyMail, beating the individual models and several reported LLMs.
Discussion (0). Continue with ORCID to comment.