REVIEW 8 cited by
Summarization is (Almost) Dead
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
How well can large language models (LLMs) generate summaries? We develop new datasets and conduct human evaluation experiments to evaluate the zero-shot generation capability of LLMs across five distinct summarization tasks. Our findings indicate a clear preference among human evaluators for LLM-generated summaries over human-written summaries and summaries generated by fine-tuned models. Specifically, LLM-generated summaries exhibit better factual consistency and fewer instances of extrinsic hallucinations. Due to the satisfactory performance of LLMs in summarization tasks (even surpassing the benchmark of reference summaries), we believe that most conventional works in the field of text summarization are no longer necessary in the era of LLMs. However, we recognize that there are still some directions worth exploring, such as the creation of novel datasets with higher quality and more reliable evaluation methods.
Forward citations
Cited by 8 Pith papers
-
VA-Blueprint: Uncovering Building Blocks for Visual Analytics System Design
A semi-automated methodology and public knowledge base catalog the building blocks of 101 urban visual analytics systems, using GPT-4 for extraction with human-in-the-loop correction and expert validation.
-
From Alerts to Intelligence: A Novel LLM-Aided Framework for Host-based Intrusion Detection
SHIELD, an LLM-aided pipeline combining a masked autoencoder, deterministic data augmentation, and multi-level prompting, detects host-based attacks with high precision on three public datasets.
-
Fair Document Valuation in LLM Summaries via Shapley Values
Cluster Shapley groups semantically similar documents via embeddings and computes cluster-level Shapley values, claiming better efficiency-accuracy trade-offs than Monte Carlo and Kernel SHAP on Amazon review summarization.
-
An Empirical Study of Many-Shot In-Context Learning for Machine Translation of Low-Resource Languages
BM25-retrieved many-shot examples match much larger random sets for translating English into ten truly low-resource languages, and ICL still helps after fine-tuning.
-
Rethinking Hate Speech Detection on Social Media: Can LLMs Replace Traditional Models?
On three hate speech datasets, including a new code-mixed IndoHateMix benchmark, fine-tuned LLMs such as LLaMA-3.1 beat multilingual BERT models, with the largest gains on code-mixed Indian text.
-
A Framework for Generating Conversational Recommendation Datasets from Behavioral Interactions
ConvRecStudio generates roughly 38K synthetic multi-turn recommendation dialogs across three domains from historical interactions, and a cross-attention transformer fusing history with dialog beats dialog-only and his...
-
Summarization for Generative Relation Extraction in the Microbiome Domain
Using LLM-generated summaries as input improves instruction-tuned generative relation extraction in the low-resource microbiome domain, though fine-tuned BERT models remain more accurate.
-
(Towards) Scalable Reliable Automated Evaluation with Large Language Models
Multi-LLM pairwise Elo ranking with adjustable consensus thresholds produces rankings of competency profiles that average Spearman ρ≈0.83 with expert judgments.
Discussion (0). Sign in to comment.