REVIEW 5 cited by
SAMSum Corpus: A Human-annotated Dialogue Dataset for Abstractive Summarization
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
This paper introduces the SAMSum Corpus, a new dataset with abstractive dialogue summaries. We investigate the challenges it poses for automated summarization by testing several models and comparing their results with those obtained on a corpus of news articles. We show that model-generated summaries of dialogues achieve higher ROUGE scores than the model-generated summaries of news -- in contrast with human evaluators' judgement. This suggests that a challenging task of abstractive dialogue summarization requires dedicated models and non-standard quality measures. To our knowledge, our study is the first attempt to introduce a high-quality chat-dialogues corpus, manually annotated with abstractive summarizations, which can be used by the research community for further studies.
Forward citations
Cited by 5 Pith papers
-
SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer
Selective three-path KV cache transfer (profile-proactive, parallel on-demand, speculative) cuts time-to-second-token up to 4.3× vs full transfer on bandwidth-limited cloud GPUs with near-baseline accuracy and decode speed.
-
CompilerKV: Risk-Adaptive KV Compression via Offline Experience Compilation
Offline-learned head-reliability and risk-threshold tables make prefill-only KV compression recover about 97.7% of uncompressed LongBench accuracy at a 512-token-per-layer memory budget.
-
OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning
A shared-backbone transformer with pairwise modality training reports top results across 25 datasets spanning 12 modalities.
-
SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning
SafeTuneBed is a plugin-based benchmark and toolkit that standardizes the data, defenses, and metrics used to evaluate how well LLM fine-tuning preserves safety alignment.
-
Zero-Shot Adaptation of Parameter-Efficient Fine-Tuning in Diffusion Models
ProLoRA transfers pre-trained LoRA, DoRA, and FouRA adapters between diffusion models in a single closed-form projection step, without retraining on the target model.
Discussion (0). Continue with ORCID to comment.