REVIEW 3 major objections 6 minor 120 references
Survey on Abstractive Text Summarization: Dataset, Models, and Metrics
T0 review · 3 major / 6 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read Small test: AI summary fact errors down from 30 percent
desk verdict Useful but uneven survey; the empirical claim of reduced factual inconsistency rests on a metric mismatch. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing machinery is the transformer encoder-decoder architecture with pretraining objectives such as masked language modeling, gap-sentence generation, and causal language modeling, plus the sparse and local-global attention variants that extend input length to about 16,000 tokens. On top of this, the experiment uses a GPT-2-based fact-checking classifier that returns the probability that a generated summary, treated as a claim, is entailed by an evidence text. For long and multi-document cases, the evidence is substituted with the reference summary when the source is too long, and that substitution is what allows the conclusion about reduced factual inconsistency to be drawn at all.
What would settle it
Run the same fact-checking protocol on a much larger sample and, for long and multi-document inputs, compare FactCheck scores when the evidence is the original source document versus the reference summary; if source-evidence scores fall substantially below reference-evidence scores, the claimed reduction in factual inconsistency is largely an artifact of the evidence substitution.
Extended reading notes
Core claim
The paper claims that, on its test cases, factual inconsistency in abstractive summarization has fallen to a small fraction of the roughly 30 percent rate reported by earlier studies, with average FactCheck scores of 0.93 to 1.00 for short-document summaries, 0.795 to 0.986 for long-document summaries, and 0.835 to 0.93 for multi-document summaries. It also claims that model size, knowledge distillation, and finetuning on multiple domains tend to improve scores, and that pretrained transformer models such as BART, PEGASUS, Longformer, LongT5, PRIMERA, CENTRUM, and REFLECT can produce fluent summaries, with repetition largely controlled by no-repeat n-gram settings. The survey's broader claim is that the field can now be mapped by task type, dataset domain, and evaluation dimension, even though factual faithfulness remains an open challenge for long and multi-document inputs.
Load-bearing premise
The conclusion that factual inconsistency has dropped depends on assuming the reference summary is factually consistent and can serve as evidence for the source, and on a very small number of test documents.
Editorial extensions
If this is right
- If the reported FactCheck scores hold up, modern pretrained summarizers may have reduced the hallucination problem to a small fraction of its earlier level, at least for short news-like documents.
- Finetuning on multiple domains and using distilled versions of large models appear to be reliable routes to better summarization scores, not just larger parameter counts.
- Factual consistency of long and multi-document summaries cannot yet be measured directly, so better evidence-based checkers would be needed before deploying these models.
- ROUGE-style overlap metrics remain the default evaluation, with semantic and factual metrics playing a supporting role, so model rankings could shift if factuality were weighted more heavily.
- The taxonomy of tasks by input length and document count gives practitioners a way to choose a model family appropriate to their use case.
Reading between the lines
- The 'reduced significantly' conclusion is fragile because it rests on only 7, 2, and 10 test documents, and a larger evaluation could move the averages considerably.
- For long and multi-document outputs, using the reference summary as evidence likely inflates FactCheck scores, because the reference is already a clean paraphrase; comparing source-evidence scores would quantify this inflation.
- A testable extension would be to run the same protocol across a broad multi-domain benchmark to see whether the apparent drop in factual inconsistency is domain-dependent, especially outside news.
- If the evidence-substitution gap turns out to be large, the practical takeaway changes from 'hallucination is mostly solved' to 'we still lack a reliable way to measure faithfulness for long inputs.'
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper is a survey of abstractive text summarization, covering task definitions, extractive/abstractive/hybrid approaches, transformer-based models for short, long, and multi-document inputs, datasets, and automatic evaluation metrics. It also reports small-scale experiments on publicly available fine-tuned checkpoints (BART, PEGASUS, Longformer, LongT5, PRIMERA, CENTRUM, REFLECT) evaluated with ROUGE, METEOR, CHRF, BertScore, and a FactCheck probability score. The central empirical claim, stated in Sections 6.1, 6.2, and 7, is that factual inconsistency in generated summaries has "reduced significantly" relative to the approximately 30% inconsistency rate reported by references [46-49].
Significance. If the survey's map of the field is accurate, it provides a useful organized introduction to abstractive summarization models, datasets, and metrics, particularly for readers outside the area. A notable strength is that the authors release the code, data, and per-sample results, which supports reproducibility. The empirical sections are best read as illustrative demonstrations of public checkpoints rather than as rigorous comparative evaluations. The claimed large reduction in factual inconsistency would be significant if supported, but it is not supported by the current evidence because the FactCheck score is an uncalibrated average probability that is not commensurable with the 30% inconsistency baseline, and the sample sizes of 7, 2, and 10 documents are far too small to support the word "significantly".
major comments (3)
- [§6.1, §6.2, §7; Tables 2-4] The conclusion that factual inconsistency "has reduced significantly compared to the ~30% factual inconsistency reported by [46-49]" is not supported by the reported FactCheck numbers. The FactCheck column is an average probability assigned by a FEVER-trained GPT-2 NLI model, while a 30% inconsistency rate is a proportion of summaries classified as inconsistent under some scoring scheme. An average probability of 0.93 can coexist with a high proportion of low-probability summaries, so the comparison is meaningful only if the paper specifies and applies a threshold (e.g., probability below 0.5 counts as inconsistent) and reports the resulting rate, along with some check of the model's calibration. Without this, the average FactCheck value in Tables 2, 3, and 4 cannot be converted into an inconsistency percentage, and the central claim in Section 7 is not established.
- [§6] For long and multi-document summaries, the reference summary of the source document is used as evidence for the factuality check, under the stated assumption that the reference summary is factually consistent in entities and entity relations. This assumption is load-bearing for the long- and multi-document claims in Tables 3 and 4. Reference summaries are not guaranteed to be factually complete or consistent, and the paper provides no validation of this assumption. Consequently, the FactCheck scores for long and multi-document outputs measure consistency with the reference summary rather than faithfulness to the source documents, and the strong claim of reduced factual inconsistency cannot be drawn from them.
- [§6.1 and §6.2] The sample sizes are 7 short documents, 2 long documents, and 10 multi-document clusters. The paper reports only means, with no confidence intervals, standard deviations, per-model statistical tests, or per-sample dispersion. Statements such as "M s2 shows 100% factually consistent summaries score" and "the factuality problem ... has reduced significantly" are therefore not statistically supported; the study is too small to detect meaningful differences among models or to compare reliably with historical rates.
minor comments (6)
- [Abstract and author affiliation] The affiliation for Flavio Bertini is given as "University of Parma" in the affiliation block but the contact line says "flavio.bertini@unipr.it"; the reader report uses "Favio" as a first name. Please verify the spelling and institutional attribution.
- [Keyword line] The keyword line contains the typo "Estractive" (should be "Extractive").
- [Section 2.3, paragraph on abstractive summarization] The sentence "but the factuality of the summries produce is still a challenge" contains spelling errors and should be rewritten, e.g., "but the factuality of the summaries produced is still a challenge."
- [Section 6.2] The text states that "M l2 outperforms other finetuned model" when referring to the multi-document results; the model labels are M m1, M m2, and M m3, so this should read "M m2."
- [Sections 3.1 and 5] The paper uses inconsistent capitalization and spacing for model names (e.g., "P EGASU SLARGE", "Tranformer", "Rouge"). A consistent formatting pass across the text and tables would improve readability.
- [Table 1] The table caption calls the table "The breakdown of the experimented finetuned models," but several listed models are not covered by the later experiments (e.g., Gigaword, Wikihow, Reddit TIFU, BookSum); consider either expanding the experiments or revising the table caption to indicate which models and datasets were actually tested.
Circularity Check
No circularity found: the paper performs no derivation or parameter fitting, and its empirical claim, while methodologically fragile, is not defined in terms of the conclusion it asserts.
full rationale
This is a survey with small illustrative experiments, not a derivation chain. No model is fitted by the authors; they evaluate publicly available Hugging Face checkpoints (BART, PEGASUS, Longformer, LongT5, PRIMERA, CENTRUM, REFLECT) with external metrics (ROUGE, METEOR, CHRF, BertScore) and an external fact-checking model (fractalego/fact-checking, a GPT-2 model trained on FEVER). The paper therefore contains no self-definitional step, no fitted parameter renamed as a prediction, and no uniqueness claim imported from the authors' own prior work. The only self-citation is to the authors' GitHub repository for data and code, which is data-availability boilerplate and is not load-bearing for any stated conclusion. The central empirical claim that factual inconsistency 'has reduced significantly compared to the ~30% factual inconsistency reported by [46-49]' is unsupported because an average continuous FactCheck probability is compared to a proportion of inconsistent summaries without applying a classification threshold, and the FactCheck model's calibration for summarization is unknown. That is a validity or correctness problem, not circularity: the measured score is not defined in terms of the claim, and no input quantity is reused as the output. The paper also explicitly discloses that for long and multi-document cases the reference summary was used as evidence under an assumption of factual consistency; this weakens the faithfulness measurement but does not make the reported score equivalent to the conclusion by construction. Per the stated rules, an unsupported comparison is not a circular step, and the honest finding is therefore no significant circularity.
Assumptions & free parameters
assumptions (4)
- domain assumption The GPT-2-based fact-checking model from [85] provides a valid measure of summary factual consistency.
- ad hoc to paper The reference summary of a document is factually consistent and can serve as evidence for long and multi-document factuality checks.
- ad hoc to paper Seven short, two long, and ten multi-document samples are representative enough to compare model performance.
- domain assumption Automatic metrics (ROUGE, METEOR, CHRF, BERTScore) and the fact-checking score capture the quality dimensions claimed in Section 4.
Cite this review
Pith. "Pith review of Survey on Abstractive Text Summarization: Dataset, Models, and Metrics." pith.science (2026). https://pith.science/paper/N3E2FOF3
@misc{pith2026241217165,
author = {Pith},
title = {Pith review of: Survey on Abstractive Text Summarization: Dataset, Models, and Metrics},
year = {2026},
howpublished = {\url{https://pith.science/paper/N3E2FOF3}},
note = {Machine review of arXiv:2412.17165}
}
read the original abstract
The advancements in deep learning, particularly the introduction of transformers, have been pivotal in enhancing various natural language processing (NLP) tasks. These include text-to-text applications such as machine translation, text classification, and text summarization, as well as data-to-text tasks like response generation and image-to-text tasks such as captioning. Transformer models are distinguished by their attention mechanisms, pretraining on general knowledge, and fine-tuning for downstream tasks. This has led to significant improvements, particularly in abstractive summarization, where sections of a source document are paraphrased to produce summaries that closely resemble human expression. The effectiveness of these models is assessed using diverse metrics, encompassing techniques like semantic overlap and factual correctness. This survey examines the state of the art in text summarization models, with a specific focus on the abstractive summarization approach. It reviews various datasets and evaluation metrics used to measure model performance. Additionally, it includes the results of test cases using abstractive summarization models to underscore the advantages and limitations of contemporary transformer-based models. The source codes and the data are available at https://github.com/gospelnnadi/Text-Summarization-SOTA-Experiment.
Figures
Figures from the paper (1 more)
Reference graph
Works this paper leans on
-
[1]
Text summarization sota experiment
Github repository. Text summarization sota experiment. https://github.com/data-lang/ Text-Summarization-SOTA-Experiment.git , 2023
2023
-
[2]
World report: Citable documents, 2022
ScimagoJR. World report: Citable documents, 2022. https://www.scimagojr.com/worldreport.php, 2022
2022
-
[3]
Report of all scientific articles on covid published from 2015 to 2024
Scopus. Report of all scientific articles on covid published from 2015 to 2024. https://www.scopus.com/, 2024
2015
-
[4]
Congbo Ma, Wei Emma Zhang, Mingyu Guo, Hu Wang, and Quan Z. Sheng. Multi-document summarization via deep learning techniques: A survey. arXiv preprint arXiv:, 2021
2021
-
[5]
H. P. Luhn. The automatic creation of literature abstracts. IBM Journal of Research and Development, 2(2):159– 165, 1958
1958
-
[6]
Bart: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension
Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Ves Stoyanov, and Luke Zettlemoyer. Bart: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension. 2019. 18 Survey on Abstractive Text Summarization: Dataset, Models, and Metrics
2019
-
[7]
Jingqing Zhang, Yao Zhao, Mohammad Saleh, and Peter J. Liu. Pegasus: Pre-training with extracted gap- sentences for abstractive summarization. 2020
2020
-
[8]
Peters, and Arman Cohan
Iz Beltagy, Matthew E. Peters, and Arman Cohan. Longformer: The long-document transformer. arXiv preprint arXiv:, 2020
2020
Show all 120 references
-
[9]
Big bird: Transformers for longer sequences
Manzil Zaheer, Guru Guruganesh, Avinava Dubey, Joshua Ainslie, Chris Alberti, Santiago Ontanon, Philip Pham, Anirudh Ravula, Qifan Wang, Li Yang, and Amr Ahmed. Big bird: Transformers for longer sequences. arXiv preprint arXiv:, 2021
2021
-
[10]
Longt5: Efficient text-to-text transformer for long sequences
Mandy Guo, Joshua Ainslie, David Uthus, Santiago Ontañón, Jianmo Ni, Yun-Hsuan Sung, and Yinfei Yang. Longt5: Efficient text-to-text transformer for long sequences. In Association for Computational Linguistics, 2022
2022
-
[11]
Gomez, Łukasz Kaiser, and Illia Polosukhin
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. arXiv preprint arXiv:, 2017
2017
-
[12]
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. Bert: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Langu...
2019
-
[13]
Primera: Pyramid-based masked sentence pre-training for multi-document summarization
Wen Xiao, Iz Beltagy, Giuseppe Carenini, and Arman Cohan. Primera: Pyramid-based masked sentence pre-training for multi-document summarization. arXiv preprint arXiv:, 2022
2022
-
[14]
Antognini and B
D. Antognini and B. Faltings. Learning to create sentence semantic relation graphs for multi-document summa- rization. In L. Wang, J. C. K. Cheung, G. Carenini, and F. Liu, editors, Proceedings of the 2nd Workshop on New Frontiers in Summarization, pages 32–41, Hong Kong, Chin...
2019
-
[15]
M. T. Nayeem, T. A. Fuad, and Y . Chali. Abstractive unsupervised multi-document summarization using paraphrastic sentence fusion. In E. M. Bender, L. Derczynski, and P. Isabelle, editors, Proceedings of the 27th International Conference on Computational Linguistics, pages 119...
2018
-
[16]
Yasunaga, R
M. Yasunaga, R. Zhang, K. Meelu, A. Pareek, K. Srinivasan, and D. Radev. Graph-based neural multi-document summarization. In R. Levy and L. Specia, editors, Proceedings of the 21st Conference on Computational Natural Language Learning (CoNLL 2017), pages 452–462, Vancouver, Ca...
2017
-
[17]
Song, Y .-S
Y .-Z. Song, Y .-S. Chen, and H.-H. Shuai. Improving multi-document summarization through referenced flexible extraction with credit-awareness. 2022
2022
-
[18]
Efficiently summarizing text and graph encodings of multi-document clusters
Ramakanth Pasunuru, Mengwen Liu, Mohit Bansal, Sujith Ravi, and Markus Dreyer. Efficiently summarizing text and graph encodings of multi-document clusters. In Association for Computational Linguistics, 2021
2021
-
[19]
Learning to extract coherent summary via deep reinforcement learning
Yuxiang Wu and Baotian Hu. Learning to extract coherent summary via deep reinforcement learning. In Thirty-Second AAAI Conference on Artificial Intelligence, 2018
2018
-
[20]
Banditsum: Extractive summarization as a contextual bandit
Yue Dong, Yikang Shen, Eric Crawford, Herke van Hoof, and Jackie Chi Kit Cheung. Banditsum: Extractive summarization as a contextual bandit. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 3739–3748, Brussels, Belgium, 2018. Ass...
2018
-
[21]
Neural document summa- rization by jointly learning to score and select sentences
Qingyu Zhou, Nan Yang, Furu Wei, Shaohan Huang, Ming Zhou, and Tiejun Zhao. Neural document summa- rization by jointly learning to score and select sentences. In ACL 2018, pages 654–663, Melbourne, Australia,
2018
-
[22]
Neural latent extractive document summarization
Xingxing Zhang, Mirella Lapata, Furu Wei, and Ming Zhou. Neural latent extractive document summarization. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 779–784, Brussels, Belgium, 2018. Association for Computational Linguistics
2018
-
[23]
Cohen, and Mirella Lapata
Shashi Narayan, Shay B. Cohen, and Mirella Lapata. Ranking sentences for extractive summarization with reinforcement learning. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Vol...
2018
-
[24]
Williams
Ronald J. Williams. Simple statistical gradient-following algorithms for connectionist reinforcement learning. Machine Learning, 8(3-4):229–256, 1992
1992
-
[25]
Neural extractive text summarization with syntactic compression
Jiacheng Xu and Greg Durrett. Neural extractive text summarization with syntactic compression. In EMNLP- IJCNLP 2019, pages 3292–3303, Hong Kong, China, 2019. Association for Computational Linguistics. 19 Survey on Abstractive Text Summarization: Dataset, Models, and Metrics
2019
-
[26]
Strass: A light and effective method for extractive summarization based on sentence embeddings
Leo Bouscarrat, Antoine Bonnefoy, Thomas Peel, and Cecile Pereira. Strass: A light and effective method for extractive summarization based on sentence embeddings. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics: Student Research Works...
2019
-
[27]
A novel extractive multi-document text summarization system using quantum-inspired genetic algorithm: Mtsqiga
Mohammad Mojrian and Seyedabolghasem Mirroshandel. A novel extractive multi-document text summarization system using quantum-inspired genetic algorithm: Mtsqiga. Expert Systems with Applications, 171:114555, 01 2021
2021
-
[28]
Alguliev, R.M
R.M. Alguliev, R.M. Aliguliyev, and N.R. Isazade. Cdds: Constraint-driven document summarization models. Expert Systems with Applications, 40:458–465, 2013
2013
-
[29]
Alguliyev, R
R. Alguliyev, R. Aliguliyev, and N. Isazade. An unsupervised approach to generating generic summaries of documents. Applied Soft Computing, 34, 2015
2015
-
[30]
Alguliyev, R
R. Alguliyev, R. Aliguliyev, M. Hajirahimova, and C. Mehdiyev. Mcmr: Maximum coverage and minimum redundant text summarization model. Expert Systems with Applications, 38:14514–14522, Nov 2011
2011
-
[31]
Y . Liu, X. Wang, J. Zhang, and H. Xu. Personalized pagerank based multi-document summarization. InIEEE International Workshop on Semantic Computing and Systems, pages 169–173, 2008
2008
-
[32]
Alguliyev, R
R. Alguliyev, R. Aliguliyev, and C. Mehdiyev. Sentence selection for generic document summarization using an adaptive differential evolution algorithm. Swarm and Evolutionary Computation, 1:213–222, Dec 2011
2011
-
[33]
Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer
Matthew E. Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer. Deep contextualized word representations. arXiv preprint arXiv:, 2018
2018
-
[34]
Abstractive text summarization using sequence-to-sequence rnns and beyond
Ramesh Nallapati, Bowen Zhou, Caglar Gulcehre, and Bing Xiang. Abstractive text summarization using sequence-to-sequence rnns and beyond. In Proceedings of SIGNLL Conference on Computational Natural Language Learning (CoNLL), 2016
2016
-
[35]
Newsroom: A dataset of 1.3 million summaries with diverse extractive strategies
Max Grusky, Mor Naaman, and Yoav Artzi. Newsroom: A dataset of 1.3 million summaries with diverse extractive strategies. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, volume 1 ...
2018
-
[36]
Narayan, S
S. Narayan, S. B. Cohen, and M. Lapata. Don’t give me the details, just the summary! topic-aware convolutional neural networks for extreme summarization. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 1797–1807, Brussels, Belgi...
2018
-
[37]
An entity-driven framework for abstractive summarization
Eva Sharma, Luyang Huang, Zhe Hu, and Lu Wang. An entity-driven framework for abstractive summarization. In Association for Computational Linguistics, 2019
2019
-
[38]
Hierarchical transformers for multi-document summarization
Yang Liu and Mirella Lapata. Hierarchical transformers for multi-document summarization. In Association for Computational Linguistics (ACL 2019), pages 5070–5081, Florence, Italy, 2019
2019
-
[39]
Text summarization with pretrained encoders
Yang Liu and Mirella Lapata. Text summarization with pretrained encoders. 2019
2019
-
[40]
Liu, and Christopher D
Abigail See, Peter J. Liu, and Christopher D. Manning. Get to the point: Summarization with pointer-generator networks. 2017
2017
-
[41]
Cohan, F
A. Cohan, F. Dernoncourt, D. S. Kim, T. Bui, S. Kim, W. Chang, and N. Goharian. A discourse-aware attention model for abstractive summarization of long documents. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguisti...
2018
-
[42]
Soft layer-specific multi-task summarization with entailment and question generation
Han Guo, Ramakanth Pasunuru, and Mohit Bansal. Soft layer-specific multi-task summarization with entailment and question generation. In Association for Computational Linguistics, 2018
2018
-
[43]
Improving abstraction in text summarization
Wojciech Kry´sci´nski, Romain Paulus, Caiming Xiong, and Richard Socher. Improving abstraction in text summarization. In Association for Computational Linguistics, 2018
2018
-
[44]
Raffel, N
C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y . Zhou, W. Li, and P. J. Liu. Exploring the limits of transfer learning with a unified text-to-text transformer. 2023, 2023
2023
-
[45]
Distillation knowledge applied on pegasus for summarization, 2019–2020
Lorenzo Niccolai and Andrea Asperti. Distillation knowledge applied on pegasus for summarization, 2019–2020
2019
-
[46]
Faithful to the original: Fact aware neural abstractive summarization
Ziqiang Cao, Furu Wei, Wenjie Li, and Sujian Li. Faithful to the original: Fact aware neural abstractive summarization. In Proceedings of the Thirty Second AAAI Conference on Artificial Intelligence (AAAI-18), pages 4784–4791, New Orleans, Louisiana, USA, 2018. 20 Survey on Ab...
2018
-
[47]
Liu, and Mohammad Saleh
Ben Goodrich, Vinay Rao, Peter J. Liu, and Mohammad Saleh. Assessing the factual accuracy of generated text. In Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining (KDD 2019), pages 166–175, Anchorage, AK, USA, 2019
2019
-
[48]
Tobias Falke, Leonardo F. R. Ribeiro, Prasetya Ajie Utama, Ido Dagan, and Iryna Gurevych. Ranking generated summaries by correctness: An interesting but challenging application for natural language inference. In Proceedings of the 57th Conference of the Association for Computa...
2019
-
[49]
Neural text summarization: A critical evaluation
Wojciech Kry´sci´nski, Nitish Shirish Keskar, Bryan McCann, Caiming Xiong, and Richard Socher. Neural text summarization: A critical evaluation. Salesforce Research, 2019
2019
-
[50]
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber. Long short-term memory. Neural Computation, 9(8):1735–1780, 1997
1997
-
[51]
Review summarization with pointer gen and bert, 2020
Matthew Martin and Marjolein Pawlus. Review summarization with pointer gen and bert, 2020
2020
-
[52]
A deep reinforced model for abstractive summarization
Romain Paulus, Caiming Xiong, and Richard Socher. A deep reinforced model for abstractive summarization. arXiv preprint arXiv:1705.04304, 2017
2017 arXiv
-
[53]
Multi-reward reinforced summarization with saliency and entailment
Ramakanth Pasunuru and Mohit Bansal. Multi-reward reinforced summarization with saliency and entailment. In Association for Computational Linguistics, 2018
2018
-
[54]
Closed-book training to improve summarization encoder memory
Yichen Jiang and Mohit Bansal. Closed-book training to improve summarization encoder memory. InAssociation for Computational Linguistics, 2018
2018
-
[55]
Ziegler, Nisan Stiennon, Jeffrey Wu, Tom B
Daniel M. Ziegler, Nisan Stiennon, Jeffrey Wu, Tom B. Brown, Alec Radford, Dario Amodei, Paul Christiano, and Geoffrey Irving. Fine-tuning language models from human preferences. arXiv preprint arXiv:, 2020
2020
-
[56]
Unified language model pre-training for natural language understanding and generation
Li Dong, Nan Yang, Wenhui Wang, Furu Wei, Xiaodong Liu, Yu Wang, Jianfeng Gao, Ming Zhou, and Hsiao- Wuen Hon. Unified language model pre-training for natural language understanding and generation. arXiv preprint arXiv:, 2019
2019
-
[57]
Deep communicating agents for abstractive summarization
Asli Celikyilmaz, Antoine Bosselut, Xiaodong He, and Yejin Choi. Deep communicating agents for abstractive summarization. In Association for Computational Linguistics, 2018
2018
-
[58]
W. Li, X. Xiao, J. Liu, H. Wu, H. Wang, and J. Du. Leveraging graph to improve abstractive multi-document summarization. In D. Jurafsky, J. Chai, N. Schluter, and J. Tetreault, editors, Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pa...
2020
-
[59]
Fast abstractive summarization with reinforce-selected sentence rewriting
Yen-Chun Chen and Mohit Bansal. Fast abstractive summarization with reinforce-selected sentence rewriting. In Association for Computational Linguistics, 2018
2018
-
[60]
Sebastian Gehrmann, Yuntian Deng, and Alexander M. Rush. Bottom-up abstractive summarization. 2018
2018
-
[61]
A unified model for extractive and abstractive summarization using inconsistency loss
Wan-Ting Hsu, Chieh-Kai Lin, Ming-Ying Lee, Kerui Min, Jing Tang, and Min Sun. A unified model for extractive and abstractive summarization using inconsistency loss. arXiv preprint arXiv:, 2018
2018
-
[62]
Improving multi-document summarization through referenced flexible extraction with credit-awareness
Yun-Zhu Song, Yi-Syuan Chen, and Hong-Han Shuai. Improving multi-document summarization through referenced flexible extraction with credit-awareness. arXiv preprint arXiv:, 2022
2022
-
[63]
Noah A. Smith. Contextual word representations: Putting words into computers, 2020
2020
-
[64]
Efficient estimation of word representations in vector space
Tomas Mikolov, Kai Chen, Greg Corrado, and Jeffrey Dean. Efficient estimation of word representations in vector space. arXiv preprint arXiv:, 2013
2013
-
[65]
Jeffrey Pennington, Richard Socher, and Christopher D. Manning. Glove: Global vectors for word representation. In Association for Computational Linguistics, 2014
2014
-
[66]
An introduction to convolutional neural networks
Keiron O’Shea and Ryan Nash. An introduction to convolutional neural networks. ArXiv e-prints, 11 2015
2015
-
[67]
Mike Schuster and Kuldip K. Paliwal. Bidirectional recurrent neural networks. IEEE Transactions on Signal Processing, 45(11):2673–2681, 1997
1997
-
[68]
Huggingface
HuggingFace. Huggingface. https://huggingface.co. Accessed: 2024-12-20
2024
-
[69]
Survey of the state of the art in natural language generation: Core tasks, applications and evaluation
Albert Gatt and Emiel Krahmer. Survey of the state of the art in natural language generation: Core tasks, applications and evaluation. arXiv preprint arXiv:1703.09902v4, 2018
2018 arXiv
-
[70]
Rouge: A package for automatic evaluation of summaries, 2004
Chin-Yew Lin. Rouge: A package for automatic evaluation of summaries, 2004
2004
-
[71]
Learning to score system summaries for better content selection evaluation
Maxime Peyrard, Teresa Botschen, and Iryna Gurevych. Learning to score system summaries for better content selection evaluation. In Association for Computational Linguistics, 2017. 21 Survey on Abstractive Text Summarization: Dataset, Models, and Metrics
2017
-
[72]
Weinberger, and Yoav Artzi
Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger, and Yoav Artzi. Bertscore: Evaluating text generation with bert. arXiv preprint arXiv:1904.09675v3, 2020
1904 arXiv
-
[73]
Meyer, and Steffen Eger
Wei Zhao, Maxime Peyrard, Fei Liu, Yang Gao, Christian M. Meyer, and Steffen Eger. Moverscore: Text genera- tion evaluating with contextualized embeddings and earth mover distance. In Association for Computational Linguistics, 2019
2019
-
[74]
Speeding up word mover’s distance and its variants via properties of distances between embeddings, 12 2019
Matheus Werner and Eduardo Laber. Speeding up word mover’s distance and its variants via properties of distances between embeddings, 12 2019
2019
-
[75]
Elizabeth Clark, Asli Celikyilmaz, and Noah A. Smith. Sentence mover’s similarity: Automatic evaluation for multi-sentence texts. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 2748–2760, Florence, Italy, 2019. Association for...
2019
-
[76]
Answers unite! unsupervised metrics for reinforced summarization models
Thomas Scialom, Sylvain Lamprier, Benjamin Piwowarski, and Jacopo Staiano. Answers unite! unsupervised metrics for reinforced summarization models. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Confere...
2019
-
[77]
Vasilyev, Vedant Dharnidharka, and John Bohannon
Oleg V . Vasilyev, Vedant Dharnidharka, and John Bohannon. Fill in the blanc: human-free quality estimation of document summaries. CoRR, abs/2002.09836, 2020
2002 arXiv
-
[78]
Supert: towards new frontiers in unsupervised evaluation metrics for multi document summarization
Yang Gao, Wei Zhao, and Steffen Eger. Supert: towards new frontiers in unsupervised evaluation metrics for multi document summarization. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 1347–1354, Online, 2020. Association for C...
2020
-
[79]
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. Bleu: a method for automatic evaluation of machine translation. In Association for Computational Linguistics, 2002
2002
-
[80]
Chrf: Character n-gram f-score for automatic mt evaluation
Maja Popovi´c. Chrf: Character n-gram f-score for automatic mt evaluation. In Association for Computational Linguistics, 2015
2015
-
[81]
Meteor: An automatic metric for mt evaluation with high levels of correlation with human judgments
Alon Lavie and Abhaya Agarwal. Meteor: An automatic metric for mt evaluation with high levels of correlation with human judgments. In Proceedings of the Second Workshop on Statistical Machine Translation , pages 228–231, Prague, Czech Republic, 2007. Association for Computatio...
2007
-
[82]
Lawrence Zitnick, and Devi Parikh
Ramakrishna Vedantam, C. Lawrence Zitnick, and Devi Parikh. Cider: Consensus-based image description evaluation. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 4566–4575, 2015
2015
-
[83]
Evaluating the factual consistency of abstractive text summarization
Wojciech Kry´sci´nski, Bryan McCann, Caiming Xiong, and Richard Socher. Evaluating the factual consistency of abstractive text summarization. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 9332–9346, Online, 2020. Assoc...
2020
-
[84]
Entity-level factual consistency of abstractive text summarization
Feng Nan, Ramesh Nallapati, Zhiguo Wang, Cicero Nogueira dos Santos, Henghui Zhu, Dejiao Zhang, Kathleen McKeown, and Bing Xiang. Entity-level factual consistency of abstractive text summarization. In Proceedings of the 16th Conference of the European Chapter of the Associatio...
2021
-
[85]
Fact-checking
Fractalego. Fact-checking. https://huggingface.co/fractalego/fact-checking, 2021
2021
-
[86]
Fabbri, Wojciech Kry´sci´nski, Bryan McCann, Caiming Xiong, Richard Socher, and Dragomir Radev
Alexander R. Fabbri, Wojciech Kry´sci´nski, Bryan McCann, Caiming Xiong, Richard Socher, and Dragomir Radev. Summeval: Re-evaluating summarization evaluation. In Association for Computational Linguistics, 2021
2021
-
[87]
Generating representative headlines for news stories
Xiaotao Gu, Yuning Mao, Jiawei Han, Jialu Liu, Hongkun Yu, You Wu, Cong Yu, Daniel Finnie, Jiaqi Zhai, and Nichol Zukoski. Generating representative headlines for news stories. arXiv preprint arXiv:2001.09386v4, 2020
2001 arXiv
-
[88]
K. M. Hermann, T. Kocisky, E. Grefenstette, L. Espeholt, W. Kay, M. Suleyman, and P. Blunsom. Teaching machines to read and comprehend. In Advances in Neural Information Processing Systems, pages 1693–1701, 2015
2015
-
[89]
Newsroom: A dataset of 1.3 million summaries with diverse extractive strategies
Max Grusky, Mor Naaman, and Yoav Artzi. Newsroom: A dataset of 1.3 million summaries with diverse extractive strategies. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, volume 1 ...
2018
-
[90]
A. M. Rush, S. Chopra, and J. Weston. A neural attention model for abstractive sentence summarization. In Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing, pages 379–389, Lisbon, Portugal, 2015. Association for Computational Linguistics
2015
-
[91]
Graff, J
D. Graff, J. Kong, K. Chen, and K. Maeda. English gigaword. Linguistic Data Consortium, 4(1):34, 2003. 22 Survey on Abstractive Text Summarization: Dataset, Models, and Metrics
2003
-
[92]
Sharma, C
E. Sharma, C. Li, and L. Wang. Bigpatent: A large scale dataset for abstractive and coherent summarization. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 2204–2213, Florence, Italy, 2019. Association for Computational Linguistics
2019
-
[93]
Koupaee and W
M. Koupaee and W. Y . Wang. Wikihow: A large scale text summarization dataset. arXiv preprint arXiv:1810.09305, 2018
2018 arXiv
-
[94]
B. Kim, H. Kim, and G. Kim. Abstractive summarization of reddit posts with multi-level memory networks. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, volume 1 (Long and Short P...
2019
-
[95]
Zhang and J
R. Zhang and J. Tetreault. This email could save your life: Introducing the task of email subject line generation. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 446–456, Florence, Italy, 2019. Association for Computational Li...
2019
-
[96]
The enron corpus: A new dataset for email classification research
Bryan Klimt and Yiming Yang. The enron corpus: A new dataset for email classification research. In Jean- François Boulicaut, Floriana Esposito, Fosca Giannotti, and Dino Pedreschi, editors, Machine Learning: ECML 2004, volume 3201 of Lecture Notes in Computer Science, Berlin, ...
2004
-
[97]
Kornilova and V
A. Kornilova and V . Eidelman. Billsum: A corpus for automatic summarization of us legislation. InProceedings of the 2nd Workshop on New Frontiers in Summarization, pages 48–56, Hong Kong, China, 2019. Association for Computational Linguistics
2019
-
[98]
Booksum: A collection of datasets for long-form narrative summarization
Wojciech Kry´sci´nski, Nazneen Rajani, Divyansh Agarwal, Caiming Xiong, and Dragomir Radev. Booksum: A collection of datasets for long-form narrative summarization. arXiv preprint arXiv:2105.08209v1, 2021
2021 arXiv
-
[99]
Efficient attentions for long document summarization
Luyang Huang, Shuyang Cao, Nikolaus Parulian, Heng Ji, and Lu Wang. Efficient attentions for long document summarization. In Kristina Toutanova, Anna Rumshisky, Luke Zettlemoyer, Dilek Hakkani-Tur, Iz Beltagy, Steven Bethard, Ryan Cotterell, Tanmoy Chakraborty, and Yichao Zhou...
2021
-
[100]
Multi-lexsum: Real-world summaries of civil rights lawsuits at multiple granularities, 06 2022
Zejiang Shen, Kyle Lo, Lauren Yu, Nathan Dahlberg, Margo Schlanger, and Doug Downey. Multi-lexsum: Real-world summaries of civil rights lawsuits at multiple granularities, 06 2022
2022
-
[101]
Ms2: Multi-document summarization of medical studies
Jay DeYoung, Iz Beltagy, Madeleine van Zuylen, Bailey Kuehl, and Lucy Lu Wang. Ms2: Multi-document summarization of medical studies. arXiv preprint arXiv:2104.06486v3, 2021
2021 arXiv
-
[102]
Multi-xscience: A large-scale dataset for extreme multi-document summarization of scientific articles
Yao Lu, Yue Dong, and Laurent Charlin. Multi-xscience: A large-scale dataset for extreme multi-document summarization of scientific articles. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), 2020
2020
-
[103]
Fabbri, Irene Li, Tianwei She, Suyi Li, and Dragomir R
Alexander R. Fabbri, Irene Li, Tianwei She, Suyi Li, and Dragomir R. Radev. Multi-news: a large-scale multi-document summarization dataset and abstractive hierarchical model. arXiv preprint arXiv:1906.01749v3, 2019
1906 arXiv
-
[104]
A large-scale multi-document summarization dataset from the wikipedia current events portal
Demian Gholipour Ghalandari, Chris Hokamp, Nghia The Pham, John Glover, and Georgiana Ifrim. A large-scale multi-document summarization dataset from the wikipedia current events portal. In Proceedings of the 2020 Annual Meeting of the Association for Computational Linguistics ...
2020
-
[105]
Neural network-based abstract generation for opinions and arguments
Lu Wang and Wang Ling. Neural network-based abstract generation for opinions and arguments. In Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (HLT-NAACL 2016), pages 47–57, San Dieg...
2016
-
[106]
Liu, Mohammad Saleh, Etienne Pot, Ben Goodrich, Ryan Sepassi, Lukasz Kaiser, and Noam Shazeer
Peter J. Liu, Mohammad Saleh, Etienne Pot, Ben Goodrich, Ryan Sepassi, Lukasz Kaiser, and Noam Shazeer. Generating wikipedia by summarizing long sequences. In Proceedings of the 6th International Conference on Learning Representations (ICLR 2018), Vancouver, Canada, 2018
2018
-
[107]
https://duc.nist.gov/duc2004/, 2004
D u c 2 0 0 4: Documents, tasks, and measures. https://duc.nist.gov/duc2004/, 2004
2004
-
[108]
Summarizing opinions: Aspect extraction meets sentiment prediction and they are both weakly supervised
Stefanos Angelidis and Mirella Lapata. Summarizing opinions: Aspect extraction meets sentiment prediction and they are both weakly supervised. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing (EMNLP), 2018
2018
-
[109]
Nutri-bullets: Summarizing health studies by composing segments
Darsh J Shah, Lili Yu, Tao Lei, and Regina Barzilay. Nutri-bullets: Summarizing health studies by composing segments. arXiv preprint arXiv:2103.11921v1, 2021
2021 arXiv
-
[110]
Aquamuse: Automatically generating datasets for query-based multi-document summarization
Sayali Kulkarni, Sheide Chammas, Wan Zhu, Fei Sha, and Eugene Ie. Aquamuse: Automatically generating datasets for query-based multi-document summarization. arXiv preprint arXiv:2010.12694v1, 2020. 23 Survey on Abstractive Text Summarization: Dataset, Models, and Metrics
2010 arXiv
-
[111]
Gamewikisum: A novel large multi-document summarization dataset
Diego Antognini and Boi Faltings. Gamewikisum: A novel large multi-document summarization dataset. In Proceedings of the 2020 Language Resources and Evaluation Conference (LREC), 2020
2020
-
[112]
Large scale abstractive multi-review summarization (lsars) via aspect alignment
Haojie Pan, Rongqin Yang, Xin Zhou, Rui Wang, Deng Cai, and Xiaozhong Liu. Large scale abstractive multi-review summarization (lsars) via aspect alignment. In Proceedings of the 43rd International ACM SIGIR Conference on Research and Development in Information Retrieval, pages...
2020
-
[113]
Opinosis: A graph based approach to abstractive summa- rization of highly redundant opinions
Kavita Ganesan, ChengXiang Zhai, and Jiawei Han. Opinosis: A graph based approach to abstractive summa- rization of highly redundant opinions. In Proceedings of the 23rd International Conference on Computational Linguistics (COLING 2010), pages 340–348, Beijing, China, 2010
2010
-
[114]
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mari...
2020
-
[115]
GitHub. Github. https://github.com/, 2024
2024
-
[116]
Hallucinations in neural machine translation
Katherine Lee, Orhan Firat, Ashish Agarwal, Clara Fannjiang, and David Sussillo. Hallucinations in neural machine translation. In NeurIPS 2018 Workshop on Interpretability and Robustness for Audio, Speech, and Language, 2018
2018
-
[117]
On faithfulness and factuality in abstractive summarization
Joshua Maynez, Shashi Narayan, Bernd Bohnet, and Ryan McDonald. On faithfulness and factuality in abstractive summarization. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 1906–1919, Online, 2020
1906
-
[118]
Object hallucination in image captioning
Anna Rohrbach, Lisa Anne Hendricks, Kaylee Burns, Trevor Darrell, and Kate Saenko. Object hallucination in image captioning. In Ellen Riloff, David Chiang, Julia Hockenmaier, and Jun’ichi Tsujii, editors, Proceedings of the 2018 Conference on Empirical Methods in Natural Langu...
2018
-
[119]
Reducing quantity hallucinations in abstractive summarization
Zheng Zhao, Shay B Cohen, and Bonnie Webber. Reducing quantity hallucinations in abstractive summarization. arXiv preprint arXiv:2009.13312, 2020. 24
2009 arXiv
-
[2018]
Association for Computational Linguistics
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.