Pith. sign in

REVIEW 4 major objections 6 minor 168 references

A Comprehensive Survey on Legal Summarization: Challenges and Future Directions

T0 review · 4 major / 6 minor · reviewed 2026-08-10 · deepseek-v4-flash

Pith's one-line read A systematic survey of 123 legal-summarization papers maps the field and finds research concentrated in a few English/common-law jurisdictions, dominated by ROUGE, and rarely validated by human or expert evaluation.

desk verdict A genuinely useful survey of transformer-era legal summarization, but the 'comprehensive' claim rests on a retrieval procedure that cannot be audited and the text contains several factual slips that a careful referee would catch. read the letter →

arxiv 2501.17830 v1 pith:LYSSZEOL submitted 2025-01-29 cs.CL

classification cs.CL
keywords legalsummarizationsystematicliteraturereviewNLPtransformereraextractiveabstractiveevaluationmetricsdatasets
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper attempts to establish a systematic, up-to-date map of automatic legal summarization: the methods, datasets, and evaluation metrics used across 123 papers published during the transformer era (2017-2024). Its headline statistics are field-level: over 95 percent of method papers report ROUGE, roughly 20 percent add any human evaluation, about 9 percent involve legal experts, and every dataset has a single reference summary. If the survey's coverage is complete, these numbers describe the legal-summarization field as a whole, not a sample. That matters because legal summaries are high-stakes outputs, and the survey's gap analysis points to concrete fixes: multi-reference and user-personalized datasets, domain-specific embeddings, better handling of long documents, and community-built benchmarks for evaluating the evaluators.

What carries the argument

The machinery is the systematic review protocol: a fixed search query over eight primary and four secondary digital libraries, explicit inclusion/exclusion criteria, two screening stages (titles and abstracts, then full text) with agreement among the first three authors, and a final corpus of 123 papers. On that corpus the survey builds a three-part taxonomy for organising the literature — region (Section 4), strategy (extractive, abstractive, hybrid; Section 5), and method (Section 6) — which generates the field-level statistics. The evaluation discussion is organised by a set of metric families (lexical-overlap, embedding-based, factuality, ranking, classification) and by a five-characteristic human-evaluation framework (content coverage, content representation, efficiency, language quality, summary impact); these are the lenses through which the survey reaches its gap analysis.

What would settle it

Re-run the stated query across an independent set of bibliographic databases and add backward and forward citation chasing, then count peer-reviewed legal-summarization papers from 2017 to 2024 that meet the survey's inclusion criteria but are absent from its 123-paper corpus; if the missed set is large enough to shift the reported figures (over 95 percent ROUGE usage, about 20 percent human evaluation), the survey's field-level conclusions are not robust.

Watch

Extended reading notes

Core claim

The central claim, stated on the paper's own terms, is that the legal-summarization work of the transformer era can be systematically organised and, once organised, reveals a recognizable state of the field. Extractive, abstractive, and hybrid strategies are built mostly on BERT-family encoders and BART-, Pegasus-, and T5-style decoders; datasets such as the US legislative corpus, Indian and UK Supreme Court corpora, Canadian case-law records, and the multilingual EU legislative corpus supply most of the training data; and evaluation is dominated by lexical-overlap metrics, with ROUGE used in more than 95 percent of method papers while human evaluation appears in only about 20 percent, of which about 9 percent use legal experts. The survey also claims that research is concentrated in a few English/common-law jurisdictions, that legal-summarization datasets all carry exactly one reference summary, and that no community benchmark exists for meta-evaluating metrics or human-evaluation protocols in the legal domain. These findings are presented as the basis for a future research agenda: user-specific and multi-reference ground truths, domain-specific embeddings, interpretable handling of lengthy interdependent documents, multimodal and dialogue-aware inputs, and community-built meta-evaluation datasets.

Load-bearing premise

The load-bearing assumption is that the repository search, the keyword query, and the manual screening actually recovered all or nearly all relevant legal-summarization papers from 2017 to 2024; if a substantial body of work was missed or screened inconsistently, the survey's trend statistics describe a convenience sample rather than the whole field.

Editorial extensions

If this is right

  • If the coverage is complete, any claimed progress in legal summarization is currently progress measured mostly by ROUGE, so new results should be re-examined against the metric's known insensitivity to legal content.
  • The regional concentration implies that cross-jurisdiction transfer is not yet reliable; work on non-English and civil-law legal systems remains underserved and cannot be assumed to inherit the existing results.
  • With only about 20 percent of papers using human evaluation, and even fewer involving legal experts, most quality claims in the literature stand on automatic scores alone, making a community benchmark for evaluation metrics a prerequisite for trustworthy comparisons.
  • The single-reference design of all current datasets rules out studying consensus summaries, so building multi-reference legal corpora would open a new line of evaluation research.
  • The survey's limitation analysis identifies the same document types again and again, leaving legal dialogues, meetings, transcripts, and multimodal or low-resource settings as open ground for future work.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Extension: the paper counts metrics per paper; a complementary analysis that weights by dataset would likely sharpen the concentration, since a small number of datasets (the US legislative corpus alone appears in 14 papers) drives much of the observed trend.
  • Extension: the human-evaluation dimensions in the survey could be turned into a testable protocol: have legal experts score a sample of generated summaries on those dimensions and correlate the scores with automatic metrics; the resulting correlation table would tell practitioners which metric to trust when experts are unavailable.
  • Extension: because the search window ends in July 2024, the survey captures only the first wave of instruction-tuned large language models; re-running the same protocol with a later cutoff could test whether ROUGE dominance and the human-evaluation gap are shrinking or persisting.
  • Implication the authors do not draw: a leaderboard built only on the existing English/common-law datasets would reward models that fit those corpora rather than models that generalize, so a cross-jurisdiction held-out benchmark is needed before the field can claim transferable progress.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. This paper presents a systematic survey of legal document summarization research from 2017 to 2024. The authors describe a paper selection methodology following Kitchenham and Charters, review 123 papers, and organize the field along several axes: datasets (Section 3), regional specifics (Section 4), extractive/abstractive/hybrid strategies (Section 5), methodological families (Section 6), and evaluation metrics including automated and human evaluation (Section 7). The paper concludes with challenges and future directions (Section 8). The central claim is that this is a comprehensive, up-to-date, and reproducible map of the legal summarization field, supported by statistics such as over 95% ROUGE usage and 20% human-evaluation adoption.

Significance. If the comprehensiveness and reproducibility claims are substantiated, this would be the current reference survey for legal summarization, filling a genuine gap between older surveys and the transformer-era literature. The taxonomy of strategies, the dataset inventory, and the structured account of human-evaluation dimensions (Section 7.2, Figure 5, Table 5) are potentially valuable for both newcomers and practitioners. The headline statistics about metric usage and evaluation practices would be useful field-level facts. However, the survey's utility is strongly coupled to the auditability of its 123-paper sample and the correctness of its regional and dataset classifications; both currently have load-bearing issues that need to be resolved before the survey can be relied upon.

major comments (4)
  1. [Section 2 (Paper selection methodology), Table 2] The reported search query is not a reproducible retrieval string. As written, ("legal" OR "law" OR "legislative") AND ("summarization" OR "summary") AND ("of" OR "about") AND ("dialogues" OR "question and answering" OR "text" OR "textual" OR "text-based" OR "document") would, if applied to titles and abstracts, exclude the paper's own reference [117] ("Legal Case Document Summarization: Extractive and Abstractive Methods and their Evaluation") because that title contains neither "of" nor "about". If applied to full text, the mandatory connector is vacuous, and the content-term disjunction omits high-yield terms such as "court", "judgment", "ruling", and "case". The manuscript reports 262 initial hits and 123 final papers but provides no per-repository hit counts, no flow diagram, and no list of the 123 included papers, so the denominator underlying the headline statistics in Sections 7.1 and 7.2 (95% ROUGE, 20% human evaluation, 9% expert evaluation) cannot be audited. Because the central claim is comprehensiveness, this is a load-bearing reproducibility gap.
  2. [Section 4 (European specifics)] The paragraph on "European specifics" states that "Multi-LexSum released in [115] is another dataset pertaining to the European landscape" and describes it as containing expert summaries based on CRLC publications. This directly contradicts Section 3.1, where Multi-LexSum is correctly described as a dataset of 40,000 source documents from roughly 4,500 federal U.S. civil rights lawsuits sourced from the Civil Rights Litigation Clearinghouse. The misclassification affects the regional analysis, including Figure 2's country distribution, and needs to be corrected.
  3. [Table 4 (Dataset overview)] Two dataset entries in Table 4 conflict with the running text. First, KorCase_Summ is listed under the domain "Privacy Policy", although Section 4 (Asia) describes it as containing precedents of the Korean Court, and the provided URL points to the Korean legal information portal law.go.kr. Second, CLSum-HK is listed with a size of 793 documents, while Section 3.1 states that "CLSum-HK contains 233 judgments and their respective press summaries from the legal reference system of Hong Kong". These inconsistencies undermine the reliability of the dataset inventory.
  4. [Section 6 (Methods for legal summarization), paragraph 'Transformers combined with LSTMs or CNNs'] The sentence "Authors of [33] fuse topic vectors into an LSTM for improving its ability to extract legal text features" misattributes the method. Reference [33] is Dong et al. (2019), "Unified Language Model Pre-training for Natural Language Understanding and Generation" (UNILM), which is a transformer pretraining paper and does not fuse topic vectors into an LSTM. The topic-vector LSTM method appears to be described elsewhere in the literature; this citation error needs to be corrected, as it directly affects the methodological taxonomy in Section 6.
minor comments (6)
  1. [Section 2 (Paper selection methodology)] The inclusion criteria bullets contain a grammar and clarity issue: "Papers with datasets and metrics, we include paper before 2017" should be rephrased to state the intended exception for pre-2017 dataset/metric papers unambiguously.
  2. [Section 3.1 (Court rulings)] There is a typo in the Multi-LexSum description: "muti-document summaries" should be "multi-document summaries".
  3. [Section 7 (Heading)] The heading "Evalaution metrics for legal summarization" contains a typo and should read "Evaluation metrics for legal summarization".
  4. [Table 9 (Appendix A.2)] The dataset name "IN-abs" in the hybrid-methods table is inconsistently capitalized relative to the "IN-Abs" and "IN-abs" usages elsewhere; standardize the capitalization.
  5. [Section 7.1 (Automated evaluation metrics)] The metric name "FActCC" appears to be a typo for "FactCC" (Kryscinski et al., 2020); please verify and correct.
  6. [References [114], [115], [116]] Multi-LexSum is cited three times with overlapping content: [114] and [115] are both the 2022 arXiv version, and [116] is the NeurIPS 2022 version. These should be merged or clearly distinguished to avoid duplicate citations.

Circularity Check

0 steps flagged · score 1.0 of 10

No material circularity: the survey's descriptive claims rest on its own selection procedure and on external benchmarks, with only two minor self-citations that are not load-bearing.

full rationale

The paper is a systematic survey rather than a derivation, so the circularity tests must be applied to its load-bearing argument: the claim to have identified 123 relevant papers and the resulting trend statistics. That claim is supported by an explicit repository list, a stated query, inclusion/exclusion criteria, and a manual screening agreement among the first three authors, none of which is defined in terms of the survey's conclusions. The self-citations are confined to background support: [4] (Akter et al., 2022) is cited together with two independent references to illustrate that ROUGE's limitations are widely explored, and [19] (Cano and Morisio, 2017) is cited only as an example of a study following the Kitchenham and Charters systematic-review guidelines, which are themselves cited as the external method [66]. Neither citation is used to forbid alternatives, to define a quantity in terms of a target result, or to supply a premise from which the survey's statistics are derived. The headline statistics (about 95% ROUGE usage; 20% human evaluation) are descriptive aggregations of the collected papers, not fitted inputs renamed as predictions. The skeptics' concerns about the non-reproducible search query and the unpublished list of 123 studies are auditability and sampling-frame concerns, not circularity: they attack whether the sample represents the field, not whether the paper's conclusions reduce by construction to its inputs. Accordingly, the appropriate circularity score is 1.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

The survey introduces no free parameters and no invented entities. Its descriptive claims rest on the domain assumptions listed above about search coverage, screening reliability, and dataset ground-truth validity.

assumptions (3)
  • domain assumption The chosen repositories, keyword query, and inclusion/exclusion criteria recover a representative and near-complete set of legal summarization papers from 2017 to 2024.
    Section 2 defines the search strategy but provides no validation against a gold-standard reference list, so coverage completeness is assumed.
  • domain assumption Manual title and abstract screening, with inclusion decisions agreed among the first three authors, is reliable enough to reproduce the final set of 123 papers.
    Section 2 reports consensus but gives no inter-annotator agreement, no duplicate resolution counts, and no screening artifact.
  • domain assumption Reference summaries in the underlying datasets, such as Ementa, headnotes, and press summaries, are valid ground-truth summaries for legal summarization tasks.
    Section 3 catalogs these datasets and treats each reference summary as gold without discussing annotation validity, bias, or quality.

how reviews work

0 comments
Cite this review

Pith. "Pith review of A Comprehensive Survey on Legal Summarization: Challenges and Future Directions." pith.science (2026). https://pith.science/paper/LYSSZEOL

@misc{pith2026250117830,
  author       = {Pith},
  title        = {Pith review of: A Comprehensive Survey on Legal Summarization: Challenges and Future Directions},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/LYSSZEOL}},
  note         = {Machine review of arXiv:2501.17830}
}
read the original abstract

This article provides a systematic up-to-date survey of automatic summarization techniques, datasets, models, and evaluation methods in the legal domain. Through specific source selection criteria, we thoroughly review over 120 papers spanning the modern `transformer' era of natural language processing (NLP), thus filling a gap in existing systematic surveys on the matter. We present existing research along several axes and discuss trends, challenges, and opportunities for future research.

Figures

Figures reproduced from arXiv: 2501.17830 by the authors.

Figure 1
Figure 1. Schematic Legal Summarization Pipeline: Legal summarization pipelines process lengthy, structurally diverse documents to [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Countries identified in the collection of papers during the survey study [PITH_FULL_IMAGE:figures/full_fig_p007_2.png] view at source ↗
Figure 3
Figure 3. Legal summarization research trends for last 5 years [PITH_FULL_IMAGE:figures/full_fig_p008_3.png] view at source ↗
Figures from the paper (2 more)
Figure 4
Figure 4. Figure 4: Taxonomy of legal summarization strategies categorized into three main approaches, each with more specific sub-categories. [PITH_FULL_IMAGE:figures/full_fig_p010_4.png]
Figure 5
Figure 5. Figure 5: Overview of key aspects and dimensions in human evaluation for legal summarization [PITH_FULL_IMAGE:figures/full_fig_p016_5.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

168 extracted references · 60 canonical work pages

  1. [33]

    2019.Unified language model pre-training for natural language understanding and generation

    Li Dong, Nan Yang, Wenhui Wang, Furu Wei, Xiaodong Liu, Yu Wang, Jianfeng Gao, Ming Zhou, and Hsiao-Wuen Hon. 2019.Unified language model pre-training for natural language understanding and generation . Curran Associates Inc., Red Hook, NY, USA

  2. [117]

    Abhay Shukla, Paheli Bhattacharya, Soham Poddar, Rajdeep Mukherjee, Kripabandhu Ghosh, Pawan Goyal, and Saptarshi Ghosh. 2022. Legal Case Document Summarization: Extractive and Abstractive Methods and their Evaluation. In AACL/IJCNLP (1), Yulan He, Heng Ji, Sujian Li, Yang Liu, and Chua-Hui Chang (Eds.). Association for Computational Linguistics, Online o...

  3. [115]

    Zejiang Shen, Kyle Lo, Lauren Yu, Nathan Dahlberg, Margo Schlanger, and Doug Downey. 2022. Multi-LexSum: Real-World Summaries of Civil Rights Lawsuits at Multiple Granularities. https://arxiv.org/abs/2206.10883

  4. [1]

    Mean Average Precision

    2009. Mean Average Precision. In Encyclopedia of Database Systems . Springer US, 1703

  5. [2]

    Flavia Achena, David Preti, Davide Venditti, Leonardo Ranaldi, Cristina Giannone, Fabio Massimo Zanzotto, Andrea Favalli, Raniero Romagnoli, et al. 2023. Legal Summarization: to Each Court its Own Model.. In CLiC-it

  6. [3]

    Kanika Agrawal. 2020. Legal case summarization: An application for text summarization. In2020 International conference on computer communication and informatics (ICCCI). IEEE, 1–6

  7. [4]

    Mousumi Akter, Naman Bansal, and Shubhra Kanti Karmaker Santu. 2022. Revisiting Automatic Evaluation of Extractive Summarization Task: Can We Do Better than ROUGE?. In ACL (Findings). Association for Computational Linguistics, 1547–1560

  8. [5]

    Ryan Amos, Gunes Acar, Eli Lucherini, Mihir Kshirsagar, Arvind Narayanan, and Jonathan Mayer. 2021. Privacy Policies over Time: Curation and Analysis of a Million-Document Dataset. In Proceedings of the Web Conference 2021 (WWW ’21) . ACM, 2165–2176. doi:10.1145/3442381.3450048

Show all 168 references
  1. [6]

    Deepa Anand and Rupali Wagh. 2022. Effective deep learning approaches for summarization of legal texts. Journal of King Saud University - Computer and Information Sciences 34, 5 (2022), 2141–2150. doi:10.1016/j.jksuci.2019.11.015

  2. [7]

    Farid Ariai and Gianluca Demartini. 2024. Natural Language Processing for the Legal Domain: A Survey of Tasks, Datasets, Models, and Challenges. arXiv preprint arXiv:2410.21306 (2024)

  3. [8]

    Dennis Aumiller, Ashish Chouhan, and Michael Gertz. 2022. EUR-Lex-Sum: A Multi- and Cross-lingual Dataset for Long-form Summarization in the Legal Domain

  4. [9]

    Dennis Aumiller, Ashish Chouhan, and Michael Gertz. 2022. EUR-Lex-Sum: A Multi- and Cross-lingual Dataset for Long-form Summarization in the Legal Domain. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing . Association for Computational ...

  5. [10]

    Ahsaas Bajaj, Pavitra Dangati, Kalpesh Krishna, Pradhiksha Ashok Kumar, Rheeya Uppaal, Bradford Windsor, Eliot Brenner, Dominic Dotterrer, Rajarshi Das, and Andrew McCallum. 2021. Long Document Summarization in a Low Resource Setting using Pretrained Language Models. In Procee...

  6. [11]

    Ahsaas Bajaj, Pavitra Dangati, Kalpesh Krishna, Pradhiksha Ashok Kumar, Rheeya Uppaal, Bradford Windsor, Eliot Brenner, Dominic Dotterrer, Rajarshi Das, and Andrew McCallum. 2021. Long Document Summarization in a Low Resource Setting using Pretrained Language Models. In ACL (s...

  7. [12]

    Satanjeev Banerjee and Alon Lavie. 2005. METEOR: An Automatic Metric for MT Evaluation with Improved Correlation with Human Judgments. In IEEvaluation@ACL. Association for Computational Linguistics, 65–72

  8. [13]

    Vinayshekhar Bannihatti Kumar, Kasturi Bhattacharjee, and Rashmi Gangadharaiah. 2022. Towards Cross-Domain Transferability of Text Generation Models for Legal Text. In Proceedings of the Natural Legal Language Processing Workshop 2022 , Nikolaos Aletras, Ilias Chalkidis, Lesli...

  9. [14]

    Emmanuel Bauer, Dominik Stammbach, Nianlong Gu, and Elliott Ash. 2023. Legal Extractive Summarization of U.S. Court Opinions. In LIRAI@HT (CEUR Workshop Proceedings, Vol. 3594). CEUR-WS.org, 10–30

  10. [15]

    Peters, and Arman Cohan

    Iz Beltagy, Matthew E. Peters, and Arman Cohan. 2020. Longformer: The Long-Document Transformer. CoRR abs/2004.05150 (2020)

  11. [16]

    Paheli Bhattacharya, Kaustubh Hiware, Subham Rajgaria, Nilay Pochhi, Kripabandhu Ghosh, and Saptarshi Ghosh. 2019. A Comparative Study of Summarization Algorithms Applied to Legal Case Judgments. In ECIR (1) (Lecture Notes in Computer Science, Vol. 11437) . Springer, 413–428

  12. [18]

    Purnima Bindal, Vikas Kumar, Vasudha Bhatnagar, Parikshet Sirohi, and Ashwini Siwal. 2023. Citation-Based Summarization of Landmark Judgments. In Proceedings of the 20th International Conference on Natural Language Processing (ICON) , Jyoti D. Pawar and Sobha Lalitha Devi (Eds...

  13. [19]

    Erion Çano and Maurizio Morisio. 2017. Hybrid recommender systems: A systematic literature review. Intelligent Data Analysis 21, 6 (2017), 1487–1524. doi:10.3233/IDA-163209

  14. [20]

    Ilias Chalkidis, Manos Fergadiotis, Prodromos Malakasiotis, Nikolaos Aletras, and Ion Androutsopoulos. 2020. LEGAL-BERT: The Muppets straight out of Law School. https://arxiv.org/abs/2010.02559

  15. [21]

    Ilias Chalkidis and Dimitrios Kampas. 2019. Deep learning in law: early adaptation and legal word embeddings trained on large corpora. Artif. Intell. Law 27, 2 (2019), 171–198

  16. [22]

    Zhe Chen, Lin Ye, and Hongli Zhang. 2023. Enhancing LSTM and Fusing Articles of Law for Legal Text Summarization. In ICONIP (14) (Communi- cations in Computer and Information Science, Vol. 1968) . Springer, 110–124

  17. [23]

    Jianpeng Cheng and Mirella Lapata. 2016. Neural Summarization by Extracting Sentences and Words. In ACL (1). The Association for Computer Linguistics. Manuscript submitted to ACM A Comprehensive Survey on Legal Summarization 23

  18. [24]

    Jingpei Dan, Weixuan Hu, and Yuming Wang. 2023. Enhancing legal judgment summarization with integrated semantic and structural information. Artificial Intelligence and Law (2023), 1–22

  19. [25]

    Debtanu Datta, Shubham Soni, Rajdeep Mukherjee, and Saptarshi Ghosh. 2023. MILDSum: A Novel Benchmark Dataset for Multilingual Summarization of Indian Legal Case Judgments. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing , Houda Bouamo...

  20. [26]

    Leonardo de Andrade and Karin Becker. 2023. BB25HLegalSum: Leveraging BM25 and BERT-Based Clustering for the Summarization of Legal Documents. In Proceedings of the 14th International Conference on Recent Advances in Natural Language Processing , Ruslan Mitkov and Galia Angelo...

  21. [27]

    Diego de Vargas Feijó and Viviane Pereira Moreira. 2018. Rulingbr: A summarization dataset for legal texts. In International Conference on Computational Processing of the Portuguese Language . Springer, 255–264. https://link.springer.com/chapter/10.1007/978-3-319-99722-3_26

  22. [28]

    Diego de Vargas Feijó and Viviane P. Moreira. 2023. Improving abstractive summarization of legal rulings through textual entailment. Artif. Intell. Law 31, 1 (2023), 91–113

  23. [29]

    Aniket Deroy, Paheli Bhattacharya, Kripabandhu Ghosh, and Saptarshi Ghosh. 2021. An Analytical Study of Algorithmic and Expert Summaries of Legal Cases. In International Conference on Legal Knowledge and Information Systems . https://api.semanticscholar.org/CorpusID:245009600

  24. [30]

    Aniket Deroy, Kripabandhu Ghosh, and Saptarshi Ghosh. 2023. How ready are pre-trained abstractive models and LLMs for legal case judgement summarization? arXiv preprint arXiv:2306.01248 (2023)

  25. [31]

    Aniket Deroy, Kripabandhu Ghosh, and Saptarshi Ghosh. 2024. Applicability of Large Language Models and Generative Models for Legal Case Judgement Summarization. CoRR abs/2407.12848 (2024)

  26. [32]

    Aniket Deroy, Kripabandhu Ghosh, and Saptarshi Ghosh. 2024. Ensemble methods for improving extractive summarization of legal case judgements. Artif. Intell. Law 32, 1 (2024), 231–289. doi:10.1007/s10506-023-09349-8

  27. [34]

    Xinyu Duan, Yating Zhang, Lin Yuan, Xin Zhou, Xiaozhong Liu, Tianyi Wang, Ruocheng Wang, Qiong Zhang, Changlong Sun, and Fei Wu. 2019. Legal Summarization for Multi-role Debate Dialogue via Controversy Focus Mining and Multi-task Learning. In CIKM. ACM, 1361–1370

  28. [35]

    Mohamed Elaraby and Diane Litman. 2022. ArgLegalSumm: Improving abstractive summarization of legal documents with argument mining. arXiv preprint arXiv:2209.01650 (2022)

  29. [36]

    Mohamed Elaraby, Huihui Xu, Morgan Gray, Kevin D Ashley, and Diane Litman. 2024. Adding Argumentation into Human Evaluation of Long Document Abstractive Summarization: A Case Study on Legal Opinions. In Proceedings of the Fourth Workshop on Human Evaluation of NLP Systems (Hum...

  30. [37]

    Mohamed Elaraby, Yang Zhong, and Diane Litman. 2023. Towards Argument-Aware Abstractive Summarization of Long Legal Opinions with Summary Reranking. In Findings of the Association for Computational Linguistics: ACL 2023 , Anna Rogers, Jordan Boyd-Graber, and Naoaki Okazaki (Ed...

  31. [38]

    Ahmed Elnaggar, Christoph Gebendorfer, Ingo Glaser, and Florian Matthes. 2018. Multi-Task Deep Learning for Legal Document Translation, Summarization and Multi-Label Classification. https://arxiv.org/abs/1810.07513

  32. [39]

    Fabbri, Wojciech Kryscinski, Bryan McCann, Caiming Xiong, Richard Socher, and Dragomir R

    Alexander R. Fabbri, Wojciech Kryscinski, Bryan McCann, Caiming Xiong, Richard Socher, and Dragomir R. Radev. 2021. SummEval: Re-evaluating Summarization Evaluation. Trans. Assoc. Comput. Linguistics 9 (2021), 391–409

  33. [40]

    Cristhian Figueroa, Iacopo Vagliano, Oscar Rodriguez Rocha, and Maurizio Morisio. 2015. A systematic literature review of Linked Data-based recommender systems. Concurr. Comput. Pract. Exp. 27, 17 (2015), 4659–4684. doi:10.1002/CPE.3449

  34. [41]

    Filippo Galgani. 2010. Legal Case Reports. UCI Machine Learning Repository. DOI: https://doi.org/10.24432/C5ZS41

  35. [42]

    Filippo Galgani and Achim Hoffmann. 2011. LEXA: Towards Automatic Legal Citation Classification. In AI 2010: Advances in Artificial Intelligence , Jiuyong Li (Ed.). Springer Berlin Heidelberg, Berlin, Heidelberg, 445–454

  36. [43]

    Aashish Ghimire, Raj Shrestha, and John Edwards. 2023. Too Legal; Didn’t Read (TLDR): Summarization of Court Opinions. In 2023 Intermountain Engineering, Technology and Computing (IETC) . IEEE, 164–169

  37. [44]

    Satyajit Ghosh, Mousumi Dutta, and Tanaya Das. 2022. Indian legal text summarization: A text normalization-based approach. In 2022 IEEE 19th India Council International Conference (INDICON) . IEEE, 1–4

  38. [45]

    Ingo Glaser, Sebastian Moser, and Florian Matthes. 2021. Summarization of German Court Rulings. In Proceedings of the Natural Legal Language Processing Workshop 2021. Association for Computational Linguistics, Punta Cana, Dominican Republic, 180–189. https://aclanthology.org/2...

  39. [46]

    Claire Grover, Ben Hachey, and Ian Hughson. 2004. The HOLJ Corpus. Supporting Summarisation of Legal Texts. In Proceedings of the 5th International Workshop on Linguistically Interpreted Corpora , Silvia Hansen-Schirra, Stephan Oepen, and Hans Uszkoreit (Eds.). COLING, Geneva,...

  40. [47]

    Karunya Harikrishnan, Malathi M, and Sundharakumar K B. 2024. Topic-Driven Contractual Language Understanding and Summarization: An Integrated Approach for Simplifying Legal Documents. In 2024 4th International Conference on Intelligent Technologies (CONIT) . IEEE, 1–6. doi:10...

  41. [48]

    S Hochreiter. 1997. Long Short-term Memory. Neural Computation MIT-Press (1997). Manuscript submitted to ACM 24 Mousumi Akter, Erion Çano, Erik Weber, Dennis Dobler, and Ivan Habernal

  42. [49]

    Yu-Xiang Hong and Chia-Hui Chang. 2023. Improving colloquial case legal judgment prediction via abstractive text summarization. Comput. Law Secur. Rev. 51 (2023), 105863

  43. [50]

    Yue Huang, Lijuan Sun, Chong Han, and Jian Guo. 2023. A high-precision two-stage legal judgment summarization. Mathematics 11, 6 (2023), 1320. doi:10.3390/math11061320

  44. [51]

    Yuxin Huang, Zhengtao Yu, Junjun Guo, Yan Xiang, and Yantuan Xian. 2021. Element graph-augmented abstractive summarization for legal public opinion news with graph transformer. Neurocomputing 460 (2021), 166–180

  45. [52]

    Yuxin Huang, Zhengtao Yu, Junjun Guo, Zhiqiang Yu, and Yantuan Xian. 2020. Legal public opinion news abstractive summarization by incorporating topic information. Int. J. Mach. Learn. Cybern. 11, 9 (2020), 2039–2050

  46. [53]

    Deepali Jain, Malaya Dutta Borah, and Anupam Biswas. 2020. Fine-Tuning Textrank for Legal Document Summarization: A Bayesian Optimization Based Approach. In FIRE. ACM, 41–48

  47. [54]

    Deepali Jain, Malaya Dutta Borah, and Anupam Biswas. 2021. Automatic summarization of legal bills: A comparative analysis of classical extractive approaches. In 2021 international conference on computing, communication, and intelligent systems (ICCCIS) . IEEE, 394–400

  48. [55]

    Deepali Jain, Malaya Dutta Borah, and Anupam Biswas. 2021. CAWESumm: A Contextual and Anonymous Walk Embedding Based Extractive Summarization of Legal Bills. In Proceedings of the 18th International Conference on Natural Language Processing (ICON 2021), National Institute of T...

  49. [56]

    Deepali Jain, Malaya Dutta Borah, and Anupam Biswas. 2021. Summarization of legal documents: Where are we now and the way forward.Comput. Sci. Rev. 40 (2021), 100388

  50. [57]

    Deepali Jain, Malaya Dutta Borah, and Anupam Biswas. 2023. Bayesian Optimization based Score Fusion of Linguistic Approaches for Improving Legal Document Summarization. Knowl. Based Syst. 264 (2023), 110336

  51. [58]

    Deepali Jain, Malaya Dutta Borah, and Anupam Biswas. 2024. A sentence is known by the company it keeps: Improving Legal Document Summarization Using Deep Clustering. Artif. Intell. Law 32, 1 (2024), 165–200

  52. [59]

    Deepali Jain, Malaya Dutta Borah, and Anupam Biswas. 2024. Summarization of Lengthy Legal Documents via Abstractive Dataset Building: An Extract-then-Assign Approach. Expert Syst. Appl. 237, Part B (2024), 121571

  53. [60]

    Kalervo Järvelin and Jaana Kekäläinen. 2002. Cumulated gain-based evaluation of IR techniques. ACM Trans. Inf. Syst. 20, 4 (2002), 422–446

  54. [61]

    Changzhen Ji, Yating Zhang, Adam Jatowt, and Haipang Wu. 2023. CDD: A Large Scale Dataset for Legal Intelligence Research. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing: Industry Track , Mingxuan Wang and Imed Zitouni (Eds.). Associa...

  55. [62]

    Prathamesh Kalamkar, Aman Tiwari, Astha Agarwal, Saurabh Karn, Smita Gupta, Vivek Raghavan, and Ashutosh Modi. 2022. Corpus for Automatic Structuring of Legal Documents. In Proceedings of the Thirteenth Language Resources and Evaluation Conference , Nicoletta Calzolari, Frédér...

  56. [63]

    Ambedkar Kanapala, Sukomal Pal, and Rajendra Pamula. 2019. Text summarization from legal documents: a survey. Artif. Intell. Rev. 51, 3 (2019), 371–402

  57. [64]

    Georgia Kapitsaki and Maria Papoutsoglou. 2024. A privacy policies dataset in Greek in the GDPR era. In Proceedings of the 27th Pan-Hellenic Conference on Progress in Computing and Informatics (Lamia, Greece) (PCI ’23). Association for Computing Machinery, New York, NY, USA, 1...

  58. [65]

    Moniba Keymanesh, Micha Elsner, and Srinivasan Sarthasarathy. 2020. Toward Domain-Guided Controllable Summarization of Privacy Policies.. In NLLP@ KDD. 18–24

  59. [66]

    Barbara Ann Kitchenham and Stuart Charters. 2007. Guidelines for performing Systematic Literature Reviews in Software Engineering . Technical Report EBSE 2007-001. Keele University and Durham University Joint Report. https://www.elsevier.com/__data/promis_misc/ 525444systemati...

  60. [67]

    Svea Klaus, Ria Van Hecke, Kaweh Djafari Naini, Ismail Sengor Altingovde, Juan Bernabé-Moreno, and Enrique Herrera-Viedma. 2022. Summarizing Legal Regulatory Documents using Transformers. In SIGIR. ACM, 2426–2430

  61. [68]

    Svea Klaus, Ria Van Hecke, Kaweh Djafari Naini, Ismail Sengor Altingovde, Juan Bernabé-Moreno, and Enrique Herrera-Viedma. 2022. Summarizing Legal Regulatory Documents using Transformers. In Proceedings of the 45th International ACM SIGIR Conference on Research and Development...

  62. [69]

    Marios Koniaris, Dimitris Galanis, Eugenia Giannini, and Panayiotis Tsanakas. 2023. Evaluation of Automatic Legal Text Summarization Techniques for Greek Case Law. Information 14, 4 (2023). doi:10.3390/info14040250

  63. [70]

    Anastassia Kornilova and Vladimir Eidelman. 2019. BillSum: A Corpus for Automatic Summarization of US Legislation. In Proceedings of the 2nd Workshop on New Frontiers in Summarization, Lu Wang, Jackie Chi Kit Cheung, Giuseppe Carenini, and Fei Liu (Eds.). Association for Compu...

  64. [71]

    Alex Krizhevsky, Ilya Sutskever, and Geoffrey E. Hinton. 2017. ImageNet classification with deep convolutional neural networks. Commun. ACM 60, 6 (2017), 84–90

  65. [72]

    Wojciech Kryscinski, Nitish Shirish Keskar, Bryan McCann, Caiming Xiong, and Richard Socher. 2019. Neural Text Summarization: A Critical Evaluation. In EMNLP/IJCNLP (1). Association for Computational Linguistics, 540–551. Manuscript submitted to ACM A Comprehensive Survey on L...

  66. [73]

    Wojciech Kryscinski, Bryan McCann, Caiming Xiong, and Richard Socher. 2020. Evaluating the Factual Consistency of Abstractive Text Summa- rization. In EMNLP (1). Association for Computational Linguistics, 9332–9346

  67. [74]

    Kurisinkel and Nancy Chen

    Litton J. Kurisinkel and Nancy Chen. 2022. Tractable & Coherent Multi-Document Summarization: Discrete Optimization of Multiple Neural Modeling Streams via Integer Linear Programming. In EMNLP (Industry Track). Association for Computational Linguistics, 237–243

  68. [75]

    Bennett, and Marti A

    Philippe Laban, Tobias Schnabel, Paul N. Bennett, and Marti A. Hearst. 2022. SummaC: Re-Visiting NLI-based Models for Inconsistency Detection in Summarization. Trans. Assoc. Comput. Linguistics 10 (2022), 163–177

  69. [76]

    Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer. 2020. BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension. In ACL. Association for C...

  70. [78]

    Chin-Yew Lin. 2004. ROUGE: A Package for Automatic Evaluation of Summaries. InText Summarization Branches Out. Association for Computational Linguistics, Barcelona, Spain, 74–81. https://aclanthology.org/W04-1013

  71. [79]

    Shuaiqi Liu, Jiannong Cao, Yicong Li, Ruosong Yang, and Zhiyuan Wen. 2024. Low-resource court judgment summarization for common law systems. Information Processing & Management 61, 5 (2024), 103796. doi:10.1016/j.ipm.2024.103796

  72. [80]

    Emilia Lukose, Suparna De, and Jon Johnson. 2022. Privacy Pitfalls of Online Service Terms and Conditions: a Hybrid Approach for Classification and Summarization. In NLLP@EMNLP. Association for Computational Linguistics, 65–75

  73. [81]

    Yinglong Ma, Peng Zhang, and Jiangang Ma. 2018. An Ontology Driven Knowledge Block Summarization Approach for Chinese Judgment Document Classification. IEEE Access 6 (2018), 71327–71338

  74. [82]

    Manuj Malik, Zheng Zhao, Marcio Fonseca, Shrisha Rao, and Shay B. Cohen. 2024. CivilSum: A Dataset for Abstractive Summarization of Indian Court Decisions. In SIGIR. ACM, 2241–2250

  75. [83]

    Arpan Mandal, Paheli Bhattacharya, Sekhar Mandal, and Saptarshi Ghosh. 2021. Improving Legal Case Summarization Using Document-Specific Catchphrases. In JURIX (Frontiers in Artificial Intelligence and Applications, Vol. 346) . IOS Press, 76–81

  76. [84]

    Laura Manor and Junyi Jessy Li. 2019. Plain English Summarization of Contracts. In Proceedings of the Natural Legal Language Processing Workshop

  77. [85]

    McDonald

    Joshua Maynez, Shashi Narayan, Bernd Bohnet, and Ryan T. McDonald. 2020. On Faithfulness and Factuality in Abstractive Summarization. In ACL. Association for Computational Linguistics, 1906–1919

  78. [86]

    Marie Chantelle Cruz Medina, Lucas Matheus Da Silva Oliveira, Jean Felipe Coelho Ferreira, Leandro Honorato de Souza Silva, Cleyton Mário de Oliveira Rodrigues, João Fausto Lorenzato de Oliveira, Paulo Christiano Sobral, Bruno Souza, Dionizio Feitosa, and Bruno J. T. Fernandes...

  79. [87]

    Kaiz Merchant and Yash Pande. 2018. NLP Based Latent Semantic Analysis for Legal Text Summarization. In ICACCI. IEEE, 1803–1807

  80. [88]

    Gianluca Moro, Nicola Piscaglia, Luca Ragazzi, and Paolo Italiani. 2024. Multi-language transfer learning for low-resource legal case summarization. Artif. Intell. Law 32, 4 (2024), 1111–1139

  81. [89]

    Akheel Muhammed, Hamna Muslihuddeen, Shalaka Sankar, and M Anand Kumar. 2024. Impact of Rhetorical Roles in Abstractive Legal Document Summarization. In 2024 5th International Conference on Innovative Trends in Information Technology (ICITIIT) . IEEE, 1–6

  82. [90]

    Raghav, and Roshni Kar

    Ankan Mullick, Abhilash Nandy, Manav Nitin Kapadnis, Sohan Patnaik, R. Raghav, and Roshni Kar. 2022. An Evaluation Framework for Legal Document Summarization. In LREC. European Language Resources Association, 4747–4753

  83. [91]

    Duy-Hung Nguyen, Bao-Sinh Nguyen, Nguyen-Viet-Dung Nghiem, Dung Tien Le, Mim Amina Khatun, Minh-Tien Nguyen, and Hung Le. 2021. Robust Deep Reinforcement Learning for Extractive Legal Summarization. In ICONIP (6) (Communications in Computer and Information Science, Vol. 1517)....

  84. [93]

    Joel Niklaus and Daniele Giofré. 2023. Can we Pretrain a SotA Legal Language Model on a Budget From Scratch?. In SustaiNLP. Association for Computational Linguistics, 158–182

  85. [94]

    Joel Niklaus, Veton Matoshi, Matthias Stürmer, Ilias Chalkidis, and Daniel Ho. 2024. MultiLegalPile: A 689GB Multilingual Legal Corpus. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , Lun-Wei Ku, Andre Martin...

  86. [95]

    OpenAI. 2023. GPT-4 Technical Report. CoRR abs/2303.08774 (2023)

  87. [96]

    Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002. Bleu: a Method for Automatic Evaluation of Machine Translation. InACL. ACL, 311–318

  88. [97]

    Vedant Parikh, Vidit Mathur, Parth Mehta, Namita Mittal, and Prasenjit Majumder. 2021. LawSum: A weakly supervised approach for Indian Legal Document Summarization. CoRR abs/2110.01188 (2021). https://arxiv.org/abs/2110.01188

  89. [98]

    Shounak Paul, Arpan Mandal, Pawan Goyal, and Saptarshi Ghosh. 2023. Pre-trained Language Models for the Legal Domain: A Case Study on Indian Law. https://arxiv.org/abs/2209.06049 Manuscript submitted to ACM 26 Mousumi Akter, Erion Çano, Erik Weber, Dennis Dobler, and Ivan Habernal

  90. [99]

    Seth Polsley, Pooja Jhunjhunwala, and Ruihong Huang. 2016. CaseSummarizer: A System for Automated Summarization of Legal Texts. In Proceedings of COLING 2016, the 26th International Conference on Computational Linguistics: System Demonstrations , Hideo Watanabe (Ed.). The COLI...

  91. [100]

    Priyanka Prabhakar, Deepa Gupta, and Peeta Basa Pati. 2022. Abstractive Summarization of Indian Legal Judgments. In OCIT. IEEE, 256–261. doi:10.1109/OCIT56763.2022.00056

  92. [101]

    Bindal Purnima, Kumar Vikas, Bhatnagar Vasudha, Sirohi Parikshet, and Siwal Ashwini. 2023. Citation-Based Summarization of Landmark Judgments. In Proceedings of the 20th International Conference on Natural Language Processing (ICON) . 588–593

  93. [102]

    Weijian Qin and Xudong Luo. 2023. A Legal News Summarisation Model Based on RoBERTa, T5 and Dilated Gated CNN. In ICTAI. IEEE, 889–897

  94. [103]

    Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020. Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer. J. Mach. Learn. Res. 21 (2020), 140:1–140:67

  95. [104]

    Thomson Reuters. 2023. Legal Drafting Challenges, Risks, and Opportunities. https://legal.thomsonreuters.com/blog/legal-drafting-challenges- risks-and-opportunities/. Accessed: 2023-02-04

  96. [105]

    Rush, Sumit Chopra, and Jason Weston

    Alexander M. Rush, Sumit Chopra, and Jason Weston. 2015. A Neural Attention Model for Abstractive Sentence Summarization. In EMNLP. The Association for Computational Linguistics, 379–389

  97. [106]

    Siddhartha Rusiya and Anupam Jamatia. 2023. Implementation of legal documents text summarization and classification by applying neural network techniques. In Machine Intelligence Techniques for Data Analysis and Signal Processing: Proceedings of the 4th International Conferenc...

  98. [107]

    Clacio, Rey Edison R

    Ria Ambrocio Sagum, Patrick Anndwin C. Clacio, Rey Edison R. Cayetano, and Airis Dale F. Lobrio. 2023. Philippine Court Case Summarizer using Latent Semantic Analysis. In ICCSCI (Procedia Computer Science, Vol. 227) . Elsevier, 474–481

  99. [108]

    Olivier Salaün, Aurore Clément Troussel, Sylvain Longhais, Hannes Westermann, Philippe Langlais, and Karim Benyekhlef. 2022. Conditional Abstractive Summarization of Court Decisions for Laymen and Insights from Human Evaluation. In JURIX (Frontiers in Artificial Intelligence a...

  100. [109]

    Abhilasha Sancheti, Aparna Garimella, Balaji Vasan Srinivasan, and Rachel Rudinger. 2023. What to Read in a Contract? Party-Specific Summarization of Legal Obligations, Entitlements, and Prohibitions. In EMNLP. Association for Computational Linguistics, 14708–14725

  101. [110]

    Ajay V Saravade and Pratiksha R Deshmukh. 2023. Improving Feature Vector for Extractive Text Summarization of Legal Judgement. In 2023 International Conference on Electrical, Electronics, Communication and Computers (ELEXCOM) . IEEE, 1–5

  102. [111]

    Ayesha Sarwar, Seemab Latif, Rabia Irfan, Adnan Ul-Hasan, and Faisal Shafait. 2022. Text Summarization from Judicial Records using Deep Neural Machines. In 2022 International Conference on Electrical, Computer, Communications and Mechatronics Engineering (ICECCME) . IEEE, 1–6

  103. [112]

    Marijn Schraagen, Floris Bex, Nick Van De Luijtgaarden, and Daniël Prijs. 2022. Abstractive Summarization of Dutch Court Verdicts Using Sequence-to-sequence Models. In Proceedings of the Natural Legal Language Processing Workshop 2022 , Nikolaos Aletras, Ilias Chalkidis, Lesli...

  104. [113]

    Saloni Sharma, Surabhi Srivastava, Pradeepika Verma, Anshul Verma, and Sachchida Nand Chaurasia. 2023. A comprehensive analysis of indian legal documents summarization techniques. SN Computer Science 4, 5 (2023), 614. doi:10.1007/s42979-023-01983-y

  105. [114]

    Zejiang Shen, Kyle Lo, Lauren Yu, Nathan Dahlberg, Margo Schlanger, and Doug Downey. 2022. Multi-LexSum: Real-World Summaries of Civil Rights Lawsuits at Multiple Granularities. CoRR abs/2206.10883 (2022). doi:10.48550/arXiv.2206.10883

  106. [116]

    Zejiang Shen, Kyle Lo, Lauren Yu, Nathan Dahlberg, Margo Schlanger, and Doug Downey. 2024. Multi-LexSum: real-world summaries of civil rights lawsuits at multiple granularities. In Proceedings of the 36th International Conference on Neural Information Processing Systems (New O...

  107. [118]

    Samarth Singhal, Siddhant Singh, Sandeep Yadav, and Anil Singh Parihar. 2023. LTSum: Legal Text Summarizer. In ICCCNT. IEEE, 1–6

  108. [119]

    Dezhao Song, Sally Gao, Baosheng He, and Frank Schilder. 2022. On the Effectiveness of Pre-Trained Language Models for Legal Natural Language Processing: An Empirical Study. IEEE Access 10 (2022), 75835–75858. doi:10.1109/ACCESS.2022.3190408

  109. [120]

    Bianca Steffes, Piotr Rataj, Luise Burger, and Lukas Roth. 2023. On evaluating legal summaries with ROUGE. In Proceedings of the Nineteenth International Conference on Artificial Intelligence and Law (Braga, Portugal) (ICAIL ’23). Association for Computing Machinery, New York,...

  110. [121]

    Sheetal A Takale, Sandeep A Thorat, and Rajani S Sajjan. 2022. Legal document summarization using ripple down rules. In 2022 IEEE International Women in Engineering (WIE) Conference on Electrical and Computer Engineering (WIECON-ECE) . IEEE, 78–83

  111. [122]

    Tran, Minh Le Nguyen, Satoshi Tojo, and Ken Satoh

    Vu D. Tran, Minh Le Nguyen, Satoshi Tojo, and Ken Satoh. 2020. Encoded summarization: summarizing documents into continuous vector space for legal case retrieval. Artif. Intell. Law 28, 4 (2020), 441–467

  112. [123]

    Mehta, and Jenish Dhanani

    Aashka Trivedi, Anya Trivedi, Sourabh Varshney, Vidhey Joshipura, Rupa G. Mehta, and Jenish Dhanani. 2020. Extracted Summary Based Recommendation System for Indian Legal Documents. In ICCCNT. IEEE, 1–6. Manuscript submitted to ACM A Comprehensive Survey on Legal Summarization 27

  113. [124]

    Santosh T.y.s.s., Mahmoud Aly, and Matthias Grabmair. 2024. LexAbSumm: Aspect-based Summarization of Legal Decisions. In Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024) , Nicoletta Calzol...

  114. [125]

    Surisetty Hima Varshini, Gottimukkala Sarayu Varma, Priyanka Prabhakar, and Peeta Basa Pati. 2024. Comparative Analysis of Legal Outcome Prediction with Detailed and Summarized Text. In 2024 IEEE 9th International Conference for Convergence in Technology (I2CT) . 1–6. doi:10.1...

  115. [126]

    Gomez, Lukasz Kaiser, and Illia Polosukhin

    Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017. Attention Is All You Need. In Advances in Neural Information Processing Systems 30 . Curran Associates, Inc., Long Beach, CA, USA, 5998–6008

  116. [127]

    Huihui Xu and Kevin Ashley. 2023. A Question-Answering Approach to Evaluating Legal Summaries . IOS Press. doi:10.3233/faia230977

  117. [128]

    Huihui Xu and Kevin D. Ashley. 2023. Argumentative Segmentation Enhancement for Legal Summarization. In ASAIL@ICAIL (CEUR Workshop Proceedings, Vol. 3441). CEUR-WS.org, 141–150

  118. [129]

    Huihui Xu, Jaromir Savelka, and Kevin D. Ashley. 2021. Toward summarizing case decisions via extracting argument issues, reasons, and conclusions. In Proceedings of the Eighteenth International Conference on Artificial Intelligence and Law (São Paulo, Brazil) (ICAIL ’21). Asso...

  119. [130]

    Huihui Xu, Jaromír Šavelka, and Kevin D. Ashley. 2020. Using argument mining for legal text summarization. In IOS Press, Vol. 334. IOS Press, 184. doi:10.3233/faia200862

  120. [131]

    Hiroaki Yamada, Simone Teufel, and Takenobu Tokunaga. 2017. Annotation of argument structure in Japanese legal documents. In ArgMin- ing@EMNLP. Association for Computational Linguistics, 22–31

  121. [132]

    Jian Yuan, Zhongyu Wei, Yixu Gao, Wei Chen, Yun Song, Donghua Zhao, Jinglei Ma, Zhen Hu, Shaokun Zou, Donghai Li, and Xuanjing Huang

  122. [133]

    Weizhe Yuan, Graham Neubig, and Pengfei Liu. 2021. BARTScore: Evaluating Generated Text as Text Generation. In NeurIPS. 27263–27277

  123. [134]

    Jingqing Zhang, Yao Zhao, Mohammad Saleh, and Peter J. Liu. 2020. PEGASUS: Pre-training with Extracted Gap-sentences for Abstractive Summarization. In ICML (Proceedings of Machine Learning Research, Vol. 119) . PMLR, 11328–11339

  124. [135]

    Weinberger, and Yoav Artzi

    Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger, and Yoav Artzi. 2020. BERTScore: Evaluating Text Generation with BERT. InICLR. OpenReview.net

  125. [136]

    Ashley, and Matthias Grabmair

    Linwu Zhong, Ziyi Zhong, Zinian Zhao, Siyuan Wang, Kevin D. Ashley, and Matthias Grabmair. 2019. Automatic Summarization of Legal Decisions using Iterative Masking of Predictive Sentences. In Proceedings of the Seventeenth International Conference on Artificial Intelligence an...

  126. [137]

    Yang Zhong and Diane J. Litman. 2022. Computing and Exploiting Document Structure to Improve Unsupervised Extractive Summarization of Legal Case Decisions. In NLLP@EMNLP. Association for Computational Linguistics, 322–337

  127. [138]

    Yang Zhong and Diane J. Litman. 2023. STRONG - Structure Controllable Legal Opinion Summary Generation. In IJCNLP (Findings). Association for Computational Linguistics, 431–448

  128. [139]

    May Myo Zin, Ha-Thanh Nguyen, Ken Satoh, Saku Sugawara, and Fumihito Nishino. 2023. Information Extraction from Lengthy Legal Contracts: Leveraging Query-Based Summarization and GPT-3.5. In JURIX (Frontiers in Artificial Intelligence and Applications, Vol. 379) . IOS Press, 17...

  129. [142]

    2024 DCESumm Identify and rank sentence relevance in a document using Legal BERT, refine scores with deep clustering, and select sorted sen- tences based on the desired summary length BillSum, Forum ROUGE No

  130. [143]

    2023 Italian-LEGAL- BERT Fine-tuned Italian-BERT, Italian-LEGAL- BERT, and Italian-LEGAL-BERT-SC models to predict the most relevant sentences ITA CaseHold ROUGE Yes

  131. [144]

    BillSum ROUGE Yes

    2023 BB25HLegalSum Leverage BERT and BM25 to rank and cluster unique, relevant sentences, aggregate them into candidate summaries, and present the most representative summary by highlight- ing the most representative sentences. BillSum ROUGE Yes

  132. [145]

    Landmark judgments attract pub- lic attention and gather numerous citations, highlighting arguments and precedents that reinforce the cited ruling

    2023 CB-JSumm Used InLegalBERT to obtain embeddings for citation and judgment sentences, then ap- plied cosine similarity to select summary sentences. Landmark judgments attract pub- lic attention and gather numerous citations, highlighting arguments and precedents that reinfo...

  133. [146]

    2023 GPT-3.5 Query-based summarization extracts key sentences relevant to predefined queries, which are then processed by GPT-3.5 for in- formation extraction CUAD F1 No

  134. [147]

    2023 Proposed method comprises a Sentence En- coder, Topic Model, Position Encoder, TS- LSTM Network, Law Article Processor, and Sentence Classifier CAIL ROUGE No

  135. [148]

    court opinions.ion U.S

    2023 Lawformer Utilized the Longformer encoder, which was pre-trained on the LegalBART objective us- ing 6 million U.S. court opinions.ion U.S. court opinions BillSum ROUGE No

  136. [149]

    2023 Compared BERT, RoBERTa, and XLNet for extractive summarization in the legal domain AILA ROUGE No

  137. [150]

    2023 Bayesian Opti- mization Score Fusion Current extractive and graph-based meth- ods such as TextRank, LSA, and KLSum have been enhanced with a technique for scoring significant sentences based on linguistic fea- tures BillSum, Gov- Report, FIRE, AILA ROUGE No

  138. [151]

    2022 Unsupervised Graph-based Ranking model HipoRank employs a reweighting algorithm that considers the history of previously se- lected sentences to iteratively update sen- tence importance scores and then select top- k candidates for extractive summary CanLII ROUGE, BERTScore Yes

  139. [152]

    2021 DELSumm An Integer Linear Programming (ILP) objec- tive maximizes summaries by selecting in- formative sentences, balancing thematic seg- ment representation, and minimizing redun- dancy, ensuring comprehensive and concise content Private ROUGE No

  140. [153]

    BillSum ROUGE No Table 7

    2021 CAWESumm (Contextual Anonymous Walk Embedding Summarizer) Legal-BERT sentence embeddings and Anonymous Walk Embeddings are con- catenated and input into an MLP model to learn binary classification of sentence summary-worthiness during training. BillSum ROUGE No Table 7. L...

  141. [154]

    BART is then fine- tuned for abstractive summarization using this augmented dataset BillSum, FIRE ROUGE, BERTScore No

    2024 Extract-then- Assign Extractive summaries are generated and matched to ground truth summaries, creat- ing new training samples. BART is then fine- tuned for abstractive summarization using this augmented dataset BillSum, FIRE ROUGE, BERTScore No

  142. [155]

    2024 A supervised BERTopic variant clusters con- tracts into seven groups, followed by Legal Pegasus, fine-tuned for the legal domain, to generate summaries for each document clus- ter Contract Un- derstanding Atticus Dataset (CUAD) F1 No

  143. [156]

    2023 A two-stage approach: keywords and key sentences are extracted using BERT+LSTM in the first stage, followed by abstractive sum- mary generation with UNILM and attention mechanisms in the second stage CAIL, LCRD ROUGE No

  144. [157]

    2023 STRONG (Struc- ture conTRollable OpiNion sum- mary Generation) A three-stage process: train a sentence struc- ture classifier on annotated data, predict sil- ver labels for unannotated summaries, and fine-tune the LED model using structure- guided tokens for test set infe...

  145. [158]

    2022 Two methods are proposed: a hybrid rein- forcement learning approach combining ex- tractive sentence selection with abstractive rewriting, and a transformer-based summa- rization method leveraging BART RechtspraakNL ROUGE Yes

  146. [159]

    2022 VanBART, FPT- BART Court decision summaries are generated based on a question-answer-decision triplet, designed to be intelligible for ordinary citi- zens without legal expertise JusticeBot, Can- LII ROUGE Yes

  147. [160]

    2021 BART Long documents are compressed by identi- fying salient sentences, which are then in- put into BART to generate abstractive sum- maries Amicus data ROUGE No

  148. [161]

    2021 ERG-GAT, ERG- GTN, TIG-GAT, TIG-GTN A dual-encoder model combines BERT with an Element Graph concept that encodes topic information, enhancing summarization by integrating structured topic representation LPO-news ROUGE Yes

  149. [162]

    BSLT uses BERTSUM-based Lawformer and Transformer for extraction, while LPGN combines PGN and Lawformer for accurate, high-quality abstractive summaries CAIL ROUGE No

    2020 Documents are segmented into three parts for summary generation. BSLT uses BERTSUM-based Lawformer and Transformer for extraction, while LPGN combines PGN and Lawformer for accurate, high-quality abstractive summaries CAIL ROUGE No

  150. [163]

    List of datasets, methods, and metrics related to abstractive legal document summarization

    2020 Abstractive summarization model with two encoders and one decoder, incorporating topic words to enhance performance, built on the Point-Generator Network (PTGEN) framework Private ROUGE No Table 8. List of datasets, methods, and metrics related to abstractive legal docume...

  151. [164]

    2024 Ensemble differ- ent summariza- tion algorithms Three ensemble methods are proposed: voting-based, ranked-list ensemble (using Borda count and Reciprocal Ranking), and graph-based (leveraging sentence similarity graphs to select identical sentences from con- nected compon...

  152. [165]

    2023 BSLT LPGN A hybrid legal summarization approach com- bines the extractive model BSLT and the ab- stractive model LPGN based on Lawformer to generate the summary CAIL2020 ROUGE No

  153. [166]

    2023 Extractive then abstractive A transfer learning approach combines ex- tractive and abstractive techniques for legal summarization, addressing limited labeled data by selecting sentences and generating summaries using GPT-2 Australian Le- gal Case Report Dataset ROUGE, FactCC No

  154. [167]

    2023 Extractive then abstractive Legal texts are normalized using dictionaries for abbreviations and article summaries, then processed with BART for extractive summa- rization and PEGASUS for abstractive sum- marization SCI, Indi- anKanoon, Manupatra, ILDC ROUGE No

  155. [168]

    2023 Extractive then abstractive Machine learning methods label sentences as important or not, followed by LSTMs for summarization, which are then compared to PEGASUS for performance evaluation Opinion of the supreme court of Utah, Idaho, Arizona, New Mexico, Nevada, and Color...

  156. [169]

    The resulting corpus is then fed into T5 PE- GASUS for abstractive summarization CAIL ROUGE No

    2023 Extractive then abstractive RoBERTa vectorizes document sentences, which are processed by a Dilated Gated CNN extractive model to select relevant sentences. The resulting corpus is then fed into T5 PE- GASUS for abstractive summarization CAIL ROUGE No

  157. [170]

    2022 Extractive then abstractive Ripple Down Rules classify sentences into 13 rhetorical roles using C4.5 decision trees, Naive Bayes, SVM, Conditional Random Fields, and BiLSTM algorithms for improved sentence categorization 50 documents of Bombay High Court collected from Le...

  158. [171]

    List of datasets, methods, and metrics related to hybrid legal document summarization

    2018 Ontology-based An ontology-based approach extracts seman- tic knowledge from Chinese legal documents, summarizes it into knowledge blocks, com- putes block similarity, and classifies the doc- uments into different categories CTA, CDD Accuracy No Table 9. List of datasets,...

  159. [2019]

    https://www.aclweb.org/anthology/W19-2201

    Association for Computational Linguistics, Minneapolis, Minnesota, 1–11. https://www.aclweb.org/anthology/W19-2201

  160. [2021]

    Data Intelligence 3, 2 (2021), 287–307

    Overview of SMP-CAIL2020-Argmine: The Interactive Argument-Pair Extraction in Judgement Document Challenge. Data Intelligence 3, 2 (2021), 287–307. doi:10.1162/dint_a_00094

Pith tools

Reviewed August 10, 2026 · model on record in the stance chip above.