{"id":"83a35ba1-d39b-4aed-bf7a-46ab904e79b4","arxiv_id":"2501.17830","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A systematic survey of legal summarization finds a field dominated by English common-law datasets, ROUGE-based evaluation, and few human or expert validation studies.","lead":"This paper reviews 123 studies on automatic legal text summarization published since 2017, cataloging datasets, methods, evaluation metrics, and research gaps. It is a reference map for researchers and practitioners working on legal NLP summarization.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 'comprehensive' claim rests on a non-reproducible sampling frame: the reported query contains a mandatory ('of' OR 'about') term and the 123-study list is never published, so recall and the resulting trend statistics cannot be audited.","rationale":"The reader's CONDITIONAL verdict is appropriate. My stress-test identifies the same weakest point: the survey's value depends on the completeness of the 123-paper sample, and that sample is not auditable. I add a concrete internal symptom, namely that the Section 2 query contains a mandatory 'of'/'about' term and an incomplete set of domain content words, which makes the reported retrieval process impossible to reproduce as stated, regardless of whether it is interpreted as metadata search or full-text search. Without an included-study list, the percentages in Section 7 and the gap analysis in Section 8 cannot be checked. I do not see grounds for REJECT: the taxonomy, dataset summaries, and discussion are useful, and the flaws are fixable by releasing the study list and a precise query protocol. I also credit the authors for providing a recognizable systematic-review skeleton and for discussing limitations such as ROUGE's weaknesses and the small human-evaluation samples. My concern does not change the verdict; it sharpens the condition under which the survey should be relied on. The factual dataset misclassifications noted by the reader, such as Multi-LexSum being described as European in Section 4, reinforce the need for a corrected, auditable appendix, but they are not the single most load-bearing issue.","tokens_in":28636,"tokens_out":7878,"duration_ms":81281,"concrete_test":"Ask the authors to release the full list of 123 included studies with source repository, retrieval date, and screening decision. Independently build a gold-standard set by taking every legal-summarization paper from NLLP 2019-2024, ICAIL 2017-2024, JURIX 2017-2023, plus papers cited in prior surveys [56, 63], and keeping only items meeting Table 3's inclusion criteria. Then compute recall of the 123-study list against this gold standard and recompute the ROUGE-usage and human-evaluation percentages on the union. If recall is below roughly 90 percent, or if the added papers shift either percentage by more than 5 percentage points, the trend claims should be revised to describe the retrieved sample rather than the field.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 2's reported query, ('legal' OR 'law' OR 'legislative') AND ('summarization' OR 'summary') AND ('of' OR 'about') AND ('dialogues' OR 'question and answering' OR 'text' OR 'textual' OR 'text-based' OR 'document'), is not a coherent retrieval string. If applied to title or abstract metadata, the mandatory ('of' OR 'about') connector excludes titles such as 'Legal Case Document Summarization: Extractive and Abstractive Methods and their Evaluation', which is the survey's own reference [117]. If applied to full text, the connector is vacuous, and the content-term disjunction omits high-yield domain vocabulary such as 'court', 'judgment', 'ruling', 'case', and 'opinion'; either way, the 262 initial hits cannot be reconstructed. The paper also gives no list of the 123 included studies, no per-repository hit counts, and no flow diagram, so a reader cannot check whether the sample is the field or a convenience subset. The headline statistics (Section 7.1: more than 95 percent ROUGE usage; Section 7.2: 20 percent human evaluation, 9 percent expert evaluation) depend on this unverifiable denominator. The concern is not about author intent; it is that the central claim of comprehensiveness requires an audit trail that the manuscript does not provide.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper presents a systematic survey of legal document summarization research from 2017 to 2024. The authors describe a paper selection methodology following Kitchenham and Charters, review 123 papers, and organize the field along several axes: datasets (Section 3), regional specifics (Section 4), extractive/abstractive/hybrid strategies (Section 5), methodological families (Section 6), and evaluation metrics including automated and human evaluation (Section 7). The paper concludes with challenges and future directions (Section 8). The central claim is that this is a comprehensive, up-to-date, and reproducible map of the legal summarization field, supported by statistics such as over 95% ROUGE usage and 20% human-evaluation adoption.","tokens_in":28919,"tokens_out":3593,"duration_ms":32175,"significance":"If the comprehensiveness and reproducibility claims are substantiated, this would be the current reference survey for legal summarization, filling a genuine gap between older surveys and the transformer-era literature. The taxonomy of strategies, the dataset inventory, and the structured account of human-evaluation dimensions (Section 7.2, Figure 5, Table 5) are potentially valuable for both newcomers and practitioners. The headline statistics about metric usage and evaluation practices would be useful field-level facts. However, the survey's utility is strongly coupled to the auditability of its 123-paper sample and the correctness of its regional and dataset classifications; both currently have load-bearing issues that need to be resolved before the survey can be relied upon.","major_comments":[{"comment":"The reported search query is not a reproducible retrieval string. As written, (\"legal\" OR \"law\" OR \"legislative\") AND (\"summarization\" OR \"summary\") AND (\"of\" OR \"about\") AND (\"dialogues\" OR \"question and answering\" OR \"text\" OR \"textual\" OR \"text-based\" OR \"document\") would, if applied to titles and abstracts, exclude the paper's own reference [117] (\"Legal Case Document Summarization: Extractive and Abstractive Methods and their Evaluation\") because that title contains neither \"of\" nor \"about\". If applied to full text, the mandatory connector is vacuous, and the content-term disjunction omits high-yield terms such as \"court\", \"judgment\", \"ruling\", and \"case\". The manuscript reports 262 initial hits and 123 final papers but provides no per-repository hit counts, no flow diagram, and no list of the 123 included papers, so the denominator underlying the headline statistics in Sections 7.1 and 7.2 (95% ROUGE, 20% human evaluation, 9% expert evaluation) cannot be audited. Because the central claim is comprehensiveness, this is a load-bearing reproducibility gap.","section":"Section 2 (Paper selection methodology), Table 2"},{"comment":"The paragraph on \"European specifics\" states that \"Multi-LexSum released in [115] is another dataset pertaining to the European landscape\" and describes it as containing expert summaries based on CRLC publications. This directly contradicts Section 3.1, where Multi-LexSum is correctly described as a dataset of 40,000 source documents from roughly 4,500 federal U.S. civil rights lawsuits sourced from the Civil Rights Litigation Clearinghouse. The misclassification affects the regional analysis, including Figure 2's country distribution, and needs to be corrected.","section":"Section 4 (European specifics)"},{"comment":"Two dataset entries in Table 4 conflict with the running text. First, KorCase_Summ is listed under the domain \"Privacy Policy\", although Section 4 (Asia) describes it as containing precedents of the Korean Court, and the provided URL points to the Korean legal information portal law.go.kr. Second, CLSum-HK is listed with a size of 793 documents, while Section 3.1 states that \"CLSum-HK contains 233 judgments and their respective press summaries from the legal reference system of Hong Kong\". These inconsistencies undermine the reliability of the dataset inventory.","section":"Table 4 (Dataset overview)"},{"comment":"The sentence \"Authors of [33] fuse topic vectors into an LSTM for improving its ability to extract legal text features\" misattributes the method. Reference [33] is Dong et al. (2019), \"Unified Language Model Pre-training for Natural Language Understanding and Generation\" (UNILM), which is a transformer pretraining paper and does not fuse topic vectors into an LSTM. The topic-vector LSTM method appears to be described elsewhere in the literature; this citation error needs to be corrected, as it directly affects the methodological taxonomy in Section 6.","section":"Section 6 (Methods for legal summarization), paragraph 'Transformers combined with LSTMs or CNNs'"}],"minor_comments":[{"comment":"The inclusion criteria bullets contain a grammar and clarity issue: \"Papers with datasets and metrics, we include paper before 2017\" should be rephrased to state the intended exception for pre-2017 dataset/metric papers unambiguously.","section":"Section 2 (Paper selection methodology)"},{"comment":"There is a typo in the Multi-LexSum description: \"muti-document summaries\" should be \"multi-document summaries\".","section":"Section 3.1 (Court rulings)"},{"comment":"The heading \"Evalaution metrics for legal summarization\" contains a typo and should read \"Evaluation metrics for legal summarization\".","section":"Section 7 (Heading)"},{"comment":"The dataset name \"IN-abs\" in the hybrid-methods table is inconsistently capitalized relative to the \"IN-Abs\" and \"IN-abs\" usages elsewhere; standardize the capitalization.","section":"Table 9 (Appendix A.2)"},{"comment":"The metric name \"FActCC\" appears to be a typo for \"FactCC\" (Kryscinski et al., 2020); please verify and correct.","section":"Section 7.1 (Automated evaluation metrics)"},{"comment":"Multi-LexSum is cited three times with overlapping content: [114] and [115] are both the 2022 arXiv version, and [116] is the NeurIPS 2022 version. These should be merged or clearly distinguished to avoid duplicate citations.","section":"References [114], [115], [116]"}],"recommendation":"major_revision","confidential_remarks":"The paper is potentially a valuable reference survey, but the reproducibility issues in Section 2 are fundamental to the 'comprehensive' claim and must be addressed with a working query, per-repository counts, a flow diagram, and the full list of included papers. The factual errors in Sections 3, 4, and 6 are easily correctable but should be checked systematically against the 123-paper corpus before resubmission. I would not consider acceptance until the audit trail is supplied."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know two things about this one. First, it is the best current map of transformer-era legal summarization that I have seen: 123 papers organized by region, strategy, method, and evaluation, with a usable dataset inventory and some sensible trend statistics. Second, do not trust the numbers or the table entries without checking them. The survey's central claim to comprehensiveness is real but not verifiable as written, and a few descriptive errors mar what is otherwise a solid synthesis.\n\nWhat the paper does well: it collects and organizes a scattered literature, gives a reasonable taxonomy (extractive, abstractive, hybrid; rule-based, rank-based, model-centric), and documents a real field-level problem—roughly 95% ROUGE usage, only 20% human evaluation, and 9% expert evaluation. The dataset tables, despite the errors below, are a genuinely useful starting point. The authors also cite and position themselves against prior surveys honestly; this is a legitimate extension, not a rehash.\n\nThe soft spots are real. Section 2's search query is not a coherent retrieval string. The mandatory ('of' OR 'about') connector is either vacuous in full-text search or excludes titles like 'Legal Case Document Summarization' in metadata search, and the paper provides no list of the 123 included studies, no per-repository hit counts, and no flow diagram. So the headline statistics cannot be audited. That is a load-bearing issue for a survey whose selling point is comprehensiveness, though not a fatal one. Then there are concrete factual errors. Section 4 places Multi-LexSum under 'European specifics' even though it is a U.S. civil rights dataset. Table 4 lists KorCase_Summ under 'Privacy Policy' when it contains Korean court precedents. Section 6 attributes a topic-vector LSTM method to reference [33], which is the UNILM paper; that attribution looks wrong. Each of these is fixable, but together they tell you the screening and fact-checking were not done at the level the paper promises.\n\nThe reader's conditional verdict is the right one. This is not a desk-reject; it is a survey that a serious editor should send to referees, with a request that the authors publish the full included-study list, repair the query description, and correct the dataset/attribution errors. If those revisions happen, I would cite it and point students to it. As it stands, I would read it with a pen in hand.","headline":"A genuinely useful survey of transformer-era legal summarization, but the 'comprehensive' claim rests on a retrieval procedure that cannot be audited and the text contains several factual slips that a careful referee would catch.","tokens_in":29386,"tokens_out":1543,"would_cite":true,"duration_ms":17766,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A systematic survey of 123 legal-summarization papers maps the field and finds research concentrated in a few English/common-law jurisdictions, dominated by ROUGE, and rarely validated by human or expert evaluation.","keywords":["legal summarization","systematic literature review","legal NLP","transformer era","extractive summarization","abstractive summarization","evaluation metrics","legal datasets"],"falsifier":"Re-run the stated query across an independent set of bibliographic databases and add backward and forward citation chasing, then count peer-reviewed legal-summarization papers from 2017 to 2024 that meet the survey's inclusion criteria but are absent from its 123-paper corpus; if the missed set is large enough to shift the reported figures (over 95 percent ROUGE usage, about 20 percent human evaluation), the survey's field-level conclusions are not robust.","tokens_in":28472,"feed_emoji":"⚖️","tokens_out":13080,"duration_ms":111701,"temperature":0.7,"pith_summary":"This paper attempts to establish a systematic, up-to-date map of automatic legal summarization: the methods, datasets, and evaluation metrics used across 123 papers published during the transformer era (2017-2024). Its headline statistics are field-level: over 95 percent of method papers report ROUGE, roughly 20 percent add any human evaluation, about 9 percent involve legal experts, and every dataset has a single reference summary. If the survey's coverage is complete, these numbers describe the legal-summarization field as a whole, not a sample. That matters because legal summaries are high-stakes outputs, and the survey's gap analysis points to concrete fixes: multi-reference and user-personalized datasets, domain-specific embeddings, better handling of long documents, and community-built benchmarks for evaluating the evaluators.","feed_headline":"123 legal-summarization papers mapped in one systematic survey","feed_subtitle":"ROUGE dominates, human evaluation is rare, and most work centers on a few jurisdictions.","key_machinery":"The machinery is the systematic review protocol: a fixed search query over eight primary and four secondary digital libraries, explicit inclusion/exclusion criteria, two screening stages (titles and abstracts, then full text) with agreement among the first three authors, and a final corpus of 123 papers. On that corpus the survey builds a three-part taxonomy for organising the literature — region (Section 4), strategy (extractive, abstractive, hybrid; Section 5), and method (Section 6) — which generates the field-level statistics. The evaluation discussion is organised by a set of metric families (lexical-overlap, embedding-based, factuality, ranking, classification) and by a five-characteristic human-evaluation framework (content coverage, content representation, efficiency, language quality, summary impact); these are the lenses through which the survey reaches its gap analysis.","core_discovery":"The central claim, stated on the paper's own terms, is that the legal-summarization work of the transformer era can be systematically organised and, once organised, reveals a recognizable state of the field. Extractive, abstractive, and hybrid strategies are built mostly on BERT-family encoders and BART-, Pegasus-, and T5-style decoders; datasets such as the US legislative corpus, Indian and UK Supreme Court corpora, Canadian case-law records, and the multilingual EU legislative corpus supply most of the training data; and evaluation is dominated by lexical-overlap metrics, with ROUGE used in more than 95 percent of method papers while human evaluation appears in only about 20 percent, of which about 9 percent use legal experts. The survey also claims that research is concentrated in a few English/common-law jurisdictions, that legal-summarization datasets all carry exactly one reference summary, and that no community benchmark exists for meta-evaluating metrics or human-evaluation protocols in the legal domain. These findings are presented as the basis for a future research agenda: user-specific and multi-reference ground truths, domain-specific embeddings, interpretable handling of lengthy interdependent documents, multimodal and dialogue-aware inputs, and community-built meta-evaluation datasets.","pith_inferences":["Extension: the paper counts metrics per paper; a complementary analysis that weights by dataset would likely sharpen the concentration, since a small number of datasets (the US legislative corpus alone appears in 14 papers) drives much of the observed trend.","Extension: the human-evaluation dimensions in the survey could be turned into a testable protocol: have legal experts score a sample of generated summaries on those dimensions and correlate the scores with automatic metrics; the resulting correlation table would tell practitioners which metric to trust when experts are unavailable.","Extension: because the search window ends in July 2024, the survey captures only the first wave of instruction-tuned large language models; re-running the same protocol with a later cutoff could test whether ROUGE dominance and the human-evaluation gap are shrinking or persisting.","Implication the authors do not draw: a leaderboard built only on the existing English/common-law datasets would reward models that fit those corpora rather than models that generalize, so a cross-jurisdiction held-out benchmark is needed before the field can claim transferable progress."],"forward_implications":["If the coverage is complete, any claimed progress in legal summarization is currently progress measured mostly by ROUGE, so new results should be re-examined against the metric's known insensitivity to legal content.","The regional concentration implies that cross-jurisdiction transfer is not yet reliable; work on non-English and civil-law legal systems remains underserved and cannot be assumed to inherit the existing results.","With only about 20 percent of papers using human evaluation, and even fewer involving legal experts, most quality claims in the literature stand on automatic scores alone, making a community benchmark for evaluation metrics a prerequisite for trustworthy comparisons.","The single-reference design of all current datasets rules out studying consensus summaries, so building multi-reference legal corpora would open a new line of evaluation research.","The survey's limitation analysis identifies the same document types again and again, leaving legal dialogues, meetings, transcripts, and multimodal or low-resource settings as open ground for future work."],"supporting_citations":[{"why":"Supplies the systematic literature review guidelines whose steps the paper's selection process follows.","marker":"[66]"},{"why":"Earlier legal summarization survey the paper positions itself against when claiming the field lacks a timely comprehensive survey.","marker":"[63]"},{"why":"Earlier overview of legal summarization used to justify the need for an updated systematic treatment.","marker":"[56]"},{"why":"Introduces the Indian and UK Supreme Court datasets and the extractive/abstractive methods used throughout the survey's taxonomy.","marker":"[117]"},{"why":"BillSum, the US legislative dataset the survey counts as the most-used resource and uses for its regional and method statistics.","marker":"[70]"},{"why":"Defines ROUGE, the metric behind the survey's statistic that more than 95 percent of method papers report it.","marker":"[78]"},{"why":"Empirical study questioning whether ROUGE captures legal content, underpinning the survey's evaluation critique.","marker":"[120]"},{"why":"General-domain meta-evaluation benchmark cited as the template for the legal meta-evaluation datasets the survey calls for.","marker":"[39]"}],"fun_headline_variants":["ROUGE rules legal summaries, but humans rarely check","Survey: 120+ papers, ROUGE dominates, human eval scarce","One survey, 120+ papers: the state of legal text summarization","Legal summarization: 120+ papers, few human checks, much work ahead","Legal summary survey: 95% ROUGE, only 9% with legal experts"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing assumption is that the repository search, the keyword query, and the manual screening actually recovered all or nearly all relevant legal-summarization papers from 2017 to 2024; if a substantial body of work was missed or screened inconsistently, the survey's trend statistics describe a convenience sample rather than the whole field.","fun_headline_variants_meta":{"raw":{"variants":["ROUGE rules legal summaries, but humans rarely check","Survey: 120+ papers, ROUGE dominates, human eval scarce","One survey, 120+ papers: the state of legal text summarization","Legal summarization: 120+ papers, few human checks, much work ahead","Legal summary survey: 95% ROUGE, only 9% with legal experts"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00083,"raw_usage":{"total_tokens":3575,"prompt_tokens":843,"completion_tokens":2732,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":459,"completion_tokens_details":{"reasoning_tokens":2631}},"tokens_in":459,"tokens_out":2732,"duration_ms":19197,"temperature":1.0,"reasoning_tokens":2631,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T04:31:16.901275+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the stated query across an independent set of bibliographic databases and add backward and forward citation chasing, then count peer-reviewed legal-summarization papers from 2017 to 2024 that meet the survey's inclusion criteria but are absent from its 123-paper corpus; if the missed set is large enough to shift the reported figures (over 95 percent ROUGE usage, about 20 percent human evaluation), the survey's field-level conclusions are not robust.","supporting_citations":[],"review_version":1}