{"id":"3ef71260-39e4-4936-bf01-c33b793d4547","arxiv_id":"2501.00097","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"CaseSumm, a 25.6K-pair dataset of Supreme Court opinions and official syllabuses, shows automated metrics favor fine-tuned Mistral while human experts prefer GPT-4, and LLM judges do not align with humans better than ROUGE.","lead":"CaseSumm pairs 25,600 U.S. Supreme Court opinions with the Court's official syllabuses, covering 1815 to 2019. The paper uses the dataset to show that automated summarization scores often disagree with trained human readers about which AI summaries are best, especially in high-stakes legal text.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central metric-vs-human discrepancy is not securely established: the human judgments come from 33 cases rated by ~11 law students, with no reported CIs or significance tests, and the Fig. 7 recruitment email contradicts the claimed source-blinding in §4.2.","rationale":"I read the paper in good faith as contributing a dataset and a comparative evaluation claim about legal summarization. The dataset contribution is well supported: 25.6K opinion/syllabus pairs, a 96% manual extraction check, public release, and a strict superset of Super-SCOTUS. The fine-tuning results are plausible. The make-or-break claim is the discrepancy between automatic and human evaluation, and the only direct evidence for it is the human study. That study is fragile for several concrete reasons: the sample is small, no uncertainty quantification is reported for the key pairwise comparison, the participants are law students rather than demonstrably expert legal summarizers, and the recruitment email in Figure 7 explicitly tells participants they are grading summaries from the authors' AI tool, undermining the claimed source-blinding in §4.2. The correlational evidence in Table 4 also lacks confidence intervals, so the G-Eval negative result is hard to calibrate. I do not claim the authors are dishonest; the Fig. 7 wording may be an oversight, but it matters because the human evaluation is the only support for the central discrepancy and for the claim that Mistral hallucinates more. I also note the internal inconsistency between 'roughly 20% of Mistral FT summaries have at least 1 factual error' and 'a total of 10 errors identified across all evaluations,' which suggests the error counts need verification. Because the dataset contribution and automatic metric findings would likely survive even if the human evaluation were re-run, I do not move the reader's verdict away from CONDITIONAL; the appropriate next step is a targeted blinded replication with proper confidence intervals.","tokens_in":18598,"tokens_out":6656,"duration_ms":64232,"concrete_test":"Re-run the human evaluation on the same 33 cases with fresh annotators recruited using neutral instructions that do not mention AI, the authors' tool, or the study purpose, and compute the mean rank difference between GPT-4 and Mistral FT with a 95% bootstrap CI and a paired significance test. If the CI includes zero, the central metric-human discrepancy is not robust; if GPT-4 still significantly outranks Mistral FT under neutral blinding, the concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's headline claim is that automatic metrics prefer Mistral FT while humans prefer GPT-4. The human side of that comparison is the load-bearing pillar. It rests on 57 opinion readings across 33 unique cases by second- and third-year law students (§4.2). The paper reports no confidence intervals, effect sizes, or pairwise significance tests for the GPT-4-vs-Mistral-FT rank differences in Figure 3, so we cannot tell whether the observed gap exceeds the minimum detectable effect of 0.52 rank points the authors themselves compute. In addition, the recruitment email in Figure 7 tells participants they will 'evaluate the quality of the summaries produced by our AI-based tool,' which directly conflicts with the §4.2 statement that 'Students were not told the source of each summary.' Knowing that at least one summary is AI-generated by the authors' tool can shift expectations about hallucinations and machine style. This does not necessarily break the GPT-4/Mistral pairwise comparison because both are AI outputs, but it does contaminate the claimed neutral expert evaluation and the comparisons against human-written controls that support the conclusion that GPT-4 outperforms human syllabuses. Table 4's correlations also lack confidence intervals; with n=33, the observed gaps between ROUGE and G-Eval are likely within sampling noise, so the G-Eval conclusion is underdetermined. There is also an internal inconsistency in §5.3: 'roughly 20% of Mistral FT summaries have at least 1 factual error' but 'a total of 10 errors identified across all evaluations'—with 57 readings those numbers cannot both hold, so the hallucination counts need audit. Furthermore, the participants are law students, while the abstract repeatedly calls them 'expert human annotators'; their representativeness as expert legal judgment is asserted rather than demonstrated.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces CaseSumm, a dataset of 25,642 U.S. Supreme Court opinion/syllabus pairs spanning 1815–2019, with the official Court syllabus serving as a gold-standard summary. The authors describe an extraction pipeline using PDFs from Public Resource Org and the Library of Congress, validate it on 100 randomly sampled cases (96 perfect extractions), and release the dataset publicly. They then benchmark GPT-4 Turbo, Mistral 7B (base and fine-tuned), and two human-written controls (Westlaw, Oyez) against the syllabuses using ROUGE, BERTScore, and human evaluation by second- and third-year law students. The central reported finding is a discrepancy: automatic metrics favor fine-tuned Mistral, while human evaluators most often rank GPT-4 higher, and G-Eval (an LLM-based metric) does not correlate better with human judgments than traditional automatic metrics. The paper also includes a hallucination error analysis with examples from both models.","tokens_in":18972,"tokens_out":13007,"duration_ms":121470,"significance":"The dataset contribution is substantial: CaseSumm is the largest open legal summarization dataset by document count, spans two centuries, and provides official syllabuses rather than crowd-sourced or machine-constructed summaries. The 96% manual extraction check and public release on HuggingFace are concrete strengths. If the evaluation findings hold, the paper would provide valuable evidence that automatic metrics and LLM-based judges diverge from human judgment in a high-stakes domain, which is an important result for summarization evaluation research. However, the human evaluation underpinning the headline discrepancy rests on a small sample (33 unique cases, 57 readings, ~11 students) and lacks inferential statistics, so the central comparative claims should be treated as preliminary rather than definitive.","major_comments":[{"comment":"The central claim that expert humans most commonly rank GPT-4 over Mistral FT is not statistically supported. The human evaluation consists of 57 opinion readings across 33 unique cases by second- and third-year law students (median 5 readings per student). No confidence intervals, effect sizes, or pairwise significance tests are reported for the rank differences in Figure 3, so the observed GPT-4-vs-Mistral-FT gap cannot be distinguished from sampling noise; the authors' own minimum detectable effect of 0.52 rank points is computed but never applied to the actual results. In addition, the recruitment email in Figure 7 states that participants will \"evaluate the quality of the summaries produced by our AI-based tool,\" which is in tension with the §4.2 statement that students were not told the source of each summary. This does not necessarily break the GPT-4/Mistral pairwise comparison, but it undermines the claimed neutrality of the evaluation and the comparisons against human-written controls (e.g., the conclusion that GPT-4 outperforms official syllabuses on some dimensions). Please report raw rank distributions, effect sizes with confidence intervals, and clarify exactly what participants were told about the provenance of the summaries.","section":"§4.2, §5.3, Figure 3"},{"comment":"The conclusion that \"LLM-based evaluation does not correlate with human judgments better than traditional automatic metrics\" is too strong for the evidence presented. Table 4 shows that G-Eval correlates much more strongly with human specificity judgments (adapted G-Eval specificity 0.51 vs. ROUGE-1 -0.28) and clarity (0.23–0.28 vs. ~0.0 for ROUGE-2), which the paper's own bullet points acknowledge. No confidence intervals or tests for differences between dependent correlations are reported; with 33 unique cases (or 57 opinion readings), differences such as ROUGE-1 sensitivity 0.54 vs. adapted G-Eval sensitivity 0.43 are likely within sampling noise. The abstract's blanket claim should be qualified to state that traditional metrics perform comparably or better on sensitivity/style, while G-Eval performs better on specificity/clarity, and the correlation comparisons should include uncertainty quantification.","section":"§5.3, Table 4, Abstract"}],"minor_comments":[{"comment":"GPT-4's prompt was optimized with DSPy using ROUGE-2 as the optimization metric, and the same ROUGE family is later used as an evaluation metric for GPT-4. This means GPT-4's ROUGE scores are partially fitted to the evaluation metric; the paper should acknowledge this when interpreting the automatic evaluation comparison, even though the prompt optimization likely biases in GPT-4's favor.","section":"§4.1, Appendix A.1"},{"comment":"The row for \"GovReport*\" has an asterisk but no footnote or citation identifying the dataset source; please add the reference.","section":"Table 1"},{"comment":"The statement \"doubling opinion length increases syllabus length by nearly 2/3\" appears to be based on an unreported regression; the correlation of 0.676 reported in Table 3 does not by itself imply this elasticity. Please report the underlying regression or rephrase the claim.","section":"§5.2"},{"comment":"In the reference for Kapoor et al., \"Peter Henderon\" should be \"Peter Henderson.\"","section":"References"},{"comment":"The sentence \"For all measures where the outcome is rank, we mark the mean rank identically 3) with a red dashed line\" is garbled and should be rewritten (e.g., \"we mark the mean rank, 3, with a red dashed line\").","section":"Appendix C.1"},{"comment":"The evaluators are described as \"expert human annotators\" in the abstract and as \"expert humans\" in §5, but the participants are second- and third-year law students. Please either adjust the terminology to \"law student evaluators\" or justify the \"expert\" label with information about their legal training and experience.","section":"Abstract, §4.2"}],"recommendation":"major_revision","confidential_remarks":"I want to flag two things for the editor. First, the dataset itself is a solid contribution and likely publishable after revision; the main risk is that the headline evaluation result (automatic-vs-human discrepancy) is built on a small human study with no inferential statistics. Second, the paper's abstract and conclusion make broader claims about LLM-based evaluation than Table 4 supports; the overgeneralization should be corrected regardless of whether the authors add significance tests. The \"expert\" label for law students also deserves editorial attention."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"CaseSumm is the real thing on the dataset side: 25.6K SCOTUS opinion/syllabus pairs spanning 1815-2019, a strict superset of Super-SCOTUS, with a documented extraction pipeline and a 96% manual accuracy check on 100 cases. Public release under CC BY-NC makes it a resource people will actually use. That part deserves credit.\n\nThe evaluation section is where I get cautious. The headline claim is that automatic metrics favor Mistral FT while expert humans rank GPT-4 higher, and that G-Eval doesn't correlate better with humans than ROUGE/BERTScore. The human side rests on 57 opinion readings from 33 unique cases by second- and third-year law students. That's small, and the paper doesn't report confidence intervals, effect sizes, or pairwise significance tests for the rank differences in Figure 3. The authors even compute a minimum detectable effect of 0.52 rank points, but we can't tell whether the observed GPT-4-vs-Mistral gap exceeds it. Table 4's correlations also lack CIs; with n=33, the G-Eval vs ROUGE differences are plausibly noise.\n\nThere's also a blinding problem. Section 4.2 says students were not told the source of each summary, but the recruitment email in Figure 7 tells them they will 'evaluate the quality of the summaries produced by our AI-based tool.' That doesn't necessarily break the GPT-4 vs Mistral FT pairwise comparison, but it does contaminate the comparisons against human-written syllabuses, which matter for the claim that GPT-4 beats official syllabuses.\n\nOne internal inconsistency: the text says roughly 20% of Mistral FT summaries have at least 1 factual error, with a total of 10 errors identified across all evaluations. With 57 readings those numbers can't both hold, so the error counts need an audit.\n\nCalling second- and third-year law students 'expert human annotators' in the abstract is a stretch. They are legal trainees, not experienced summarizers. That overstates the evidence.\n\nThe dataset itself is solid, and the paper is honest about limitations. The evaluation is a useful pilot but currently underpowered for the strength of the claims. I'd send this to review, expect major revisions on the human evaluation reporting, and cite it for the dataset.","headline":"A genuinely useful legal summarization dataset, paired with an evaluation section whose headline human-vs-automatic discrepancy is not yet supported by the reported statistics.","tokens_in":19537,"tokens_out":2148,"would_cite":true,"duration_ms":18376,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A 25,600-case Supreme Court dataset shows automatic metrics and expert human rankings disagree on which AI summary is best.","keywords":["legal summarization","long-context summarization","Supreme Court","dataset","human evaluation","hallucination","ROUGE","G-Eval"],"falsifier":"Re-run the human evaluation with a cohort of practicing attorneys who are told only that the summaries 'may be written by humans or machines,' and compare their rankings to the law-student rankings from the paper; if GPT-4 no longer outranks Mistral FT and the official syllabuses, the claimed human-versus-automatic discrepancy does not generalize.","tokens_in":18328,"feed_emoji":"⚖️","tokens_out":5465,"duration_ms":48800,"temperature":0.7,"pith_summary":"CaseSumm pairs 25,600 U.S. Supreme Court opinions with their official syllabuses, the Court's own summaries written by an attorney and approved by the Justices, spanning 1815-2019. Using this resource, the paper claims that standard automatic summarization metrics and LLM-based judges systematically disagree with expert human judgment about what makes a good legal summary. Fine-tuned Mistral 7B wins on ROUGE and BERTScore, but law-student readers rank GPT-4 summaries as clearer, more sensitive, and more specific, and flag Mistral as the most hallucination-prone. The paper also reports that G-Eval, a GPT-4-based evaluation, does not correlate with human rankings better than traditional metrics, and that GPT-4 summaries often outrank human-written ones including the official syllabuses except on factual correctness. If correct, the lesson is that legal summarization evaluation cannot rely on lexical or LLM scores alone, and the very notion that human-written ground-truth summaries are inherently superior is questionable.","feed_headline":"Human experts overrule automatic metrics in AI legal summaries","feed_subtitle":"A 25.6K-case Supreme Court benchmark shows auto scores and expert readers pick different winners.","key_machinery":"The central object is the dataset itself: pairs of Supreme Court majority opinions and their official summaries, or syllabuses, which serve as gold-standard references because they are written by a Court-employed attorney and approved by the Justices. The argument's mechanism is a side-by-side evaluation in which five candidate summaries (Mistral Base, Mistral FT, GPT-4 Turbo, and human-written Oyez and Westlaw summaries) are scored against the official syllabus with ROUGE, BERTScore, and G-Eval, and are also ranked by law students who read the original opinion on sensitivity, specificity, clarity, style, and factual error. Comparing the rankings produced by these different methods is what exposes the discrepancy between automatic and human judgment.","core_discovery":"The paper's central comparative claim is that evaluation method changes the winner: automatic metrics (ROUGE, BERTScore) and even LLM-based G-Eval disagree with expert human rankings when judging LLM-generated legal summaries. In the 622-case automatic evaluation, fine-tuned Mistral 7B leads on recall and F1, and its summaries are closest in length and compression to official syllabuses; but in a human evaluation where second- and third-year law students ranked five candidate summaries against the source opinion, GPT-4 most often ranks first on sensitivity, specificity, clarity, and style, while Mistral FT roughly matches the official syllabus on those dimensions yet is the only candidate with a conspicuous error rate (about 20% of Mistral FT summaries contain at least one factual error). The paper further finds that G-Eval, whether used with its default prompts or prompts adapted to the human rubric, does not correlate with human rankings better than traditional automatic metrics, and that GPT-4 summaries often outperformed human-written summaries including official syllabuses in several quality dimensions except factual correctness.","pith_inferences":["The recruitment email shown in Figure 7 tells participants they are grading summaries from the authors' AI tool, so it is untested whether the reported preference for GPT-4 would survive a fully blinded protocol; that is an inference because the paper claims participants were not told the source of each summary.","A practical consequence is that legal research tools optimized toward ROUGE-like scores may be rewarded for stylistic mimicry while still producing citation errors and fact misrepresentations that require separate verification.","Because the human evaluation rests on only 33 unique cases read by a median of 5 cases per student, the effect sizes are uncertain; a larger replication with practicing attorneys could change the ranking even if the direction holds.","The dataset's temporal structure could support tests of whether modern LLMs reproduce the historical drift in syllabus style, such as the emergence of the 'Held:' section, although the paper does not run those tests."],"forward_implications":["Legal summarization evaluation should include human reading; automatic metrics alone can pick a winner that human experts reject.","LLM-based judges such as G-Eval do not yet offer a better reference-free alternative to traditional lexical and semantic metrics.","Fine-tuning an open 7B model can match or beat much larger models on lexical overlap, but it can also increase the risk of confident factual hallucinations.","Human-written summaries, including official syllabuses and paid commercial services, are not inherently better than strong LLM summaries on clarity or style, though they remain more factually reliable.","The 200-year span of the dataset makes it possible to study how summary length, compression, and responsiveness to source length have changed over time."],"supporting_citations":[{"why":"Defines ROUGE, the primary lexical automatic metric used for evaluation.","marker":"Lin, 2004"},{"why":"Defines BERTScore, the semantic automatic metric used alongside ROUGE.","marker":"Zhang et al., 2019"},{"why":"Provides G-Eval, the LLM-based evaluator tested against human rankings.","marker":"Liu et al., 2023"},{"why":"Supplies a portion of the opinion data and the Oyez control summaries.","marker":"Fang et al., 2023"},{"why":"Introduces Mistral 7B, the open-source model that is fine-tuned and evaluated.","marker":"Jiang et al., 2023"},{"why":"Documents GPT-4 Turbo, the proprietary model benchmarked in the study.","marker":"GPT-4 Team, 2024"},{"why":"Defines BARTScore, which the paper tests and then excludes from the main analysis.","marker":"Yuan et al., 2021"},{"why":"Provides the U.S. Reports archive from which most opinions are drawn.","marker":"Public Resource Org, 2024"}],"fun_headline_variants":["Evaluation method flips winner on AI legal summaries","Human experts catch what auto metrics miss in AI legal summaries","Auto metrics pick Mistral, humans pick GPT-4 for legal summaries","In AI legal summaries, humans rank GPT-4 higher than auto metrics do","Human review overrules auto scores for AI legal summaries"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The human rankings that drive the main discrepancy come from 33 unique cases read by second- and third-year law students, and the recruitment email told them they were grading summaries from the authors' AI tool, so the load-bearing premise is that these rankings are a valid, unbiased proxy for expert legal judgment.","fun_headline_variants_meta":{"raw":{"variants":["Evaluation method flips winner on AI legal summaries","Human experts catch what auto metrics miss in AI legal summaries","Auto metrics pick Mistral, humans pick GPT-4 for legal summaries","In AI legal summaries, humans rank GPT-4 higher than auto metrics do","Human review overrules auto scores for AI legal summaries"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000227,"raw_usage":{"total_tokens":1517,"prompt_tokens":1033,"completion_tokens":484,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":649,"completion_tokens_details":{"reasoning_tokens":399}},"tokens_in":649,"tokens_out":484,"duration_ms":4865,"temperature":1.0,"reasoning_tokens":399,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T22:59:48.272607+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the human evaluation with a cohort of practicing attorneys who are told only that the summaries 'may be written by humans or machines,' and compare their rankings to the law-student rankings from the paper; if GPT-4 no longer outranks Mistral FT and the official syllabuses, the claimed human-versus-automatic discrepancy does not generalize.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the U.S. Reports archive from which most opinions are drawn."}],"review_version":1}