Pith. sign in

REVIEW 3 cited by

How Ready are Pre-trained Abstractive Models and LLMs for Legal Case Judgement Summarization?

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2306.01248 v2 pith:46DQG42C submitted 2023-06-02 cs.CL cs.IRcs.LG

classification cs.CLcs.IRcs.LG
keywords summarizationmodelsabstractivecasesummariesllmspre-trainedgenerate
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Automatic summarization of legal case judgements has traditionally been attempted by using extractive summarization methods. However, in recent years, abstractive summarization models are gaining popularity since they can generate more natural and coherent summaries. Legal domain-specific pre-trained abstractive summarization models are now available. Moreover, general-domain pre-trained Large Language Models (LLMs), such as ChatGPT, are known to generate high-quality text and have the capacity for text summarization. Hence it is natural to ask if these models are ready for off-the-shelf application to automatically generate abstractive summaries for case judgements. To explore this question, we apply several state-of-the-art domain-specific abstractive summarization models and general-domain LLMs on Indian court case judgements, and check the quality of the generated summaries. In addition to standard metrics for summary quality, we check for inconsistencies and hallucinations in the summaries. We see that abstractive summarization models generally achieve slightly higher scores than extractive models in terms of standard summary evaluation metrics such as ROUGE and BLEU. However, we often find inconsistent or hallucinated information in the generated abstractive summaries. Overall, our investigation indicates that the pre-trained abstractive summarization models and LLMs are not yet ready for fully automatic deployment for case judgement summarization; rather a human-in-the-loop approach including manual checks for inconsistencies is more suitable at present.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. RELexED: Retrieval-Enhanced Legal Summarization with Exemplar Diversity

    cs.CL 2025-01 conditional novelty 5.0 of 10

    Diverse, influence-function-scored exemplar summaries retrieved with a DPP improve legal summarization over no-exemplar and similarity-only baselines on SuperSCOTUS and CivilSum, with modest and statistically partial gains.

  2. CoPERLex: Content Planning with Event-based Representations for Legal Case Summarization

    cs.CL 2025-01 conditional novelty 5.0 of 10

    An event-based planning pipeline with content selection improves faithfulness and coherence in legal case summarization across four datasets.

  3. A Comprehensive Survey on Legal Summarization: Challenges and Future Directions

    cs.CL 2025-01 conditional novelty 4.0 of 10

    A systematic survey of legal summarization finds a field dominated by English common-law datasets, ROUGE-based evaluation, and few human or expert validation studies.

Pith tools