Pith. sign in

REVIEW 2 cited by

Corpus for Automatic Structuring of Legal Documents

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2201.13125 v2 pith:EJSZFDF7 submitted 2022-01-31 cs.CL cs.AIcs.LG

classification cs.CLcs.AIcs.LG
keywords legalcorpusdocumentsrhetoricalrolesannotatedbaselineintroduce
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In populous countries, pending legal cases have been growing exponentially. There is a need for developing techniques for processing and organizing legal documents. In this paper, we introduce a new corpus for structuring legal documents. In particular, we introduce a corpus of legal judgment documents in English that are segmented into topical and coherent parts. Each of these parts is annotated with a label coming from a list of pre-defined Rhetorical Roles. We develop baseline models for automatically predicting rhetorical roles in a legal document based on the annotated corpus. Further, we show the application of rhetorical roles to improve performance on the tasks of summarization and legal judgment prediction. We release the corpus and baseline model code along with the paper.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. AILQA: Evaluating AI-Driven Legal Question Answering Systems for the Indian Legal System

    cs.CL 2026-07 conditional novelty 4.0 of 10

    RAG with top-3 chunk retrieval lifts smaller LLMs on Indian legal QA (Llama2-70B: 45.7% to 51.7% on AIBE) but often hurts large models, and under the study's own rating protocol some AI answers outscored the reference...

  2. A Data Science Approach to Calcutta High Court Judgments: An Efficient LLM and RAG-powered Framework for Summarization and Similar Cases Retrieval

    cs.IR 2025-06 reject novelty 4.0 of 10

    Fine-tuning Pegasus on LLM-annotated headnotes improves part of the legal summarization pipeline, and a RAG framework retrieves similar Calcutta High Court cases, though retrieval quality is never measured.

Pith tools