Pith. sign in

REVIEW 1 cited by

LePaRD: A Large-Scale Dataset of Judges Citing Precedents

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2311.09356 v3 pith:4MWMT7GL submitted 2023-11-15 cs.CL

classification cs.CL
keywords legalleparddatasetpassagepredictionretrievaltaskcontext
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We present the Legal Passage Retrieval Dataset LePaRD. LePaRD is a massive collection of U.S. federal judicial citations to precedent in context. The dataset aims to facilitate work on legal passage prediction, a challenging practice-oriented legal retrieval and reasoning task. Legal passage prediction seeks to predict relevant passages from precedential court decisions given the context of a legal argument. We extensively evaluate various retrieval approaches on LePaRD, and find that classification appears to work best. However, we note that legal precedent prediction is a difficult task, and there remains significant room for improvement. We hope that by publishing LePaRD, we will encourage others to engage with a legal NLP task that promises to help expand access to justice by reducing the burden associated with legal research. A subset of the LePaRD dataset is freely available and the whole dataset will be released upon publication.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Assessing the Performance Gap Between Lexical and Semantic Models for Information Retrieval With Formulaic Legal Language

    cs.CL 2025-06 conditional novelty 5.0 of 10

    On CJEU paragraph retrieval, BM25 beats off-the-shelf dense models on most metrics, fine-tuned dense models beat BM25, and BM25 wins mainly when queries have less verbatim overlap with the target.

Pith tools