Pith. sign in

REVIEW 3 cited by

Yes, BM25 is a Strong Baseline for Legal Case Retrieval

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2105.05686 v2 pith:T6PJ2U22 submitted 2021-04-26 cs.IR cs.CLcs.LG

classification cs.IRcs.CLcs.LG
keywords bm25colieeaboveavailablebaselinecasecodedescribe
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We describe our single submission to task 1 of COLIEE 2021. Our vanilla BM25 got second place, well above the median of submissions. Code is available at https://github.com/neuralmind-ai/coliee.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. BriefMe: A Legal NLP Benchmark for Assisting with Legal Briefs

    cs.CL 2025-06 conditional novelty 6.0 of 10

    BriefMe introduces a legal brief benchmark with argument summarization, argument completion, and case retrieval, and shows LLMs beat human headings on the first two but struggle on the latter two.

  2. Assessing the Performance Gap Between Lexical and Semantic Models for Information Retrieval With Formulaic Legal Language

    cs.CL 2025-06 conditional novelty 5.0 of 10

    On CJEU paragraph retrieval, BM25 beats off-the-shelf dense models on most metrics, fine-tuned dense models beat BM25, and BM25 wins mainly when queries have less verbatim overlap with the target.

  3. ASP2LJ : An Adversarial Self-Play Laywer Augmented Legal Judgment Framework

    cs.CL 2025-06 conditional novelty 5.0 of 10

    ASP2LJ combines synthetic case generation with adversarial self-play for lawyer agents, improving legal judgment prediction on a Chinese benchmark and on a new rare-case dataset.

Pith tools