Pith. sign in

REVIEW 13 cited by

LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2308.11462 v1 pith:T4IFBEK3 submitted 2023-08-20 cs.CL cs.AIcs.CY

classification cs.CLcs.AIcs.CY
keywords legalreasoninglegalbenchllmstaskstypesbenchmarkbuilt
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The advent of large language models (LLMs) and their adoption by the legal community has given rise to the question: what types of legal reasoning can LLMs perform? To enable greater study of this question, we present LegalBench: a collaboratively constructed legal reasoning benchmark consisting of 162 tasks covering six different types of legal reasoning. LegalBench was built through an interdisciplinary process, in which we collected tasks designed and hand-crafted by legal professionals. Because these subject matter experts took a leading role in construction, tasks either measure legal reasoning capabilities that are practically useful, or measure reasoning skills that lawyers find interesting. To enable cross-disciplinary conversations about LLMs in the law, we additionally show how popular legal frameworks for describing legal reasoning -- which distinguish between its many forms -- correspond to LegalBench tasks, thus giving lawyers and LLM developers a common vocabulary. This paper describes LegalBench, presents an empirical evaluation of 20 open-source and commercial LLMs, and illustrates the types of research explorations LegalBench enables.

Discussion (0). Sign in to comment.

Forward citations

Cited by 13 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 28 citations worldwide. Full citation record

  1. BLAD: A Historically Contextualized, Multilingual Dataset of Bangladeshi Legal Acts (1799 to 2025)

    cs.CL 2026-07 conditional novelty 7.0 of 10

    BLAD releases 1,484 Bangladeshi legal acts (1799–2025) with structural annotations and historical government context.

  2. AI for Statutory Simplification: A Comprehensive State Legal Corpus and Labor Benchmark

    cs.IR 2025-08 conditional novelty 7.0 of 10

    State-of-the-art LLMs with retrieval answer simplified boolean questions about state unemployment insurance law with at best 0.69 F1, well short of reliable end-to-end code simplification.

  3. Reasoning Consensus: Structural Ensembling of LLM Reasoning via Weighted DAG Aggregation

    cs.CL 2026-07 conditional novelty 6.0 of 10

    Combining multiple LLMs' reasoning traces into weighted DAGs gives an auditable consensus graph that matches self-consistency and modestly improves on majority voting.

  4. Toward Automated Detection of Documentation Inconsistencies in Electronic Health Records

    cs.CL 2026-07 conditional novelty 6.0 of 10

    An LLM pipeline applied to 3,000 MIMIC-IV discharge summaries surfaced 3,460 candidate documentation inconsistencies, which the authors organize into a graded ontology of contradiction and ambiguity.

  5. Constraint-Aware Hierarchical Search for Regulation-Driven Fine-Grained Classification

    cs.AI 2026-07 accept novelty 6.0 of 10

    A constraint-aware hierarchical search over a regulatory tree, using local candidates plus structured rule fields, beats strong RAG baselines on four new expert-validated benchmarks.

  6. WisWheat: A Three-Tiered Vision-Language Dataset for Wheat Management

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A new wheat-specific dataset with pretraining, quantitative, and instruction-tuning layers improves VLM performance on wheat stress diagnosis and growth-stage management tasks.

  7. AIMSCheck: Leveraging LLMs for AI-Assisted Review of Modern Slavery Statements Across Jurisdictions

    cs.CY 2025-06 conditional novelty 6.0 of 10

    New UK and Canadian modern slavery statement datasets plus a three-level AI review framework, with fine-tuned models showing strong cross-jurisdictional generalization.

  8. Question Begets Question: Self-Evolving Curriculum for Reinforcement Fine-Tuning on Competition Mathematics

    cs.LG 2026-08 conditional novelty 5.0 of 10

    A self-evolving curriculum that retrains a language model on variants of problems it can mostly get right lifts AIME pass@1 from 5.6% to 16.5%, beating static augmentation under the same data budget.

  9. Confidently Wrong: Exception Chain Collapse in Frontier LLM Rule Evaluation

    cs.AI 2026-07 conditional novelty 5.0 of 10

    Nested exception-chain eligibility breaks frontier LLMs in unstable ways; an SMT execution layer makes outcomes deterministic given authored rules.

  10. Enterprise Large Language Model Evaluation Benchmark

    cs.AI 2025-06 reject novelty 5.0 of 10

    A 14-task enterprise LLM benchmark built mostly from GPT-4o-generated labels and scored by GPT-4o-as-judge shows open-source models closing the reasoning gap, but the dataset is not public and the evaluation is partly...

  11. AILQA: Evaluating AI-Driven Legal Question Answering Systems for the Indian Legal System

    cs.CL 2026-07 conditional novelty 4.0 of 10

    RAG with top-3 chunk retrieval lifts smaller LLMs on Indian legal QA (Llama2-70B: 45.7% to 51.7% on AIBE) but often hurts large models, and under the study's own rating protocol some AI answers outscored the reference...

  12. L-MARS: Legal Multi-Agent System with Agentic Search and Citation-Faithfulness Audit

    cs.AI 2025-08 reject novelty 4.0 of 10

    The abstract reports that a judge-driven multi-agent loop improves legal-citation faithfulness from 0.13 to 0.25 strict F1 and cuts the no-citation rate from 34% to 13%, but the provided manuscript text does not conta...

  13. An Integrated Framework of Prompt Engineering and Multidimensional Knowledge Graphs for Legal Dispute Analysis

    cs.AI 2025-07 conditional novelty 4.0 of 10

    A prompt-plus-knowledge-graph framework for legal dispute analysis reports improved LLM sensitivity and citation accuracy on a 100-pair test set, but with limited statistical support.

Pith tools