Pith. sign in

REVIEW 11 cited by

LawBench: Benchmarking Legal Knowledge of Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2309.16289 v1 pith:GXFUHCF5 submitted 2023-09-28 cs.CL cs.AIcs.LG

classification cs.CLcs.AIcs.LG
keywords legalllmsknowledgelawbenchtaskswhethercapabilitiesdomain
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large language models (LLMs) have demonstrated strong capabilities in various aspects. However, when applying them to the highly specialized, safe-critical legal domain, it is unclear how much legal knowledge they possess and whether they can reliably perform legal-related tasks. To address this gap, we propose a comprehensive evaluation benchmark LawBench. LawBench has been meticulously crafted to have precise assessment of the LLMs' legal capabilities from three cognitive levels: (1) Legal knowledge memorization: whether LLMs can memorize needed legal concepts, articles and facts; (2) Legal knowledge understanding: whether LLMs can comprehend entities, events and relationships within legal text; (3) Legal knowledge applying: whether LLMs can properly utilize their legal knowledge and make necessary reasoning steps to solve realistic legal tasks. LawBench contains 20 diverse tasks covering 5 task types: single-label classification (SLC), multi-label classification (MLC), regression, extraction and generation. We perform extensive evaluations of 51 LLMs on LawBench, including 20 multilingual LLMs, 22 Chinese-oriented LLMs and 9 legal specific LLMs. The results show that GPT-4 remains the best-performing LLM in the legal domain, surpassing the others by a significant margin. While fine-tuning LLMs on legal specific text brings certain improvements, we are still a long way from obtaining usable and reliable LLMs in legal tasks. All data, model predictions and evaluation code are released in https://github.com/open-compass/LawBench/. We hope this benchmark provides in-depth understanding of the LLMs' domain-specified capabilities and speed up the development of LLMs in the legal domain.

Discussion (0). Sign in to comment.

Forward citations

Cited by 11 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. AI for Statutory Simplification: A Comprehensive State Legal Corpus and Labor Benchmark

    cs.IR 2025-08 conditional novelty 7.0 of 10

    State-of-the-art LLMs with retrieval answer simplified boolean questions about state unemployment insurance law with at best 0.69 F1, well short of reliable end-to-end code simplification.

  2. What out-of-the-box LLMs can(t) do in law? A Turing test in Italian exams for lawyers, judges and notaries

    cs.CY 2026-08 conditional novelty 6.0 of 10

    Out-of-the-box LLMs can match or beat top human candidates on Italian bar and judge essay exams but all fail the notary exam, which requires constrained legal drafting and planning.

  3. Benchmarking LLM-Assisted Blue Teaming via Standardized Threat Hunting

    cs.CR 2025-09 conditional novelty 6.0 of 10

    Standardized modular threat-hunting workflows (CyberTeam) improve LLM performance on blue team tasks compared to open-ended ICL, CoT, and ToT prompting across 30 tasks and 452k samples.

  4. Towards Evaluation for Real-World LLM Unlearning

    cs.AI 2025-08 conditional novelty 6.0 of 10

    DCUE evaluates LLM unlearning by comparing core-token confidence score distributions of the unlearned model and the original model, corrected by a validation set, using the Kolmogorov-Smirnov test.

  5. Using Large Language Models for Legal Decision-Making in Austrian Value-Added Tax Law: An Experimental Study

    cs.CL 2025-07 conditional novelty 6.0 of 10

    RAG-enhanced LLMs slightly outperform fine-tuned LLMs on Austrian/EU VAT questions, but not significantly, and neither approach is ready for full automation.

  6. Measuring Faithfulness and Abstention: An Automated Pipeline for Evaluating LLM-Generated 3-ply Case-Based Legal Arguments

    cs.CL 2025-05 conditional novelty 6.0 of 10

    An automated LLM-based evaluator finds that eight LLMs rarely hallucinate factors in legal argument generation but often omit relevant factors and usually fail to abstain when no common ground exists.

  7. Enterprise Large Language Model Evaluation Benchmark

    cs.AI 2025-06 reject novelty 5.0 of 10

    A 14-task enterprise LLM benchmark built mostly from GPT-4o-generated labels and scored by GPT-4o-as-judge shows open-source models closing the reasoning gap, but the dataset is not public and the evaluation is partly...

  8. ASP2LJ : An Adversarial Self-Play Laywer Augmented Legal Judgment Framework

    cs.CL 2025-06 conditional novelty 5.0 of 10

    ASP2LJ combines synthetic case generation with adversarial self-play for lawyer agents, improving legal judgment prediction on a Chinese benchmark and on a new rare-case dataset.

  9. A Survey of Automatic Evaluation Methods on Text, Visual and Speech Generations

    cs.CL 2025-06 conditional novelty 5.0 of 10

    A unified taxonomy and comparative meta-evaluation of automatic evaluation methods across text, vision, and speech generation, concluding that LLM-based evaluators dominate current practice.

  10. When Large Language Models Meet Law: Dual-Lens Taxonomy, Technical Advances, and Ethical Governance

    cs.CL 2025-07 conditional novelty 4.0 of 10

    A literature review that classifies LLM-for-law research using a dual-lens taxonomy of Toulmin argumentation components and legal practitioner roles.

  11. Pangu Embedded: An Efficient Dual-system LLM Reasoner with Metacognition

    cs.CL 2025-05 conditional novelty 4.0 of 10

    Pangu Embedded, a 7B reasoner trained with iterative distillation, RL, and an adaptive fast/slow thinking scheme, reports superior benchmark scores to similarly sized Qwen3-8B and GLM-4-9B.

Pith tools