Pith. sign in

REVIEW 18 cited by

CUAD: An Expert-Annotated NLP Dataset for Legal Contract Review

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2103.06268 v2 pith:QCMRM4BT submitted 2021-03-10 cs.CL cs.LG

classification cs.CLcs.LG
keywords contractcuaddatasetlegalreviewatticusexpertslarge
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Many specialized domains remain untouched by deep learning, as large labeled datasets require expensive expert annotators. We address this bottleneck within the legal domain by introducing the Contract Understanding Atticus Dataset (CUAD), a new dataset for legal contract review. CUAD was created with dozens of legal experts from The Atticus Project and consists of over 13,000 annotations. The task is to highlight salient portions of a contract that are important for a human to review. We find that Transformer models have nascent performance, but that this performance is strongly influenced by model design and training dataset size. Despite these promising results, there is still substantial room for improvement. As one of the only large, specialized NLP benchmarks annotated by experts, CUAD can serve as a challenging research benchmark for the broader NLP community.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 18 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. LegalCiteTrust: Benchmarking Citation Trustworthiness in Chinese Long-Form Legal Research Reports

    cs.CL 2026-07 conditional novelty 6.0 of 10

    A new benchmark, LegalCiteTrust, measures citation existence, fidelity, and applicability in Chinese legal research reports and shows that more legal retrieval does not automatically make citations more trustworthy.

  2. Non-vacuous Generalization Bounds for Reinforcement Learning with Verifiable Rewards

    cs.LG 2026-07 reject novelty 6.0 of 10

    Non-vacuous PAC-Bayes generalization bounds for billion-parameter RLVR models, obtained by a Gumbel-max reparameterization and aggressive TinyLoRA distillation/quantization, are claimed for four tasks.

  3. Constraint-Aware Hierarchical Search for Regulation-Driven Fine-Grained Classification

    cs.AI 2026-07 accept novelty 6.0 of 10

    A constraint-aware hierarchical search over a regulatory tree, using local candidates plus structured rule fields, beats strong RAG baselines on four new expert-validated benchmarks.

  4. Evaluating Large Language Models as Expert Annotators

    cs.CL 2025-08 unverdicted novelty 6.0 of 10

    Material Fingerprinting recovers the form and parameters of hyperelastic material models by nearest-neighbor matching of test data against a simulated fingerprint database: exact at zero noise, degrading under 5% noise.

  5. Characterizing Deep Research: A Benchmark and Formal Definition

    cs.CL 2025-08 conditional novelty 6.0 of 10

    Deep research is characterized by high search and reasoning intensity; the new LiveDRBench measures claim-level precision and recall, where the best current model scores 0.55 F1.

  6. IndianBailJudgments-1200: A Multi-Attribute Dataset for Legal NLP on Indian Bail Orders

    cs.CL 2025-07 conditional novelty 6.0 of 10

    IndianBailJudgments-1200 is the first public multi-attribute dataset focused on Indian bail jurisprudence, built by LLM annotation of 1,200 High Court orders.

  7. MM-R5: MultiModal Reasoning-Enhanced ReRanker via Reinforcement Learning for Document Retrieval

    cs.AI 2025-06 conditional novelty 6.0 of 10

    MM-R5, a 7B multimodal re-ranker trained with SFT and GRPO, achieves state-of-the-art page-level recall on MMDocIR by generating per-page reasoning chains.

  8. The Common Pile v0.1: An 8TB Dataset of Public Domain and Openly Licensed Text

    cs.CL 2025-06 conditional novelty 6.0 of 10

    A new 8TB openly-licensed text corpus trains 7B LLMs that are competitive with Llama 1/2, showing that performant models need not depend on unlicensed web data.

  9. CHANCERY: Evaluating Corporate Governance Reasoning Capabilities in Language Models

    cs.AI 2025-06 conditional novelty 6.0 of 10

    A 502-question benchmark for corporate governance reasoning shows current language models reach at most 78.1 percent accuracy.

  10. Enhancing Large Multimodal Models in Key Information Extraction via Scene-Aware Document Synthesis

    cs.CV 2026-07 conditional novelty 5.0 of 10

    Scene-aware multi-agent document synthesis plus error-driven hard-example expansion improves compact Qwen3-VL models on constrained and open-category KIE, topping reported on-device baselines.

  11. LLMs for LLMs: A Structured Prompting Methodology for Long Legal Documents

    cs.AI 2025-09 reject novelty 5.0 of 10

    On CUAD legal contracts, a prompt-engineered QWEN-2 pipeline with chunking and two answer-selection heuristics reportedly outperforms the fine-tuned DeBERTa-large baseline by about 9%, reaching claimed state-of-the-ar...

  12. ACD-CLIP: Decoupling Representation and Dynamic Fusion for Zero-Shot Anomaly Detection

    cs.CV 2025-08 unverdicted novelty 5.0 of 10

    ACD-CLIP improves zero-shot anomaly detection by co-designing a convolutional low-rank adapter with a dynamic fusion gateway that modulates text prompts from visual context.

  13. LLMs for Legal Subsumption in German Employment Contracts

    cs.CL 2025-07 conditional novelty 5.0 of 10

    LLMs reach 80% weighted F1 on German employment contract clause review when given lawyer-distilled examination guidelines, but lag human lawyers when reading full legal sources.

  14. Ranking Free RAG: Replacing Re-ranking with Selection in RAG for Sensitive Domains

    cs.CL 2025-05 reject novelty 5.0 of 10

    METEORA uses DPO-tuned rationales to select and verify evidence chunks in RAG, and claims better recall, precision, evidence efficiency, and poisoning defense, though key evaluation details are missing.

  15. Reducing Hallucinations in Complex Question Answering using Simple Graph-based Retrieval-Augmented Generation (long version)

    cs.CL 2026-06 unverdicted novelty 4.0 of 10

    A graph-augmented RAG system with vector and graph query tools halves hallucinations and raises factual correctness scores on the MoNaCo complex QA benchmark.

  16. Hybrid Topic-Semantic Labeling and Graph Embeddings for Unsupervised Legal Document Clustering

    stat.ML 2025-08 reject novelty 4.0 of 10

    Concatenating Top2Vec and Node2Vec embeddings, where the Node2Vec graph encodes Top2Vec's own topic labels, yields compact clusters, but the gain is largely circular.

  17. L-MARS: Legal Multi-Agent System with Agentic Search and Citation-Faithfulness Audit

    cs.AI 2025-08 reject novelty 4.0 of 10

    The abstract reports that a judge-driven multi-agent loop improves legal-citation faithfulness from 0.13 to 0.25 strict F1 and cuts the no-citation rate from 34% to 13%, but the provided manuscript text does not conta...

  18. When Large Language Models Meet Law: Dual-Lens Taxonomy, Technical Advances, and Ethical Governance

    cs.CL 2025-07 conditional novelty 4.0 of 10

    A literature review that classifies LLM-for-law research using a dual-lens taxonomy of Toulmin argumentation components and legal practitioner roles.

Pith tools