Pith. sign in

Paper Citation Record · LEDGER

LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 42 inbound Pith citation observations for arXiv:2308.11462.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2308.11462 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 42 of 42 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 42 of 42 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T14:04:30.177659Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

28
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation a2a5231f-a917-4034-8073-9b3cf8d2e29f · inbound

AIMS.au: A Dataset for the Analysis of Modern Slavery Countermeasures in Corporate Statements cites this paper.

AIMS.au: A Dataset for the Analysis of Modern Slavery Countermeasures in Corporate Statements LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-08T14:04:30.177659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:04:30.177659Z digest=sha256:834f0a657cc6d35e6ccf38d227004ce8e33fbe798bb9bd1af2b0b7a0bafa2b05

Observation 067804ab-144d-4d0d-9762-7e1ba74c4a27 · inbound

Context Reasoner: Incentivizing Reasoning Capability for Contextualized Privacy and Safety Compliance via Reinforcement Learning cites this paper.

Context Reasoner: Incentivizing Reasoning Capability for Contextualized Privacy and Safety Compliance via Reinforcement Learning LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T15:37:26.528544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:37:26.528544Z digest=sha256:fb58929137ac883d757a0b514558446fefa03a8b767b2f02018a74a955f709d2

Observation 4457ddbc-cf62-45a7-bc22-44e191c38f93 · inbound

Towards Large Reasoning Models for Agriculture cites this paper.

Towards Large Reasoning Models for Agriculture LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T14:21:43.049112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:21:43.049112Z digest=sha256:18c9d4119d0641c5508780cc963d11e5878af3d3c1882bad885da54a93419689

Observation 26293698-8d4e-4e0a-aad9-ac0f2333dea5 · inbound

AIMSCheck: Leveraging LLMs for AI-Assisted Review of Modern Slavery Statements Across Jurisdictions cites this paper.

AIMSCheck: Leveraging LLMs for AI-Assisted Review of Modern Slavery Statements Across Jurisdictions LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:23.083854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:23.083854Z digest=sha256:9d5f244393eef533b4ad651a6b82d702dbc65997a79638c83b1cc0da642eef70

Observation a0539aad-19f8-411e-b7d9-c78ec157b0f0 · inbound

WisWheat: A Three-Tiered Vision-Language Dataset for Wheat Management cites this paper.

WisWheat: A Three-Tiered Vision-Language Dataset for Wheat Management LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:10.295568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:02:10.295568Z digest=sha256:abf68656aff8d2f3b0c559fd2477ed8cdad2fc82b271cd0f528a87ead33788c9

Observation 0c23ae7a-5ba5-4c5b-8fa2-fd50e0ddb26d · inbound

Enterprise Large Language Model Evaluation Benchmark cites this paper.

Enterprise Large Language Model Evaluation Benchmark LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T22:56:32.246086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:56:32.246086Z digest=sha256:cbf08d0f041fe940e763275b197292b0823febb2d2bc23d2c45856e045c37358

Observation 3eb6f863-281f-4072-b227-c83fca4980d6 · inbound

An Integrated Framework of Prompt Engineering and Multidimensional Knowledge Graphs for Legal Dispute Analysis cites this paper.

An Integrated Framework of Prompt Engineering and Multidimensional Knowledge Graphs for Legal Dispute Analysis LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T18:35:40.873373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:35:40.873373Z digest=sha256:464cd43756e81992d097fd845e1aa3bfcc4d7b2e4029864d6689f408066e871b

Observation 9ce16b5a-141e-4626-be52-158ad33feedf · inbound

AI for Statutory Simplification: A Comprehensive State Legal Corpus and Labor Benchmark cites this paper.

AI for Statutory Simplification: A Comprehensive State Legal Corpus and Labor Benchmark LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T15:51:46.275259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:51:46.275259Z digest=sha256:09a1c0ce0eab40a4ee0b1f4d1a81e61a12bbc370e65c63264a58c59d9e59005d

Observation 0efc6177-c0c8-4f41-8fd8-c92df459e153 · inbound

L-MARS: Legal Multi-Agent System with Agentic Search and Citation-Faithfulness Audit cites this paper.

L-MARS: Legal Multi-Agent System with Agentic Search and Citation-Faithfulness Audit LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T13:20:40.339812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:20:40.339812Z digest=sha256:5fdb454f915a348f78a4d492a84d2c392b2832a0dd3e873baae7ede2daecd99c

Observation 19fa69a8-2765-403f-8f13-90135d012582 · inbound

Nirvana: A Specialized Generalist Model With Task-Aware Memory Mechanism cites this paper.

Nirvana: A Specialized Generalist Model With Task-Aware Memory Mechanism LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T03:05:48.097934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-18T03:05:06.069642Z digest=sha256:7d746bd39476e0e1eeb7a0af08a55f9d5eccdc35e814715661ff9117950ef4cc

Observation 11d79624-30eb-429a-9918-0e82de232510 · inbound

Towards Real-World Validity in Generative AI Benchmarks: Understanding and Designing Domain-Centered Evaluations for Journalism Practitioners cites this paper.

Towards Real-World Validity in Generative AI Benchmarks: Understanding and Designing Domain-Centered Evaluations for Journalism Practitioners LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T11:01:17.029897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-18T10:59:16.139525Z digest=sha256:5afd150560c09b939168954886dab6642c89c4b1112f7938afb61d443bafce42

Observation 2fb5a109-91ee-4c48-9816-733b21e74e4e · inbound

CapTrack: Multifaceted Evaluation of Forgetting in LLM Post-Training cites this paper.

CapTrack: Multifaceted Evaluation of Forgetting in LLM Post-Training LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-25T06:45:25.582582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-25T06:40:51.046965Z digest=sha256:cf97edbd27c97425278bd8bc77ab864b4fde96b3b6a64aa0266514186e073601

Observation 0770fa1b-4680-4353-8f8d-7a5cd87c9fbc · inbound

SciPredict: Can LLMs Predict the Outcomes of Scientific Experiments in Natural Sciences? cites this paper.

SciPredict: Can LLMs Predict the Outcomes of Scientific Experiments in Natural Sciences? LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:36:03.821732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T15:55:34.768853Z digest=sha256:0942766c7c1970db5fe057925b0ca92716c6c2294f759fab37a89241e9ae81dd

Observation e0c78989-4ba3-4b36-a240-7692ee3934dc · inbound

LegalBench-BR: A Benchmark for Evaluating Large Language Models on Brazilian Legal Decision Classification cites this paper.

LegalBench-BR: A Benchmark for Evaluating Large Language Models on Brazilian Legal Decision Classification LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T11:56:26.755806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T04:27:23.053818Z digest=sha256:c4c1d4bbeda870fb5e83cac668ab1eb23fa07451679e433eda30e3fa80d5d179

Observation e4704eac-14b1-4ed5-9f66-7916bd338bd6 · inbound

KnowPilot: Your Knowledge-Driven Copilot for Domain Tasks cites this paper.

KnowPilot: Your Knowledge-Driven Copilot for Domain Tasks LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-10T06:16:20.746297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T06:15:44.360621Z digest=sha256:4e48fdc11c576b9f24f5b202196d13f35ce0d99f1e256dd95738399e12b12a29

Observation 398924d6-1683-4640-a4fd-9614f707dddd · inbound

Learning When Not to Decide: A Framework for Overcoming Factual Presumptuousness in AI Adjudication cites this paper.

Learning When Not to Decide: A Framework for Overcoming Factual Presumptuousness in AI Adjudication LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T02:11:57.255230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T02:08:24.770003Z digest=sha256:ab7cc459897cdc96a5564667dedb5c999eaf8e515856977644b4a85f8b688dd1

Observation a7818266-5c70-4bb6-b044-e164685aefa9 · inbound

Exploiting LLM-as-a-Judge Disposition on Free Text Legal QA via Prompt Optimization cites this paper.

Exploiting LLM-as-a-Judge Disposition on Free Text Legal QA via Prompt Optimization LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T13:41:10.502069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T01:10:56.318664Z digest=sha256:dd5be85b5746ba3021d161d818d578fdf705a625a9ebb0b542315092fdcec27e

Observation cf967cc5-c82b-4272-8813-f631e1ac7063 · inbound

Breaking the Secret: Economic Interventions for Combating Collusion in Embodied Multi-Agent Systems cites this paper.

Breaking the Secret: Economic Interventions for Combating Collusion in Embodied Multi-Agent Systems LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models

Reference 42

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T21:16:16.378379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T06:13:17.633146Z digest=sha256:080b9553cac01ef4f0e2fb0bd12a9b59f5480f24d19285213a5cf25bc8fb2fc6

Observation 1fd275bc-c7b0-44e8-8671-6e4c428bb6cd · inbound

Expert Evaluation of LLM's Open-Ended Legal Reasoning on the Japanese Bar Exam Writing Task cites this paper.

Expert Evaluation of LLM's Open-Ended Legal Reasoning on the Japanese Bar Exam Writing Task LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T21:16:36.010805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T05:59:54.438189Z digest=sha256:015a1e0b49dbfc38fb3ae57f696d5adeb33cb96d51cf4ab3bf3682a818fd9fa5

Observation ebcbdeb7-2d0d-4972-9bbf-2ed349ec7aad · inbound

Safety Drift After Fine-Tuning: Evidence from High-Stakes Domains cites this paper.

Safety Drift After Fine-Tuning: Evidence from High-Stakes Domains LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T23:11:19.337354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-07T17:53:57.169962Z digest=sha256:282cfb4e3e0ce2d9064896f863b895d6c61b3da85e0197f8dac56ca1779bb387

Observation 11fe66d3-1f22-46f5-8cfc-2863a2f81a9c · inbound

Query-efficient model evaluation using cached responses cites this paper.

Query-efficient model evaluation using cached responses LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T23:35:07.682388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T23:28:47.530333Z digest=sha256:309e755417ef9b3b2d5e7572c561120f0aa7de71b287c2adc4486419d0394587

Observation 2d676823-a8f4-4cee-8652-9637cce9eb42 · inbound

Byte-Exact Deduplication in Retrieval-Augmented Generation: A Three-Regime Empirical Analysis Across Public Benchmarks cites this paper.

Byte-Exact Deduplication in Retrieval-Augmented Generation: A Three-Regime Empirical Analysis Across Public Benchmarks LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:26:24.208127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-12T04:18:11.836537Z digest=sha256:bdc85e12d2101d59135b685640608a495a4c2d2c3a4c6ec15b606cff85adf6ea

Observation f34661f7-0ff0-484e-953f-b59557e57208 · inbound

Decompose-and-Refine: Structured Legal Question Answering with Parametric Retrieval cites this paper.

Decompose-and-Refine: Structured Legal Question Answering with Parametric Retrieval LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T13:24:39.984446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T13:23:16.122654Z digest=sha256:57ccc247c65571a13a64ca00705cfb35fd1d0cc447e73f22c96105c594a754ee

Observation 0a20cf91-1c6c-49fb-80ab-640f03d4acf5 · inbound

TypedCSIP: Typed Counterfactual Pretraining for Chinese Legislative Conflict Classification cites this paper.

TypedCSIP: Typed Counterfactual Pretraining for Chinese Legislative Conflict Classification LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:14:00.218650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-29T22:06:28.247167Z digest=sha256:afe021f617d42f6a3d091029f79ec458e2ea27842482f6434107aa95e907276f

Observation 762cbd88-088c-4f89-a677-07649586ee0f · inbound

UA-Legal-Bench: A Benchmark for Evaluating Large Language Models on Ukrainian Legal Reasoning cites this paper.

UA-Legal-Bench: A Benchmark for Evaluating Large Language Models on Ukrainian Legal Reasoning LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T12:23:24.527606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-29T12:15:12.570661Z digest=sha256:a74f885d3ebebe6bd62755b549e0d4c67e1f1c264a72e8949c45f921e8e0c6bb

Observation 89084092-98aa-40f3-ae4c-8eaade03f412 · inbound

Multi-Legal-Bench: Evaluating LLMs on Legal Reasoning Across Jurisdictions, Languages, and Legal Traditions cites this paper.

Multi-Legal-Bench: Evaluating LLMs on Legal Reasoning Across Jurisdictions, Languages, and Legal Traditions LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T08:33:15.935573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-29T07:24:08.269519Z digest=sha256:da3c35d8d58ae4ae1c680881053fb06e787549d2490dcacae84e18ca74131b19

Observation 0e17b22c-bbd3-4541-94e6-7a937abef3c6 · inbound

Citation Grounding: Detecting and Reducing LLM Citation Hallucinations via Legal Citation Graphs cites this paper.

Citation Grounding: Detecting and Reducing LLM Citation Hallucinations via Legal Citation Graphs LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T20:32:37.522670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T18:37:22.299236Z digest=sha256:d8c715247b111c2e337e598218577ede02a11cdc93a97edd0a62c7c157eefb51

Observation c02298e6-473d-41c6-a231-e1f3028cde94 · inbound

GIScholarBench: Benchmarking LLM Overconfidence in GIS Research cites this paper.

GIScholarBench: Benchmarking LLM Overconfidence in GIS Research LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-07-02T22:07:25.908133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T19:17:26.825374Z digest=sha256:04f0fdc0ebe7707ed1ff05b40e6b95f928f63e8eb601b9fc27f897787aebe105

Observation 11bcab69-c3de-42e6-b065-a10de6e042fb · inbound

Trace2Policy: From Expert Behavior Traces to Self-Evolving Decision Agents cites this paper.

Trace2Policy: From Expert Behavior Traces to Self-Evolving Decision Agents LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models

Reference 36

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T05:27:39.663246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-27T13:19:10.714343Z digest=sha256:881d6e4d60731cf8d622bc8316a54fcc9e9db15d7c6e6015ee3804362dc61ac0

Observation 39b03420-626a-435f-bb7b-9451693948dd · inbound

LegalHalluLens: Typed Hallucination Auditing and Calibrated Multi-Agent Debate for Trustworthy Legal AI cites this paper.

LegalHalluLens: Typed Hallucination Auditing and Calibrated Multi-Agent Debate for Trustworthy Legal AI LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-07-03T21:28:58.612291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-27T00:34:53.489694Z digest=sha256:f65142bd5f02ad213424f1c5b580a9b0a769f3c95f7b7403f56a3da30ee8c148

Observation c454b269-996a-4f29-b290-fd070686884d · inbound

The Measurement Gap in the Automation of EU Law: Benchmarking Doctrinal Legal Reasoning under the EU AI Act cites this paper.

The Measurement Gap in the Automation of EU Law: Benchmarking Doctrinal Legal Reasoning under the EU AI Act LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models

Reference 45

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T23:29:02.495106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-26T22:16:25.919667Z digest=sha256:2f16bd12982d28e10654c617d604026df04851bc9824535e0132a03177492d89

Observation 5d0442ac-aa2b-455d-935b-99fef5a8c367 · inbound

Answer Engineering: Local Trajectory Editing for Protocol-Constrained Decision Making in Large Language Models cites this paper.

Answer Engineering: Local Trajectory Editing for Protocol-Constrained Decision Making in Large Language Models LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-07-04T06:49:37.576858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-26T14:13:13.678245Z digest=sha256:9b1dda17e71a749d10903b7e9e6dea289bd98f56e3f0a346c05db32a361a073a

Observation 9675f5fb-93ac-48c5-bd49-3d74a374a686 · inbound

HAKARI-Bench: A Lightweight Benchmark for Comparing Retrieval Architectures and Efficiency Settings under Unified Conditions cites this paper.

HAKARI-Bench: A Lightweight Benchmark for Comparing Retrieval Architectures and Efficiency Settings under Unified Conditions LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models

Reference 95

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T11:59:51.263031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-26T07:22:34.547816Z digest=sha256:e46daf82717e16a0678dfc3f9f85fd092ce9d4246ebdad7c3683f12492d12b5f

Observation 77a56eeb-1020-40e3-b02d-1e0bc4c61849 · inbound

Legal Reasoning Is Not Lawyering: Rethinking Legal Benchmarks for Pro Se Access to Justice cites this paper.

Legal Reasoning Is Not Lawyering: Rethinking Legal Benchmarks for Pro Se Access to Justice LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T23:19:03.985515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-26T22:26:54.239505Z digest=sha256:5ed000cbe0ac33badd5f39a2a64dd4e7b43ff87e3a3c681f350d81e8bc642268

Observation 58c35bc4-9e48-487d-bae7-3e457539c130 · inbound

How Do Tool-Augmented LLM Agents Perform on Real-World Energy Analytics Tasks? cites this paper.

How Do Tool-Augmented LLM Agents Perform on Real-World Energy Analytics Tasks? LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T15:29:56.715847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-26T01:34:10.103638Z digest=sha256:4580f4d01a6b3edab85621c82605603ac1ad338279d28c83c3ad91a1d5103204

Observation b2379fd8-9d26-4202-822d-e816f773e61c · inbound

Constraint-Aware Hierarchical Search for Regulation-Driven Fine-Grained Classification cites this paper.

Constraint-Aware Hierarchical Search for Regulation-Driven Fine-Grained Classification LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-14T10:38:03.549728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T10:38:03.549728Z digest=sha256:6a1612f55414d8b4119838c5e79572b7a91916b409cda84e78fcf5c554768058

Observation 0f1c2e0f-2e20-40e7-897e-9605cfbc2a4f · inbound

BLAD: A Historically Contextualized, Multilingual Dataset of Bangladeshi Legal Acts (1799 to 2025) cites this paper.

BLAD: A Historically Contextualized, Multilingual Dataset of Bangladeshi Legal Acts (1799 to 2025) LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T19:01:02.854545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T19:01:02.854545Z digest=sha256:939f242a1943a8e3c3801c3e0e737fa3103a4c69907e816e8e30ecc894abf5e1

Observation 7971769f-9742-4175-9f19-32ee4b0d2009 · inbound

AILQA: Evaluating AI-Driven Legal Question Answering Systems for the Indian Legal System cites this paper.

AILQA: Evaluating AI-Driven Legal Question Answering Systems for the Indian Legal System LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-01T14:17:23.403369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T14:17:23.403369Z digest=sha256:818904b08f81c7f0ec27bfda6615d1a45890b2fbce7f1a652f6342a2c1f3981d

Observation 0b3da14e-438a-45a5-b651-7b16132da10e · inbound

Toward Automated Detection of Documentation Inconsistencies in Electronic Health Records cites this paper.

Toward Automated Detection of Documentation Inconsistencies in Electronic Health Records LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T04:04:13.822723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T04:04:13.822723Z digest=sha256:644aa13ae32dcf94ff624e5e240739eb1af2de0eb608b98d75e64ef4a98f0baa

Observation 10731c04-647d-4086-9b51-aa77a094d1c9 · inbound

Confidently Wrong: Exception Chain Collapse in Frontier LLM Rule Evaluation cites this paper.

Confidently Wrong: Exception Chain Collapse in Frontier LLM Rule Evaluation LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models

Reference 2010

Resolution
unresolved
no resolver link, observed 2026-07-30T23:41:20.109043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T23:41:20.109043Z digest=sha256:bf5f86654dec7ab3ddd3b90eb93dcdfabeafc6fca2881285d3d39fe89aed29cb

Observation 1557386e-86b7-4762-b185-713397cfc408 · inbound

Reasoning Consensus: Structural Ensembling of LLM Reasoning via Weighted DAG Aggregation cites this paper.

Reasoning Consensus: Structural Ensembling of LLM Reasoning via Weighted DAG Aggregation LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-01T01:32:21.858955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:32:21.858955Z digest=sha256:7660c1713cf4547a55877fdeaa383abe0e5953dcd597f1a3090fde57588d9376

Observation 082e6d06-398f-4600-a127-18d04123a176 · inbound

Question Begets Question: Self-Evolving Curriculum for Reinforcement Fine-Tuning on Competition Mathematics cites this paper.

Question Begets Question: Self-Evolving Curriculum for Reinforcement Fine-Tuning on Competition Mathematics LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T00:08:23.331668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:08:23.331668Z digest=sha256:973375c56db678b9b6ebacf1f2b4e407e007835184e0e69944c5e9041a249c10