Pith. sign in

Paper Citation Record · LEDGER

LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models

As of 14 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 49 inbound Pith citation observations for arXiv:2308.11462.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2308.11462 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 49 of 49 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 49 of 49 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-14T04:43:19.334168Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

28
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation c77353f1-115c-4f9d-8422-8352267a617a · inbound

AIDE: Attribute-Guided MultI-Hop Data Expansion for Data Scarcity in Task-Specific Fine-tuning cites this paper.

AIDE: Attribute-Guided MultI-Hop Data Expansion for Data Scarcity in Task-Specific Fine-tuning LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T20:03:20.597044Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:03:20.597044Z digest=sha256:cef0a753ee90d9ec5ebf04f93bf0100607268e8e389e82bba96fbe4239618448

Observation 9578b812-3808-46ea-82e2-455156ae3030 · inbound

TrimLLM: Progressive Layer Dropping for Domain-Specific LLMs cites this paper.

TrimLLM: Progressive Layer Dropping for Domain-Specific LLMs LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T15:13:23.511117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:13:23.511117Z digest=sha256:9250e46208ebfba6c0eb5f792a613220308af0a4402025bb46e68d6bbaf481a7

Observation 139bd959-449b-45b8-9680-ded324374c43 · inbound

LegalGuardian: A Privacy-Preserving Framework for Secure Integration of Large Language Models in Legal Practice cites this paper.

LegalGuardian: A Privacy-Preserving Framework for Secure Integration of Large Language Models in Legal Practice LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T18:53:52.471004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T18:53:52.471004Z digest=sha256:7ee3b0ced89028a309f03b7842e80a57aa06cbe073b3736dc6979e3b4f9a7746

Observation 8a430616-9e27-4c79-a038-68fdfff5f118 · inbound

FinanceQA: A Benchmark for Evaluating Financial Analysis Capabilities of Large Language Models cites this paper.

FinanceQA: A Benchmark for Evaluating Financial Analysis Capabilities of Large Language Models LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-10T00:53:29.674023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T00:53:29.674023Z digest=sha256:e7f89cc1bbc83f1d29941cd0f82452b957f92095d55c5c8804a7faa121877edf

Observation a2a5231f-a917-4034-8073-9b3cf8d2e29f · inbound

AIMS.au: A Dataset for the Analysis of Modern Slavery Countermeasures in Corporate Statements cites this paper.

AIMS.au: A Dataset for the Analysis of Modern Slavery Countermeasures in Corporate Statements LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-08T14:04:30.177659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:04:30.177659Z digest=sha256:461c3e88e6255cb20062db1e6110c51499e2064873997edfef84697aba6fc3ae

Observation 067804ab-144d-4d0d-9762-7e1ba74c4a27 · inbound

Context Reasoner: Incentivizing Reasoning Capability for Contextualized Privacy and Safety Compliance via Reinforcement Learning cites this paper.

Context Reasoner: Incentivizing Reasoning Capability for Contextualized Privacy and Safety Compliance via Reinforcement Learning LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T15:37:26.528544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:37:26.528544Z digest=sha256:c7667d1e3732e049a11645e612eab142db5d7da3dfa663bcfeb163ae21bdfad8

Observation 4457ddbc-cf62-45a7-bc22-44e191c38f93 · inbound

Towards Large Reasoning Models for Agriculture cites this paper.

Towards Large Reasoning Models for Agriculture LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T14:21:43.049112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:21:43.049112Z digest=sha256:2c743ba5e66dc4d9b95ae9711a857c854b6c4252f445e6db8d4860d520459ce2

Observation 26293698-8d4e-4e0a-aad9-ac0f2333dea5 · inbound

AIMSCheck: Leveraging LLMs for AI-Assisted Review of Modern Slavery Statements Across Jurisdictions cites this paper.

AIMSCheck: Leveraging LLMs for AI-Assisted Review of Modern Slavery Statements Across Jurisdictions LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:23.083854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:23.083854Z digest=sha256:49f64c6463a14d02d2a7fa4b82fa95cc99f79b7bc5edc6840b8621d9b71d1d54

Observation a0539aad-19f8-411e-b7d9-c78ec157b0f0 · inbound

WisWheat: A Three-Tiered Vision-Language Dataset for Wheat Management cites this paper.

WisWheat: A Three-Tiered Vision-Language Dataset for Wheat Management LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:10.295568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:02:10.295568Z digest=sha256:d40c42a76e593203c0484026006c7fa73b1d224b9512b84d243a33001fc3657f

Observation 0c23ae7a-5ba5-4c5b-8fa2-fd50e0ddb26d · inbound

Enterprise Large Language Model Evaluation Benchmark cites this paper.

Enterprise Large Language Model Evaluation Benchmark LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T22:56:32.246086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:56:32.246086Z digest=sha256:dd184e97d745435b1e249df2a5360e6c634ed583a20d999f494535d2de0f88e4

Observation 3eb6f863-281f-4072-b227-c83fca4980d6 · inbound

An Integrated Framework of Prompt Engineering and Multidimensional Knowledge Graphs for Legal Dispute Analysis cites this paper.

An Integrated Framework of Prompt Engineering and Multidimensional Knowledge Graphs for Legal Dispute Analysis LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T18:35:40.873373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:35:40.873373Z digest=sha256:ca7efb9ded5376d308dcdaf636ac76859fd0e17c59b19b9ede402001e3e58b6e

Observation 9ce16b5a-141e-4626-be52-158ad33feedf · inbound

AI for Statutory Simplification: A Comprehensive State Legal Corpus and Labor Benchmark cites this paper.

AI for Statutory Simplification: A Comprehensive State Legal Corpus and Labor Benchmark LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T15:51:46.275259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:51:46.275259Z digest=sha256:bcade7e9c82ad3e04c1d04d553b28aa205f9fc5cb1cb72bfa44a46ed37a4650a

Observation 0efc6177-c0c8-4f41-8fd8-c92df459e153 · inbound

L-MARS: Legal Multi-Agent System with Agentic Search and Citation-Faithfulness Audit cites this paper.

L-MARS: Legal Multi-Agent System with Agentic Search and Citation-Faithfulness Audit LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T13:20:40.339812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:20:40.339812Z digest=sha256:e2ef963a6f8bb602cc8e00745a5b27988a416ea24df8f57cb62c58678406eb32

Observation 19fa69a8-2765-403f-8f13-90135d012582 · inbound

Nirvana: A Specialized Generalist Model With Task-Aware Memory Mechanism cites this paper.

Nirvana: A Specialized Generalist Model With Task-Aware Memory Mechanism LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T03:05:48.097934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-18T03:05:06.069642Z digest=sha256:874095208428130115f62c0e41cf784c925bb40ad09c55894d4b5105eb85dde7

Observation 11d79624-30eb-429a-9918-0e82de232510 · inbound

Towards Real-World Validity in Generative AI Benchmarks: Understanding and Designing Domain-Centered Evaluations for Journalism Practitioners cites this paper.

Towards Real-World Validity in Generative AI Benchmarks: Understanding and Designing Domain-Centered Evaluations for Journalism Practitioners LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T11:01:17.029897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-18T10:59:16.139525Z digest=sha256:e666290fac44f42c0935037b8969880ae912a04d39c93606f6541f27e07d35f1

Observation 2fb5a109-91ee-4c48-9816-733b21e74e4e · inbound

CapTrack: Multifaceted Evaluation of Forgetting in LLM Post-Training cites this paper.

CapTrack: Multifaceted Evaluation of Forgetting in LLM Post-Training LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-25T06:45:25.582582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-25T06:40:51.046965Z digest=sha256:c24aa37cdfaff365ce68d94ba829d9137b391e93d40b2b04e5673fc2bdebed66

Observation 0770fa1b-4680-4353-8f8d-7a5cd87c9fbc · inbound

SciPredict: Can LLMs Predict the Outcomes of Scientific Experiments in Natural Sciences? cites this paper.

SciPredict: Can LLMs Predict the Outcomes of Scientific Experiments in Natural Sciences? LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:36:03.821732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T15:55:34.768853Z digest=sha256:0346cb9c2a2527c54f8af51f4de14dfb005ae8b8351dcd15f848905351548cf4

Observation e0c78989-4ba3-4b36-a240-7692ee3934dc · inbound

LegalBench-BR: A Benchmark for Evaluating Large Language Models on Brazilian Legal Decision Classification cites this paper.

LegalBench-BR: A Benchmark for Evaluating Large Language Models on Brazilian Legal Decision Classification LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T11:56:26.755806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T04:27:23.053818Z digest=sha256:70a21f526aac8a02cbe29158c72a247e06a97413408952fe850e7a2022b86f24

Observation e4704eac-14b1-4ed5-9f66-7916bd338bd6 · inbound

KnowPilot: Your Knowledge-Driven Copilot for Domain Tasks cites this paper.

KnowPilot: Your Knowledge-Driven Copilot for Domain Tasks LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-10T06:16:20.746297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T06:15:44.360621Z digest=sha256:2d66da604f5719ad980896bf70e92166a3c00f3fc4e221d90fecdd1a58c3bd91

Observation 398924d6-1683-4640-a4fd-9614f707dddd · inbound

Learning When Not to Decide: A Framework for Overcoming Factual Presumptuousness in AI Adjudication cites this paper.

Learning When Not to Decide: A Framework for Overcoming Factual Presumptuousness in AI Adjudication LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T02:11:57.255230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T02:08:24.770003Z digest=sha256:ddb08535516c331028c74941aa3d2652eea2274765c0270b7a28e8848a7c1114

Observation a7818266-5c70-4bb6-b044-e164685aefa9 · inbound

Exploiting LLM-as-a-Judge Disposition on Free Text Legal QA via Prompt Optimization cites this paper.

Exploiting LLM-as-a-Judge Disposition on Free Text Legal QA via Prompt Optimization LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T13:41:10.502069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T01:10:56.318664Z digest=sha256:8165ad87f836846f0b6af5f726ecf5dcaf0ecc3498994f3dc1b118e47129d194

Observation cf967cc5-c82b-4272-8813-f631e1ac7063 · inbound

Breaking the Secret: Economic Interventions for Combating Collusion in Embodied Multi-Agent Systems cites this paper.

Breaking the Secret: Economic Interventions for Combating Collusion in Embodied Multi-Agent Systems LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models

Reference 42

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T21:16:16.378379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-08T06:13:17.633146Z digest=sha256:66f1eea8a6ecd07d496e2d57859094089eea9b7a60fadbb42c4eef9933efdcbc

Observation 1fd275bc-c7b0-44e8-8671-6e4c428bb6cd · inbound

Expert Evaluation of LLM's Open-Ended Legal Reasoning on the Japanese Bar Exam Writing Task cites this paper.

Expert Evaluation of LLM's Open-Ended Legal Reasoning on the Japanese Bar Exam Writing Task LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T21:16:36.010805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-08T05:59:54.438189Z digest=sha256:40bf0ab0e8e06bfc4a8243103b3c7081f6571f9939f8a73ae0d3675ee06ddcb3

Observation ebcbdeb7-2d0d-4972-9bbf-2ed349ec7aad · inbound

Safety Drift After Fine-Tuning: Evidence from High-Stakes Domains cites this paper.

Safety Drift After Fine-Tuning: Evidence from High-Stakes Domains LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T23:11:19.337354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-07T17:53:57.169962Z digest=sha256:acfaaf89419303f9b8555666c8718d018e4524bd9883ead3ddac7f871d4bdee4

Observation 11fe66d3-1f22-46f5-8cfc-2863a2f81a9c · inbound

Query-efficient model evaluation using cached responses cites this paper.

Query-efficient model evaluation using cached responses LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T23:35:07.682388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T23:28:47.530333Z digest=sha256:44e85f11d019c80171b3d608a68bdfa4063e32394aa2d89cec968f294ea1af62

Observation 2d676823-a8f4-4cee-8652-9637cce9eb42 · inbound

Byte-Exact Deduplication in Retrieval-Augmented Generation: A Three-Regime Empirical Analysis Across Public Benchmarks cites this paper.

Byte-Exact Deduplication in Retrieval-Augmented Generation: A Three-Regime Empirical Analysis Across Public Benchmarks LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:26:24.208127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-12T04:18:11.836537Z digest=sha256:350b92aa4ffdbd486d5f18e81e0d547d3cc9da38f11d669a89fb03a2f4f08808

Observation f34661f7-0ff0-484e-953f-b59557e57208 · inbound

Decompose-and-Refine: Structured Legal Question Answering with Parametric Retrieval cites this paper.

Decompose-and-Refine: Structured Legal Question Answering with Parametric Retrieval LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T13:24:39.984446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T13:23:16.122654Z digest=sha256:65a8d30e2650fc627a06738080eb4ee998bbc3ebcac3a0d112401e0723344c2e

Observation 0a20cf91-1c6c-49fb-80ab-640f03d4acf5 · inbound

TypedCSIP: Typed Counterfactual Pretraining for Chinese Legislative Conflict Classification cites this paper.

TypedCSIP: Typed Counterfactual Pretraining for Chinese Legislative Conflict Classification LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:14:00.218650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-29T22:06:28.247167Z digest=sha256:2ddfc967e1d18f55da314b8ca7fb2f7a99f273877b8f3c3ffd00a3097650afda

Observation 762cbd88-088c-4f89-a677-07649586ee0f · inbound

UA-Legal-Bench: A Benchmark for Evaluating Large Language Models on Ukrainian Legal Reasoning cites this paper.

UA-Legal-Bench: A Benchmark for Evaluating Large Language Models on Ukrainian Legal Reasoning LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T12:23:24.527606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-29T12:15:12.570661Z digest=sha256:08d4965863bbd954e22b666a2149ce75d52e2495791e296e2d6606c93b2e71a2

Observation 89084092-98aa-40f3-ae4c-8eaade03f412 · inbound

Multi-Legal-Bench: Evaluating LLMs on Legal Reasoning Across Jurisdictions, Languages, and Legal Traditions cites this paper.

Multi-Legal-Bench: Evaluating LLMs on Legal Reasoning Across Jurisdictions, Languages, and Legal Traditions LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T08:33:15.935573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-29T07:24:08.269519Z digest=sha256:b1f4646996eac1f1f4300aabfb622619f408dae8a2a3260d8092377320c0b8ef

Observation 0e17b22c-bbd3-4541-94e6-7a937abef3c6 · inbound

Citation Grounding Measures the Oracle: Graph Coverage Determines Reported LLM Hallucination Rates in Law cites this paper.

Citation Grounding Measures the Oracle: Graph Coverage Determines Reported LLM Hallucination Rates in Law LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T20:32:37.522670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-28T18:37:22.299236Z digest=sha256:4773d87bbeaed383560ef5c6e43bb793ae8f6d00901e5eef153832aadf5bc410

Observation c02298e6-473d-41c6-a231-e1f3028cde94 · inbound

GIScholarBench: Benchmarking LLM Overconfidence in GIS Research cites this paper.

GIScholarBench: Benchmarking LLM Overconfidence in GIS Research LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-07-02T22:07:25.908133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-27T19:17:26.825374Z digest=sha256:bf72677a79f4f6edfe4eeb551c3f5fecbeaa9517c1f11374e380c5b47d033b22

Observation 11bcab69-c3de-42e6-b065-a10de6e042fb · inbound

Trace2Policy: From Expert Behavior Traces to Self-Evolving Decision Agents cites this paper.

Trace2Policy: From Expert Behavior Traces to Self-Evolving Decision Agents LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models

Reference 36

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T05:27:39.663246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-27T13:19:10.714343Z digest=sha256:dea300283ea862dba5584ad3e6eac9196cfa72d71ac44f585c6d62a5fae6ad31

Observation 39b03420-626a-435f-bb7b-9451693948dd · inbound

LegalHalluLens: Typed Hallucination Auditing and Calibrated Multi-Agent Debate for Trustworthy Legal AI cites this paper.

LegalHalluLens: Typed Hallucination Auditing and Calibrated Multi-Agent Debate for Trustworthy Legal AI LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-07-03T21:28:58.612291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-27T00:34:53.489694Z digest=sha256:b46552eed8e183a46d737413b64404a0e83e3c626396a47279f3d189dd2ec0f3

Observation c454b269-996a-4f29-b290-fd070686884d · inbound

The Measurement Gap in the Automation of EU Law: Benchmarking Doctrinal Legal Reasoning under the EU AI Act cites this paper.

The Measurement Gap in the Automation of EU Law: Benchmarking Doctrinal Legal Reasoning under the EU AI Act LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models

Reference 45

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T23:29:02.495106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-26T22:16:25.919667Z digest=sha256:53d433e55be31097f45d56ef8b001bfd1195148584a2a727447e6143119b840f

Observation 5d0442ac-aa2b-455d-935b-99fef5a8c367 · inbound

Answer Engineering: Local Trajectory Editing for Protocol-Constrained Decision Making in Large Language Models cites this paper.

Answer Engineering: Local Trajectory Editing for Protocol-Constrained Decision Making in Large Language Models LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-07-04T06:49:37.576858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-26T14:13:13.678245Z digest=sha256:a5a6601a1f787b3546762a932a07eb15e0a6fb7ce0dc19e2d85b6d4d80cfb2f5

Observation 9675f5fb-93ac-48c5-bd49-3d74a374a686 · inbound

HAKARI-Bench: A Lightweight Benchmark for Comparing Retrieval Architectures and Efficiency Settings under Unified Conditions cites this paper.

HAKARI-Bench: A Lightweight Benchmark for Comparing Retrieval Architectures and Efficiency Settings under Unified Conditions LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models

Reference 95

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T11:59:51.263031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-26T07:22:34.547816Z digest=sha256:56590f5d8cf907c77f20fe2b1bdbf463b4e1556ed02212e4a8203fa93e20e323

Observation 77a56eeb-1020-40e3-b02d-1e0bc4c61849 · inbound

Legal Reasoning Is Not Lawyering: Rethinking Legal Benchmarks for Pro Se Access to Justice cites this paper.

Legal Reasoning Is Not Lawyering: Rethinking Legal Benchmarks for Pro Se Access to Justice LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T23:19:03.985515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-26T22:26:54.239505Z digest=sha256:a4da9bee811e70297fe4b4b9181b06365417c7318dab7920cc33876fffce6db2

Observation 58c35bc4-9e48-487d-bae7-3e457539c130 · inbound

How Do Tool-Augmented LLM Agents Perform on Real-World Energy Analytics Tasks? cites this paper.

How Do Tool-Augmented LLM Agents Perform on Real-World Energy Analytics Tasks? LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T15:29:56.715847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-26T01:34:10.103638Z digest=sha256:3415eb2307ad2d2ba9d311f788b97f7abf3cda831870580895240e5eb1cb3202

Observation b2379fd8-9d26-4202-822d-e816f773e61c · inbound

Constraint-Aware Hierarchical Search for Regulation-Driven Fine-Grained Classification cites this paper.

Constraint-Aware Hierarchical Search for Regulation-Driven Fine-Grained Classification LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-14T10:38:03.549728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T10:38:03.549728Z digest=sha256:ee2299461b465401acf67b73ad1a1ac294f30cea914e1817086bc3ee2d780051

Observation 0f1c2e0f-2e20-40e7-897e-9605cfbc2a4f · inbound

BLAD: A Historically Contextualized, Multilingual Dataset of Bangladeshi Legal Acts (1799 to 2025) cites this paper.

BLAD: A Historically Contextualized, Multilingual Dataset of Bangladeshi Legal Acts (1799 to 2025) LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T19:01:02.854545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T19:01:02.854545Z digest=sha256:b5879608e36955647bca2c8b69bbb6cb01514e911e151e393de75866b8d35539

Observation 7971769f-9742-4175-9f19-32ee4b0d2009 · inbound

AILQA: Evaluating AI-Driven Legal Question Answering Systems for the Indian Legal System cites this paper.

AILQA: Evaluating AI-Driven Legal Question Answering Systems for the Indian Legal System LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-01T14:17:23.403369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T14:17:23.403369Z digest=sha256:d511a1d872c7218970e8c8d251519e191c7465bc9518192b083ee3c0f8f67d53

Observation 0b3da14e-438a-45a5-b651-7b16132da10e · inbound

Toward Automated Detection of Documentation Inconsistencies in Electronic Health Records cites this paper.

Toward Automated Detection of Documentation Inconsistencies in Electronic Health Records LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T04:04:13.822723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T04:04:13.822723Z digest=sha256:0d4a5141c97cdc24adb547654b28bca1a9d998abdd3bef5084b0a4309f715e42

Observation 10731c04-647d-4086-9b51-aa77a094d1c9 · inbound

Confidently Wrong: Exception Chain Collapse in Frontier LLM Rule Evaluation cites this paper.

Confidently Wrong: Exception Chain Collapse in Frontier LLM Rule Evaluation LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models

Reference 2010

Resolution
unresolved
no resolver link, observed 2026-07-30T23:41:20.109043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T23:41:20.109043Z digest=sha256:3d6b36b91b6bb8e83ca94aabceb84d5767600684a048976bb21add556aec4c29

Observation 1557386e-86b7-4762-b185-713397cfc408 · inbound

Reasoning Consensus: Structural Ensembling of LLM Reasoning via Weighted DAG Aggregation cites this paper.

Reasoning Consensus: Structural Ensembling of LLM Reasoning via Weighted DAG Aggregation LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-01T01:32:21.858955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:32:21.858955Z digest=sha256:c1821323211f2eacbe35c36e645f19c5fe5ddba1672dfa89bd29bdf2461480de

Observation 082e6d06-398f-4600-a127-18d04123a176 · inbound

Question Begets Question: Self-Evolving Curriculum for Reinforcement Fine-Tuning on Competition Mathematics cites this paper.

Question Begets Question: Self-Evolving Curriculum for Reinforcement Fine-Tuning on Competition Mathematics LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T00:08:23.331668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:08:23.331668Z digest=sha256:072e71190552018f7bc89862930d948c5ffe54a5f5831f522cf323f20f7703a6

Observation e9d80172-d4ba-400c-a498-5e85697ba782 · inbound

When Does Trace-Driven Evaluation Mislead MoE Expert Caching? Replay Semantics, Workload Contamination, and Operating Regimes cites this paper.

When Does Trace-Driven Evaluation Mislead MoE Expert Caching? Replay Semantics, Workload Contamination, and Operating Regimes LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T00:48:38.348560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:48:38.348560Z digest=sha256:5004b93207eddee582a46e4060cd01fcf013619e7aa0796c6b5778b52d3a592c

Observation a249d4c8-fc89-4e53-94c6-30048010c096 · inbound

When Does Trace-Driven Evaluation Mislead MoE Expert Caching? Replay Semantics, Workload Contamination, and Operating Regimes cites this paper.

When Does Trace-Driven Evaluation Mislead MoE Expert Caching? Replay Semantics, Workload Contamination, and Operating Regimes LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-14T04:43:19.334168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:43:19.334168Z digest=sha256:db2707d9524ca0b79747af5dec82640ae0e4fcf65fab92f1adcd5919893ba337

Observation a5c890be-981e-49d4-b186-c5aec3ffd088 · inbound

V-FiLLM: Verified Financial LLM Reasoning Benchmark cites this paper.

V-FiLLM: Verified Financial LLM Reasoning Benchmark LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:14.710985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T11:27:14.710985Z digest=sha256:db3b89d2e1b1462318d8c930b02e21319520cbb3c3dc27262e9e5fdbb9af5afa