Pith. sign in

Paper Citation Record · LEDGER

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models

As of 20 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 10 inbound Pith citation observations for arXiv:2505.14107.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.14107 v5

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 10 of 10 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T13:38:11.263118Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T19:27:18.623344Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 09047aae-f843-40e3-85b4-cf933f609390 · inbound

Beyond Classification Accuracy: Neural-MedBench and the Need for Deeper Reasoning Benchmarks cites this paper.

Beyond Classification Accuracy: Neural-MedBench and the Need for Deeper Reasoning Benchmarks DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-08-18T02:15:33.670502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-18T13:26:58.309566Z digest=sha256:ff1217bab7c779be940403e484b89435cfe46d3dea22c6e8accebd257538b7b4

Observation b7ec127c-9cc0-44a7-94c9-86527838c0db · inbound

MedXIAOHE: A Comprehensive Recipe for Building Medical MLLMs cites this paper.

MedXIAOHE: A Comprehensive Recipe for Building Medical MLLMs DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models

Reference 76

Resolution
verified exact
arxiv_id, observed 2026-08-18T02:15:33.670502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-15T22:52:30.992054Z digest=sha256:583b2816da65cb4d13e4c5bc8ab281b057e81c77d2baad183d9000968004b7d9

Observation 56606448-0adb-4b38-bf1b-6c8a2b8894c4 · inbound

From Exposure to Internalization: Dual-Stream Calibration for In-context Clinical Reasoning cites this paper.

From Exposure to Internalization: Dual-Stream Calibration for In-context Clinical Reasoning DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-08-18T02:15:33.670502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T19:19:15.053725Z digest=sha256:65c283faa0a70681958e9757f6a68418bd4e840ddebc9483a75799c28c07fef1

Observation eb509866-ada7-44b7-a024-1d689965352d · inbound

EpiGraph: Building Generalists for Evidence-Intensive Epilepsy Reasoning in the Wild cites this paper.

EpiGraph: Building Generalists for Evidence-Intensive Epilepsy Reasoning in the Wild DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-08-18T02:15:33.670502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-12T04:03:34.773345Z digest=sha256:af20aa4d3bc70305ec20e97617a9a7f3376bdd12e94c7af1aa2f12044ce029dc

Observation 4c17d835-2f8e-4172-ae51-856250113229 · inbound

EpiGraph: Building Generalists for Evidence-Intensive Epilepsy Reasoning in the Wild cites this paper.

EpiGraph: Building Generalists for Evidence-Intensive Epilepsy Reasoning in the Wild DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-08-18T02:15:33.670502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-14T21:25:31.662152Z digest=sha256:40b46343c281487f1078334ce4360e6eb21d7bfcc6feee0f4cdb3bcbb3fb43e7

Observation 7eb391fb-79aa-4e59-ba84-88b6f2ba5083 · inbound

EHRBench: An Automated and Reliable EHR-based Benchmark for Clinical Decision Making with LLMs cites this paper.

EHRBench: An Automated and Reliable EHR-based Benchmark for Clinical Decision Making with LLMs DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models

Reference 131

Resolution
verified exact
arxiv_id, observed 2026-08-18T02:15:33.670502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-29T06:41:06.828814Z digest=sha256:fc52008b65d94b7cafec551d632d1bf1d147109829229f6575fd644292398999

Observation 96c78045-c903-4f3b-afa0-f6f2d46e03af · inbound

Beyond English benchmarks: clinical llm evaluation in Brazilian Portuguese cites this paper.

Beyond English benchmarks: clinical llm evaluation in Brazilian Portuguese DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-08-18T02:15:33.670502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-27T21:40:16.052239Z digest=sha256:a5ba7945fef1afca96e281fbd3fae2925d2385ef5a3a3732234d46bd462e341c

Observation d98ffec1-12da-4a37-ab73-142cbea8de42 · inbound

RareDxR1: Autonomous Medical Reasoning for Rare Disease Diagnosis Beyond Human Annotation cites this paper.

RareDxR1: Autonomous Medical Reasoning for Rare Disease Diagnosis Beyond Human Annotation DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-08-18T02:15:33.670502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-02T19:21:44.653877Z digest=sha256:7e9b9eb07e6513fcbf5af23724b9621c2ecff1fc48e89f9abe3ed0e95da2c365

Observation fbbe7ad0-62de-42bd-a871-203c3e9afd2d · inbound

DeepLens Diagnosis Agent: Agentic Workflow Design Lets a Small Reasoning Model Compete with Frontier LLMs cites this paper.

DeepLens Diagnosis Agent: Agentic Workflow Design Lets a Small Reasoning Model Compete with Frontier LLMs DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-02T13:38:11.263118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:38:11.263118Z digest=sha256:c8db0c194c34edb54f1254f9c2bf9cacbdd0dbcd8c7fd270eae92ad86ee0e1ce

Observation 3530e665-cd62-480a-96b0-5471f5a47e1d · inbound

Evaluating Multi-Turn Multimodal Diagnostic Reasoning on Challenging Real-World Clinical Cases cites this paper.

Evaluating Multi-Turn Multimodal Diagnostic Reasoning on Challenging Real-World Clinical Cases DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-01T01:07:59.356644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T01:07:59.356644Z digest=sha256:1c9cf512a3f66145285382fa89b4aa310b7e06011089d047344de8f1429e7b35