Pith. sign in

Paper Citation Record · LEDGER

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models

As of 18 August 2026, this Paper Citation Record lists 67 of 67 outbound references and 10 inbound Pith citation observations for arXiv:2505.14107.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.14107 v4

Coverage vector

measured 67 of 67 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:42:26.489072Z

measured 77 of 77 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T13:38:11.263118Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T19:27:18.623344Z

Reference resolution

67 of 67 outbound references displayed

  • verified exact1
  • verified fuzzy12
  • unresolved51
  • parse uncertain3
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 56360b9f-9e2f-49b4-8488-d891ecf33e37 · outbound

This paper cites an unresolved cited work.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:42:37.297073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:42:20.680647Z digest=sha256:750ad8bb882240e627ddadea48580e074d461d464209b7af6fbb25bb9ff1b87d

Observation add5754f-9f3e-4d26-9342-ff214d10c0a2 · outbound

This paper cites an unresolved cited work.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:20.789835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:20.789835Z digest=sha256:4deeeccd17808055fcfac4df03d68588d0506e25666467ccbf9e74ace9b35035

Observation 29a188e5-a87d-4848-bc63-f8bab946a0fa · outbound

This paper cites HuatuoGPT-o1, Towards Medical Complex Reasoning with LLMs.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models HuatuoGPT-o1, Towards Medical Complex Reasoning with LLMs

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:20.896208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:20.896208Z digest=sha256:23a82345fded6d9bb759942b9d402e4240f6ea6fc9ca3582001dcf60e89b0157

Observation 1026facc-1e92-45cb-be99-89ba8001c0c8 · outbound

This paper cites An Empirical Study on Eliciting and Improving R1-like Reasoning Models.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models An Empirical Study on Eliciting and Improving R1-like Reasoning Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:20.985769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:20.985769Z digest=sha256:bfdba116e65d8ed2d85220668e3f59cd8d903d2e49cc8c243a8ef0a166d5ef43

Observation 11adcf86-c5af-4869-896e-1239510694c2 · outbound

This paper cites an unresolved cited work.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:42:37.158163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:42:21.068068Z digest=sha256:dba0b505b4c4716da44313017a284959d687eabe447cd8d6dc446ec9313b9815

Observation 0a1a466e-dcb9-4c1c-8e0c-34cc4d7966a2 · outbound

This paper cites an unresolved cited work.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:42:36.899551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:42:21.167084Z digest=sha256:c639d5c4ca801ea9c977d49ac3efcec5061db3ec7881927fd2b3646a33a4e90c

Observation 23b4fb40-e8a3-4f52-9475-dc012462c6f2 · outbound

This paper cites an unresolved cited work.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:42:36.704308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:42:21.281936Z digest=sha256:62d37f926dce5f11390e97bae63c92e3ed5dc5f37c037898ad2d1abe05de86e6

Observation 50cd3b33-c295-4122-838b-a37faeff1e13 · outbound

This paper cites rStar-Math: Small LLMs Can Master Math Reasoning with Self-Evolved Deep Thinking.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models rStar-Math: Small LLMs Can Master Math Reasoning with Self-Evolved Deep Thinking

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:21.412901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:21.412901Z digest=sha256:6e8a7dc66a1883839c9b06a97964d76be48cac753fdcf55496b1478c0c1bcb57

Observation 2f8a5f54-5954-4049-8445-244175106aa9 · outbound

This paper cites an unresolved cited work.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:42:36.512527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:42:21.517209Z digest=sha256:b7385d505312d989d49184cc4c04bdcc6c09ae6e2472ccdee80646cf05fdae89

Observation 6ec658c3-c211-4854-bccd-9d19e1538875 · outbound

This paper cites C-Eval: A Multi-Level Multi-Discipline Chinese Evaluation Suite for Foundation Models.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models C-Eval: A Multi-Level Multi-Discipline Chinese Evaluation Suite for Foundation Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:21.748095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:21.748095Z digest=sha256:8b125a6700946b938c31e4198ab9aa35213d07cf3f7715ac8428897abb1209ff

Observation 13095c2b-6ad1-4105-8fc5-ab73c61da6ea · outbound

This paper cites O1 Replication Journey -- Part 2: Surpassing O1-preview through Simple Distillation, Big Progress or Bitter Lesson?.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models O1 Replication Journey -- Part 2: Surpassing O1-preview through Simple Distillation, Big Progress or Bitter Lesson?

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:21.824586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:21.824586Z digest=sha256:64a8262600f0624c6ad68703821146cee2c296f2c912a1579c25695ab16bee60

Observation d13da93f-2bb2-43c4-8a63-ad77e29ae2cb · outbound

This paper cites O1 Replication Journey -- Part 3: Inference-time Scaling for Medical Reasoning.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models O1 Replication Journey -- Part 3: Inference-time Scaling for Medical Reasoning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:21.902894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:21.902894Z digest=sha256:5f84fe7aa9efa72bc9b3e8a9ed227f2a782e067e0c2465452303850fd87f56e4

Observation 5a3d7ed5-6f5a-4d77-bb14-e3551883f2f3 · outbound

This paper cites an unresolved cited work.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Unresolved cited work

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:21.963612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:21.963612Z digest=sha256:98945afe6e13055cee201a3ecfd5c3e44af4149dac117c3d1e399117e8c6b94c

Observation 7749b063-e0fb-4649-b3c8-43bf035ed78e · outbound

This paper cites Learning Planning-based Reasoning by Trajectories Collection and Process Reward Synthesizing.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Learning Planning-based Reasoning by Trajectories Collection and Process Reward Synthesizing

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-08-07T15:42:27.076360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:42:22.042716Z digest=sha256:f18e0f1a57c01a24f62ee59f4156ec1a5a10e5560dc770423058d8f168d8ce4d

Observation 8acdc184-0a55-44fe-a814-696c7e73685e · outbound

This paper cites an unresolved cited work.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Unresolved cited work

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:22.163572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:22.163572Z digest=sha256:14b391edeca7e117dc9d59ea15cbd4e3a672851d7fa86565b07de1c336445874

Observation 3f1a278f-7fe7-45b2-84a6-794c38506942 · outbound

This paper cites PubMedQA: A Dataset for Biomedical Research Question Answering.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models PubMedQA: A Dataset for Biomedical Research Question Answering

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:22.245082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:22.245082Z digest=sha256:b1479548bbc699fe52310a2d16c03d7889ab5a3bd5b15acace2ca86e267437a8

Observation 252428e3-ca33-4985-8889-e7b155bfcaf2 · outbound

This paper cites an unresolved cited work.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:42:36.278415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:42:22.318276Z digest=sha256:a8e5006522cea45b1b02701f70a4bef2f03afdd0587365cd3a94f1217adcfb31

Observation 5fced3ee-b9e0-408b-9371-699cdb69fb9e · outbound

This paper cites an unresolved cited work.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:42:36.014916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:42:22.418200Z digest=sha256:0a4943c64c50953bd39df42bd60d11f50aa4d3723cc97545071a2583b0f25228

Observation 50458228-0e58-4fb2-bd2e-832ae2561881 · outbound

This paper cites an unresolved cited work.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:42:35.775621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:42:22.509608Z digest=sha256:46ef52e4ddb9802364e630ee3a468dc452b3cbdd9e3f6027e321a04810d2acb2

Observation a7a65e49-e29f-4dbf-8a23-0a860b2877ba · outbound

This paper cites Imitate, Explore, and Self-Improve: A Reproduction Report on Slow-thinking Reasoning Systems.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Imitate, Explore, and Self-Improve: A Reproduction Report on Slow-thinking Reasoning Systems

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:22.583653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:22.583653Z digest=sha256:769e383ab549a92a60083d7d6d0e4d2380d41518c852035b491debd82eaf84ec

Observation e988c9fd-a9fc-4a70-80cb-4fd7071f6588 · outbound

This paper cites From Medprompt to o1: Exploration of Run-Time Strategies for Medical Challenge Problems and Beyond.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models From Medprompt to o1: Exploration of Run-Time Strategies for Medical Challenge Problems and Beyond

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:22.658240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:22.658240Z digest=sha256:eb48cc6b2ac98619efa58e3840af82d8f3f3864d242884497192da9dae91c4b1

Observation ed0aff7f-6605-46a4-80eb-1c7516581bba · outbound

This paper cites an unresolved cited work.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:42:35.540492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:42:22.738801Z digest=sha256:f40a2eed67e6cb5cfe6ddba57cf291ed68dd25eda20fcfe9f6d7c3b58c694704

Observation adb21f5c-2b62-4a06-b4c6-294df9f95d0c · outbound

This paper cites an unresolved cited work.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:42:35.305787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:42:22.826660Z digest=sha256:d3ce93294eab80391cdfe425f4eb3e2157a92d0e7af352e3cda33ae60f9f2346

Observation ef8e2c66-11bc-490a-96b2-52c7ef4f4a05 · outbound

This paper cites an unresolved cited work.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:42:35.060820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:42:22.923453Z digest=sha256:174a68b33ae7ada577746444723a3b8fb6dae237286b898bed3cc0ad97e9917f

Observation 86d4f465-4b4c-4223-91a8-df8860c40727 · outbound

This paper cites an unresolved cited work.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:42:34.892989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:42:23.014155Z digest=sha256:31710821a294150e478bc252441cdfecb6b663f34210b701225397ed0bce0e64

Observation d753f713-4f60-428d-adfe-305ab0f63ce4 · outbound

This paper cites an unresolved cited work.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:42:34.737803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:42:23.136683Z digest=sha256:8ef2279acfe9e0df55a9b1fc687d7ca3b84c54d53143f49b07aa7699acdf0633

Observation e0f2a2b3-80f0-445d-ad20-e698540706d5 · outbound

This paper cites O1 Replication Journey: A Strategic Progress Report -- Part 1.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models O1 Replication Journey: A Strategic Progress Report -- Part 1

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:23.212775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:23.212775Z digest=sha256:3472bfa4c5bde1ebcf162335ead4b5e4864a21f3e146460dbb0d4ac553c1c482

Observation c52ef2d4-c7d1-47c7-8cb6-135021e83135 · outbound

This paper cites Quantifying the Reasoning Abilities of LLMs on Real-world Clinical Cases.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Quantifying the Reasoning Abilities of LLMs on Real-world Clinical Cases

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:23.341138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:23.341138Z digest=sha256:fc6ab3e24bf28bf7435be40590ae5d11b023b9613e87cd4bae931ccc17f44516

Observation 7b82acba-893c-41df-b3ed-d825fbb530bc · outbound

This paper cites an unresolved cited work.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:42:34.518832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:42:23.423821Z digest=sha256:3e32a45cb62a454d272cd1143a7d64f2870b1e7f33f586b0c8746cceadfafedc

Observation 143ce6d6-de68-44ed-84b8-5b90c3a4a49a · outbound

This paper cites Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:23.541553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:23.541553Z digest=sha256:8cfc4ce3b5f95169be882316e47146b529cf05153b5ee1adf24e933a7197c8f3

Observation 03d0adc0-8f96-4fce-9a3e-a7ade274a3dd · outbound

This paper cites Qwen2.5 Technical Report.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Qwen2.5 Technical Report

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:23.635382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:23.635382Z digest=sha256:60677e8b25394f8b49e814df53322728dfc54bab01650c0cf5e34bab9db4a0d9

Observation dfb2780e-d949-4312-b15e-44740ee94f1b · outbound

This paper cites an unresolved cited work.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Unresolved cited work

Reference 32

Resolution
parse uncertain
raw_fallback, observed 2026-08-07T15:42:34.329538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:42:23.729601Z digest=sha256:06669737e58a5560cb86094ca72e2ba8882866b7bc72aac7d41d9de7f2fdf8de

Observation b4bf1c95-d465-4baf-bf40-c34333cbcdc8 · outbound

This paper cites an unresolved cited work.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:42:34.165599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:42:23.825815Z digest=sha256:64a74f281a69307186bc76b740d7d1254055721d36371e15e4cba3ca9207aa48

Observation ec7c2477-4f05-4195-bce0-ed7d2be78f22 · outbound

This paper cites an unresolved cited work.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:42:33.999351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:42:23.921620Z digest=sha256:51b6e51c8d8459cf5b4acfb2a552300104dbfc6373510962a343b6b6a9d6b86d

Observation c8024d9a-767f-437d-91b5-064645d21feb · outbound

This paper cites Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:24.112367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:24.112367Z digest=sha256:b752d67768081c65a5a7679a06395c716f9e9fd3ef970b8f2770a752cd01cf8a

Observation 1f349e7e-9776-4a43-a14d-1feee6535dca · outbound

This paper cites CMB: A Comprehensive Medical Benchmark in Chinese.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models CMB: A Comprehensive Medical Benchmark in Chinese

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:24.190092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:24.190092Z digest=sha256:ade75aa655e29b5006ccf960b874b9993a7774a8e074f86c8e2c056137403a23

Observation 6ca70e4c-e4fc-44f5-9602-498d4d853e7c · outbound

This paper cites an unresolved cited work.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Unresolved cited work

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:24.277569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:24.277569Z digest=sha256:4865633130abe97de148a7e42ca313d8dded75963d7e46fe5a0bf4a35ed51e01

Observation b70375a0-142c-47ef-b790-23de3e044cce · outbound

This paper cites an unresolved cited work.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Unresolved cited work

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:24.437348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:24.437348Z digest=sha256:748d6b03c07ff40f556df18a094fe0317d58f5d48ed3e33dbbde8969174dcd84

Observation a1a2fb92-3829-4cc4-b075-bef9db82819a · outbound

This paper cites an unresolved cited work.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:42:33.625676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:42:24.591939Z digest=sha256:0bd8f953c6e62c3853c8a0051dcfed0b5624ffb36de9180f759c0cfffcedd7ae

Observation c0d1f68f-f404-453c-94e5-88b1a5f37ff6 · outbound

This paper cites an unresolved cited work.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Unresolved cited work

Reference 40

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:42:33.452128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:42:24.700756Z digest=sha256:a52a6bdd82eb7476e6d27f8274606e4589c9cae533d0382ba3fcbf49b4e27c20

Observation 8e0f9c04-a308-4aa7-ab73-06a7f94f0cf9 · outbound

This paper cites FineMedLM-o1: Enhancing Medical Knowledge Reasoning Ability of LLM from Supervised Fine-Tuning to Test-Time Training.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models FineMedLM-o1: Enhancing Medical Knowledge Reasoning Ability of LLM from Supervised Fine-Tuning to Test-Time Training

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:24.782003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:24.782003Z digest=sha256:0a01fd893fc01f7f54cd6471e9f9697fd5b7102b1899bf93b108a60a6c714264

Observation 9ecad0b3-74e9-43dd-a2f4-f0a0ed2b5f08 · outbound

This paper cites o1-Coder: an o1 Replication for Coding.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models o1-Coder: an o1 Replication for Coding

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:24.855422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:24.855422Z digest=sha256:da41582bdb65017de1ee9956f1dd003949d03597c57fc5903e9040793e5a9d4b

Observation bdbf2e02-346c-4d5c-80a6-1ea18350862f · outbound

This paper cites an unresolved cited work.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:42:33.281918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:42:24.949716Z digest=sha256:9765d4f8f02be7d5b05da0b173756e9c5ab64bed4678def754489351314738cb

Observation ee58effb-e067-41e3-bf66-4c5f270dcafe · outbound

This paper cites MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:25.004807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:25.004807Z digest=sha256:56984a2078a77aee3a79f0919576b9502fccd572be12ccdfe2ff1cc60eab9b66

Observation 4f9f9bad-5ba3-4f51-a0fb-1237c3b00e20 · outbound

This paper cites an unresolved cited work.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Unresolved cited work

Reference 45

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:42:33.096586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:42:25.067031Z digest=sha256:2ae566c18e1effd4b9b0707a554e754e414ece2a202764db99a3a09362010534

Observation cb762cc9-3243-49ea-adb5-5a1fa467d0ce · outbound

This paper cites - Use appropriate Markdown syntax for headings based on the original heading hierarchy (e.g., ‘#‘, ‘##‘, ‘###‘).

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models - Use appropriate Markdown syntax for headings based on the original heading hierarchy (e.g., ‘#‘, ‘##‘, ‘###‘)

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:42:32.955211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:42:25.150270Z digest=sha256:24cc59a4c6d3b887104ca67b117bd54bd3285e22c3c78fcb1c5939df3443805d

Observation cbd173d9-fbc5-4850-9362-586f09d2e2d9 · outbound

This paper cites - Completely delete citations and footnotes, including their in-text markers.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models - Completely delete citations and footnotes, including their in-text markers

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:42:32.770494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:42:25.216414Z digest=sha256:0bea98f43306c56e62aabb6b9d12fc9740345129c1bfa14f6a96143064eca852

Observation f00fc92c-ce94-42eb-95ef-b59b375244f5 · outbound

This paper cites an unresolved cited work.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Unresolved cited work

Reference 51

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:42:32.596816Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:42:25.292911Z digest=sha256:cc8114d76c9c7f2543d5e2d6e5aa0385b40879566c9970b1b304462da5056edc

Observation 0bf3ae46-805e-4e72-a9d4-2a191cdb0763 · outbound

This paper cites - Only perform formatting and cleaning adjustments without modifying the original content.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models - Only perform formatting and cleaning adjustments without modifying the original content

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:42:32.345656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:42:25.349954Z digest=sha256:e39b6bf719cdc7e0f49c8269fe2a62f2e7561bf00fc239fcc42fa38646a27773

Observation 3b025b0b-76e8-4009-a527-6f4414c86a56 · outbound

This paper cites an unresolved cited work.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Unresolved cited work

Reference 53

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:42:31.993742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:42:25.419255Z digest=sha256:18030ac9215f87d4bcb0fd70f1a17200e029b7681968e64bd7a228f965dfc938

Observation fe1a29ce-ff78-487b-b7e8-383fd5d81997 · outbound

This paper cites - **Prohibited Content**: Any direct diagnostic statements involving disease names.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models - **Prohibited Content**: Any direct diagnostic statements involving disease names

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:42:31.554814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:42:25.491524Z digest=sha256:e9077fe929bcb95ed9720b3e0bfd4a40012da5b09021375abafdd0bc87867a2f

Observation cc00d117-d35b-460e-9deb-1529740beb7b · outbound

This paper cites an unresolved cited work.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Unresolved cited work

Reference 55

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:42:31.228417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:42:25.553291Z digest=sha256:89836266ff81a53bf922545a0c0897ed99c143a97d9b70680fb7483d3ac3d9a9

Observation 7ec7ea28-49a4-404e-84cf-37fc9aa3e8ac · outbound

This paper cites an unresolved cited work.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Unresolved cited work

Reference 56

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:42:31.015175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:42:25.614692Z digest=sha256:dae6e2e3abb2325e4c866ac2321684bf2ed59600a7ddbfca5c2370d8c4f9dd20

Observation 8ca475ba-f535-4936-94d2-b5a53b8afc8b · outbound

This paper cites DiagnosisArena-MCQ Evaluation Prompt You are an expert in the field of rare diseases.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models DiagnosisArena-MCQ Evaluation Prompt You are an expert in the field of rare diseases

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:42:30.685315Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:42:25.692156Z digest=sha256:330a2c703d61e0a0b87ef22d82142d38ece575aae9718215c46db45036ab6fd6

Observation a19227c1-4dd3-450b-8594-6b35a146e585 · outbound

This paper cites an unresolved cited work.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Unresolved cited work

Reference 58

Resolution
parse uncertain
raw_fallback, observed 2026-08-07T15:42:30.306619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:42:25.769469Z digest=sha256:6acff9eaa9efbc8ad2fb0f32a06959969c5c9807cf56b98bd60d50eb3c769207

Observation 0ae94c43-e39a-435b-b59a-1fd19ee19465 · outbound

This paper cites C Cases C.1 Benchmark Cases In this case, Case Information, Physical Examination, and Diagnostic Tests are components of the patient’s medical record.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models C Cases C.1 Benchmark Cases In this case, Case Information, Physical Examination, and Diagnostic Tests are components of the patient’s medical record

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:42:30.090590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:42:25.833573Z digest=sha256:30e683cd7ff19fff9da7364e671218fb792a17791912ed97b5bd4e4ac934c14b

Observation 83fdf6e5-502b-413a-8de7-2fbffa51c1bf · outbound

This paper cites Highly mobile, filamentous, causing embolic symptoms or arrhythmias.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Highly mobile, filamentous, causing embolic symptoms or arrhythmias

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:42:29.800933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:42:25.898275Z digest=sha256:680aa2e93077f518f8a5121050abbd9b521827ae0356e2557256ad34802f1055

Observation 406b80e0-21cc-494b-a5d9-0dee71f9d2ab · outbound

This paper cites Could be in LVOT, leading to similar symptoms.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Could be in LVOT, leading to similar symptoms

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:42:29.553124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:42:25.969036Z digest=sha256:a6564c7b2bc3dd9683310a5d4e8abace3538d28d62db843a76a8c03f861212c2

Observation dda23e77-1902-4c30-988e-5ec19956f861 · outbound

This paper cites Mobile and can cause obstruction.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Mobile and can cause obstruction

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:42:29.300596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:42:26.033021Z digest=sha256:0be64f4702e9b65a8b3de323e82f63bbe8999e77474ca563faa20cf0daf5c900

Observation a5b39362-9d2f-44e3-8ac8-e07a2071d6d2 · outbound

This paper cites an unresolved cited work.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Unresolved cited work

Reference 63

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:42:28.890974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:42:26.090955Z digest=sha256:4a02e7fa56b40c16ed1b669c85faaadd16133e2587bf96aa7aae12f7d5627338

Observation cac5c5c8-f692-4bac-abe8-81a1ebf3cb64 · outbound

This paper cites Wait, but the movement pattern and CT findings might make fibroelastoma more likely than Lambl’s.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Wait, but the movement pattern and CT findings might make fibroelastoma more likely than Lambl’s

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:42:28.599013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:42:26.152160Z digest=sha256:c0ce6d9315b161e1b9296185e0c0071f03d10a43b7a3630df99869c3d2f04f50

Observation 9b4dd91b-2b5a-48e0-a433-31f02bff4043 · outbound

This paper cites The key is the mobile mass in LVOT causing possible embolic events (leading to AFib) or obstruction.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models The key is the mobile mass in LVOT causing possible embolic events (leading to AFib) or obstruction

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:42:27.907921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:42:26.229267Z digest=sha256:fb4dab1e06e7da1ef31ca912753e8a70bd7a99d1556db84d2a93e4216faf09b7

Observation d33dae21-4284-4cfc-8e96-560b5b0f64f2 · outbound

This paper cites an unresolved cited work.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Unresolved cited work

Reference 67

Resolution
parse uncertain
raw_fallback, observed 2026-08-07T15:42:28.214479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:42:26.328940Z digest=sha256:1e003fa763a69d9f9e0416c5a0f7cfd6d9438b9bd476ff21394b20574fa003aa

Observation 5ffe65f1-7028-494e-b52a-c5be7314a30a · outbound

This paper cites an unresolved cited work.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Unresolved cited work

Reference 68

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:42:27.623039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:42:26.489072Z digest=sha256:e0b70e6123bdac2ffbbdf5e454a903620b368e24e00928a2df245a1c9ec18e95

Observation bde34791-774d-430d-b597-946b2915abad · outbound

This paper cites Measuring Massive Multitask Language Understanding.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Measuring Massive Multitask Language Understanding

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:21.630549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:21.630549Z digest=sha256:e23be300a4cd9bfaca9280b24779571d5e7a7d59f88bcf6da576cf04a5bd3393

Observation 5ef940ea-ed06-44c9-ae68-bc31d44abad0 · outbound

This paper cites Advances in neural information processing systems, 35:24824–24837.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Advances in neural information processing systems, 35:24824–24837

Reference 2022

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:42:33.810735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:42:24.508298Z digest=sha256:a86accc974bcecee77f631583d7cc05e2278ebadc06e0b9bdf8ee9bfae7ce57a

Observation 4c32e658-c443-46e1-aeed-e7fb5d426ae3 · outbound

This paper cites Baichuan-M1: Pushing the Medical Capability of Large Language Models.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Baichuan-M1: Pushing the Medical Capability of Large Language Models

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:24.011816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:24.011816Z digest=sha256:b4cd70c8ad6157a361f5637af135bba7d5ef07d46fee4c256e329b8fe24094b4

Pith citing papers

Observation 09047aae-f843-40e3-85b4-cf933f609390 · inbound

Beyond Classification Accuracy: Neural-MedBench and the Need for Deeper Reasoning Benchmarks cites this paper.

Beyond Classification Accuracy: Neural-MedBench and the Need for Deeper Reasoning Benchmarks DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-18T13:31:25.606181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-18T13:26:58.309566Z digest=sha256:c98249307e1693676998a295092018910ffb692c5cc8c4e3b9cf47ce29bc2836

Observation b7ec127c-9cc0-44a7-94c9-86527838c0db · inbound

MedXIAOHE: A Comprehensive Recipe for Building Medical MLLMs cites this paper.

MedXIAOHE: A Comprehensive Recipe for Building Medical MLLMs DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models

Reference 76

Resolution
verified exact
arxiv_id, observed 2026-05-15T22:56:50.369713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-15T22:52:30.992054Z digest=sha256:511dec506ff5ecf13ef4d11e6a482d1599aca02f139b956ee96a318d0dfc0332

Observation 56606448-0adb-4b38-bf1b-6c8a2b8894c4 · inbound

From Exposure to Internalization: Dual-Stream Calibration for In-context Clinical Reasoning cites this paper.

From Exposure to Internalization: Dual-Stream Calibration for In-context Clinical Reasoning DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:10:50.561904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T19:19:15.053725Z digest=sha256:5649117d72de59feff54ed86f4155e827bc513964a3a31832a435e595ae66086

Observation eb509866-ada7-44b7-a024-1d689965352d · inbound

EpiGraph: Building Generalists for Evidence-Intensive Epilepsy Reasoning in the Wild cites this paper.

EpiGraph: Building Generalists for Evidence-Intensive Epilepsy Reasoning in the Wild DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:41:37.715348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-12T04:03:34.773345Z digest=sha256:f2d4e44806f73448d822466d081d54d74c6d0dec198e3a97182a10437496b7cc

Observation 4c17d835-2f8e-4172-ae51-856250113229 · inbound

EpiGraph: Building Generalists for Evidence-Intensive Epilepsy Reasoning in the Wild cites this paper.

EpiGraph: Building Generalists for Evidence-Intensive Epilepsy Reasoning in the Wild DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:27:59.758936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-14T21:25:31.662152Z digest=sha256:06f7e244269434c69c7c9aa60f7a57b2da41764f31e870569989f7e206a6cb13

Observation 7eb391fb-79aa-4e59-ba84-88b6f2ba5083 · inbound

EHRBench: An Automated and Reliable EHR-based Benchmark for Clinical Decision Making with LLMs cites this paper.

EHRBench: An Automated and Reliable EHR-based Benchmark for Clinical Decision Making with LLMs DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models

Reference 131

Resolution
verified exact
arxiv_id, observed 2026-06-29T06:43:10.455707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-29T06:41:06.828814Z digest=sha256:eb45db5e3a66246d569796bb4ba7c3b6d26b6bb21f78470656cc9e30573f0630

Observation 96c78045-c903-4f3b-afa0-f6f2d46e03af · inbound

Beyond English benchmarks: clinical llm evaluation in Brazilian Portuguese cites this paper.

Beyond English benchmarks: clinical llm evaluation in Brazilian Portuguese DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-07-02T19:07:17.921236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-27T21:40:16.052239Z digest=sha256:05b43425e95ccaf8d9a9119c9280f4e44ec13c69af49f9237687fe9d4662d6f7

Observation d98ffec1-12da-4a37-ab73-142cbea8de42 · inbound

RareDxR1: Autonomous Medical Reasoning for Rare Disease Diagnosis Beyond Human Annotation cites this paper.

RareDxR1: Autonomous Medical Reasoning for Rare Disease Diagnosis Beyond Human Annotation DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-07-02T19:27:18.624793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-07-02T19:21:44.653877Z digest=sha256:711008c2c8663bb67bc91340e1b4d2ded8586454d7af4a59c2032d2d926c726d

Observation fbbe7ad0-62de-42bd-a871-203c3e9afd2d · inbound

DeepLens Diagnosis Agent: Agentic Workflow Design Lets a Small Reasoning Model Compete with Frontier LLMs cites this paper.

DeepLens Diagnosis Agent: Agentic Workflow Design Lets a Small Reasoning Model Compete with Frontier LLMs DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-02T13:38:11.263118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:38:11.263118Z digest=sha256:ddeaaf5c4c32f3a452d5a7164a8cbc9f5c5cea4e839fbfb31c9837a7ea195f35

Observation 3530e665-cd62-480a-96b0-5471f5a47e1d · inbound

Evaluating Multi-Turn Multimodal Diagnostic Reasoning on Challenging Real-World Clinical Cases cites this paper.

Evaluating Multi-Turn Multimodal Diagnostic Reasoning on Challenging Real-World Clinical Cases DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-01T01:07:59.356644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T01:07:59.356644Z digest=sha256:009dc3c2a7afc6934c1c3c4a853b6874d2240950555dded7642d76f698ea12b0