Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T13:55:57.054222Z
Paper Citation Record · LEDGER
As of 14 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 29 inbound Pith citation observations for arXiv:2505.23802.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T13:55:57.054222Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-11T04:18:04.903375Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T02:28:24.338817Z
27 of 27 outbound references displayed
External citation measurements
9
pith, observed 2026-08-05T02:28:24.338817Z
Observation 7b74bc32-96c1-4b7f-8669-1bc0d9968802 · outbound
MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks https://paperswithcode.com/sota/question-answering-on-medqa-usmle
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 568138c2-2b31-4ee5-a09a-b0700b205b85 · outbound
MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks Health Services Research and Managerial Epidemiology11, 23333928241234863 (2024) https://doi.org/10.1177/23333928241234863
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 7b3a2322-be3d-46d0-924c-858f51700a29 · outbound
MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks HealthTech Magazines (2024)
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 7d2b719d-8dc4-45c8-8fe1-b5924d099dd6 · outbound
MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks BJU International (2025) https://doi.org/10.1111/bju.16676
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 61c56e8e-d78d-48f4-b1d8-87f74ddf7fc9 · outbound
MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks Capabilities of GPT-4 on Medical Challenge Problems
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 61ba22e9-9f08-4106-a1fe-02714bf0ac59 · outbound
MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks NEJM AI2(2) (2025) https://doi.org/10.1056/AIe2401235
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 69ff72cd-ced6-40ec-aeea-fdba3dc77e35 · outbound
MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks Unresolved cited work
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 1b78878b-a970-4333-b291-6e3e91eb25ca · outbound
MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks JAMA333(4), 319–328 (2025) https://doi.org/10.1001/jama.2024.21700
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0764be9c-7c4e-4c8e-bd99-f59e60c3cc90 · outbound
MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks Nature Medicine30, 2613–2622 (2024)
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 53de62e3-c945-498e-a4cf-6801f6846bde · outbound
MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks https://cdn.openai.com/pdf/ 27 bd7a39d5-9e9f-47b3-903c-8b847ca650c7/healthbench_paper.pdf
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation ee5a201a-ea4d-4a08-898e-a9ed49716b6a · outbound
MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks Transactions on Machine Learning ResearchTMLR(2023)
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 97519918-38ee-4b73-a07e-8445f8444089 · outbound
MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks medRxiv (2025) https://doi.org/10.1101/2025.04.22.25326219
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f6d0eac4-e0cd-498c-bd2c-e20e8168f87f · outbound
MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks What Disease does this Patient Have? A Large-scale Open Domain Question Answering Dataset from Medical Exams
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c9a08a34-5230-429a-84b7-f606d98e64af · outbound
MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks Context Clues: Evaluating Long Context Models for Clinical Prediction Tasks on EHRs
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4555d654-60d5-4db3-8eee-3727717d765a · outbound
MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 39d9ae3e-d96b-486d-8980-1520ef343181 · outbound
MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks Large Language Models in the Clinic: A Comprehensive Benchmark
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8be8adeb-4f02-43d0-82ef-6a01a5923047 · outbound
MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks Nature Communications15(1), 8384 (2024)
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation dedf2ba1-a605-4bae-82a4-68ccf9f18939 · outbound
MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks BigBIO: A Framework for Data-Centric Biomedical Natural Language Processing
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 292e2100-bec1-4525-b1d0-377a389f695c · outbound
MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks A Survey on LLM-as-a-Judge
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5cfad890-8fb4-4fca-b335-0d262294c47f · outbound
MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks Evidently AI
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation bc38713e-f5ab-4a85-86d9-645a25d093ce · outbound
MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks Quantifying Variance in Evaluation Benchmarks
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 92e6df23-cbb8-4759-af2d-1c5dc548eb64 · outbound
MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks https://arxiv.org/abs/2404
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 24f9a1af-fbe8-479d-bf54-d98ce4e236d5 · outbound
MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks https: //github.com/confident-ai/deepeval
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 64b6275b-cf4b-4de5-9aed-525cc8e0aa02 · outbound
MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks SelfCheckGPT: Zero-Resource Black-Box Hallucination Detection for Generative Large Language Models
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4e2ebdfd-c285-49c6-9ae9-70387e27c35c · outbound
MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks FActScore: Fine-grained Atomic Evaluation of Factual Precision in Long Form Text Generation
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6d7d5f55-4cf7-4690-8208-b40292fd357d · outbound
MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks MPRA Paper No
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation d093a4a1-64ad-4481-bf40-2ec07e16a1a7 · outbound
MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks Nature Medicine30(4), 1134–1142 (2024) https://doi.org/ 10.1038/s41591-024-02855-5 31
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 45f4b5c5-feec-4703-b60f-3914bc00a949 · inbound
Introducing Answered with Evidence -- a framework for evaluating whether LLM responses to biomedical questions are founded in evidence MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4cd591a4-3e82-4aee-904a-b4e9ae152850 · inbound
A Novel Evaluation Benchmark for Medical LLMs: Illuminating Safety and Effectiveness in Clinical Domains MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 11e99d4d-feee-4287-ac54-ff0cd2a5ef8f · inbound
A global log for medical AI MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc95d819-ab69-4289-91ed-443b94ba985f · inbound
Scalable Stewardship of an LLM-Assisted Clinical Benchmark with Physician Oversight MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation fddf0ff6-8b95-4fdb-ae1e-8ecceac2429b · inbound
Automatic Replication of LLM Mistakes in Medical Conversations MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation efeacace-d7dc-428f-b398-55a72c9469c7 · inbound
Using LLM-as-a-Judge/Jury to Advance Scalable, Clinically-Validated Safety Evaluations of Model Responses to Users Demonstrating Psychosis MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 176be442-07bd-44c1-a6a2-1d62c26af238 · inbound
How Robust Are Large Language Models for Clinical Numeracy? An Empirical Study on Numerical Reasoning Abilities in Clinical Contexts MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 327c5508-4d08-437b-88e3-682553db506d · inbound
Green Shielding: A User-Centric Approach Towards Trustworthy AI MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 6e0e4844-4077-49ab-a851-21f86558be09 · inbound
SymptomAI: Toward a Conversational AI Agent for Everyday Symptom Assessment MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation d561dc02-1a1b-4291-9152-86bdca298f24 · inbound
SymptomAI: Toward a Conversational AI Agent for Everyday Symptom Assessment MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 1cb41f50-496a-47e3-abb4-b9fbe9bf493a · inbound
Measuring What Matters: Benchmarking Generative, Multimodal, and Agentic AI in Healthcare MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation e8838776-c841-4e93-b451-6ca2540bee98 · inbound
Event Fields: Learning Latent Event Structure for Waveform Foundation Models MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation aad9683c-c57b-4bdb-b23b-a2ceec53f2c0 · inbound
CLR-voyance: Reinforcing Open-Ended Reasoning for Inpatient Clinical Decision Support with Outcome-Aware Rubrics MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 5a1ce8a6-9fc1-44fc-a5ce-6d63febeb983 · inbound
CodeClinic: Evaluating Automation of Coding Skills for Clinical Reasoning Agents MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation c6de4e66-d417-4eeb-aa60-cd40fb5d06c2 · inbound
WISTERIA: Learning Clinical Representations from Noisy Supervision via Multi-View Consistency in Electronic Health Records MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation fa76b768-a8be-47a8-b2fb-57d57c8c3083 · inbound
Instructions Shape Production of Language, not Processing MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks
Reference 121
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation d81f6e33-e0ee-4a38-82fb-3d17b56c1898 · inbound
Instructions Shape Production of Language, not Processing MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks
Reference 121
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 744597a7-4cad-4739-9e57-cd78a33353af · inbound
CHI-Bench: Can AI Agents Automate End-to-End, Long-Horizon, Policy-Rich Healthcare Workflows? MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 191a7621-23b8-420d-ac62-46b885acc01f · inbound
AURORA: Contextual Orthogonalization for Geometric Representation Learning in Healthcare Foundation Models MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 09695644-9950-4a37-a589-a75502a21ca0 · inbound
AutoMedBench: Towards Medical AutoResearch with Agentic AI Models MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 62754d03-7ef2-4445-b55c-ad591e931ff6 · inbound
Beyond English benchmarks: clinical llm evaluation in Brazilian Portuguese MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 26c2190a-afd8-435c-a37b-521b4d7a93bf · inbound
CLExEval: A Human-in-the-Loop Framework for Qualitative Evaluation of LLM Clinical Reasoning MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation d830393c-29e0-477b-af91-f4f255af879b · inbound
A rubric-based controlled comparison of frontier language models on expert-authored clinical reasoning tasks MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 4975a1b2-199d-4813-9899-ab15e37aa937 · inbound
Aligning Clinical Needs and AI Capabilities: A Survey on LLMs for Medical Reasoning MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks
Reference 156
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation c54336c8-775b-4811-a61c-ccb61518a1c0 · inbound
MentalHospital: A Virtual Environment for Evaluating Psychiatric Clinical Encounters MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation da873bf5-463f-41dd-8348-386ac3616068 · inbound
The Path to Self-Evolving Clinical Systems: Scaling Medical Agents from Assistance to Autonomy MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b7b69def-859d-41e1-bfc4-b7e23a9c5660 · inbound
MedDDC-Eval: Diagnosis-Decoupled Evaluation of Multi-Turn Medical Consultation Agents MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98e3ffa5-fa46-45a7-86ad-a53a6fa0da2a · inbound
MedPRESS: A Multi-turn Benchmark for Patient-Pressure-Induced Medical Sycophancy in LLMs MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9abbd290-9d57-47a7-b1ab-ed397e779b6f · inbound
CliniCARE-Bench: Clinical Calibrated Audit of Medical Reasoning in EHR MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.