Pith. sign in

Paper Citation Record · LEDGER

MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks

As of 14 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 29 inbound Pith citation observations for arXiv:2505.23802.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.23802 v2

Coverage vector

measured 27 of 27 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:55:57.054222Z

measured 56 of 56 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 29 of 29 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T04:18:04.903375Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

27 of 27 outbound references displayed

  • verified exact2
  • verified fuzzy10
  • unresolved14
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

9
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 7b74bc32-96c1-4b7f-8669-1bc0d9968802 · outbound

This paper cites https://paperswithcode.com/sota/question-answering-on-medqa-usmle.

MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks https://paperswithcode.com/sota/question-answering-on-medqa-usmle

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:56:00.250449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T13:55:55.037178Z digest=sha256:1a6d84d6b790e39f30094ef9a0b5d62e49940c47428cfcd91f904c2cf8161da8

Observation 568138c2-2b31-4ee5-a09a-b0700b205b85 · outbound

This paper cites Health Services Research and Managerial Epidemiology11, 23333928241234863 (2024) https://doi.org/10.1177/23333928241234863.

MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks Health Services Research and Managerial Epidemiology11, 23333928241234863 (2024) https://doi.org/10.1177/23333928241234863

Reference 2

Resolution
verified exact
doi, observed 2026-08-07T13:55:57.476349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T13:55:55.109132Z digest=sha256:f1cbccc7c5d930b56fd4c7454b61cb08652b74b9654c75640dbbeada4dd39bdb

Observation 7b3a2322-be3d-46d0-924c-858f51700a29 · outbound

This paper cites HealthTech Magazines (2024).

MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks HealthTech Magazines (2024)

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:56:00.063684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T13:55:55.154334Z digest=sha256:0ad31e7c3af9e356074d7661d9529f10902063752f5a243d2b0d9957861db15d

Observation 7d2b719d-8dc4-45c8-8fe1-b5924d099dd6 · outbound

This paper cites BJU International (2025) https://doi.org/10.1111/bju.16676.

MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks BJU International (2025) https://doi.org/10.1111/bju.16676

Reference 4

Resolution
verified exact
doi, observed 2026-08-07T13:55:57.288067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T13:55:55.201731Z digest=sha256:968290d15c9daa924c1c18dfb773ae6d62fb4ea6634b1d960296e4de4f383c3e

Observation 61c56e8e-d78d-48f4-b1d8-87f74ddf7fc9 · outbound

This paper cites Capabilities of GPT-4 on Medical Challenge Problems.

MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks Capabilities of GPT-4 on Medical Challenge Problems

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T13:55:55.235160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:55:55.235160Z digest=sha256:21e95cd4f52c7fe1759c071cbf0bd3ad89f4bb56a12b8cd5bf4b7e20036f152b

Observation 61ba22e9-9f08-4106-a1fe-02714bf0ac59 · outbound

This paper cites NEJM AI2(2) (2025) https://doi.org/10.1056/AIe2401235.

MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks NEJM AI2(2) (2025) https://doi.org/10.1056/AIe2401235

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T13:55:55.293538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:55:55.293538Z digest=sha256:e313e46952b7496f2c930b463dbf25e7bf6780dba54ab8ab883a8db5b9ac9adb

Observation 69ff72cd-ced6-40ec-aeea-fdba3dc77e35 · outbound

This paper cites an unresolved cited work.

MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:55:59.861448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T13:55:55.401905Z digest=sha256:c529ed93cdc9e7be41764a901d712d271399221e5c7ffb3fb16dfdc3ba82e175

Observation 1b78878b-a970-4333-b291-6e3e91eb25ca · outbound

This paper cites JAMA333(4), 319–328 (2025) https://doi.org/10.1001/jama.2024.21700.

MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks JAMA333(4), 319–328 (2025) https://doi.org/10.1001/jama.2024.21700

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T13:55:55.482216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:55:55.482216Z digest=sha256:fd77727d89c7bf395d54072bdc92c8af2655ec52523ae151dd54f6fa6802f95f

Observation 0764be9c-7c4e-4c8e-bd99-f59e60c3cc90 · outbound

This paper cites Nature Medicine30, 2613–2622 (2024).

MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks Nature Medicine30, 2613–2622 (2024)

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:55:59.596835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T13:55:55.559160Z digest=sha256:66c739577b18bccd4b37b5d0366529b47055f8b11b9d83055d6bcdc6783eb31e

Observation 53de62e3-c945-498e-a4cf-6801f6846bde · outbound

This paper cites https://cdn.openai.com/pdf/ 27 bd7a39d5-9e9f-47b3-903c-8b847ca650c7/healthbench_paper.pdf.

MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks https://cdn.openai.com/pdf/ 27 bd7a39d5-9e9f-47b3-903c-8b847ca650c7/healthbench_paper.pdf

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:55:59.328508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T13:55:55.647958Z digest=sha256:4f9eee3131791a2ff9b7459082177c6db7eabe6a63a80febe9202ef8e8b3f9b4

Observation ee5a201a-ea4d-4a08-898e-a9ed49716b6a · outbound

This paper cites Transactions on Machine Learning ResearchTMLR(2023).

MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks Transactions on Machine Learning ResearchTMLR(2023)

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:55:59.077154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T13:55:55.729671Z digest=sha256:2c95fc8cd1458ca8192072fde3b0f26cbaaeb834400b8102890de289aac46913

Observation 97519918-38ee-4b73-a07e-8445f8444089 · outbound

This paper cites medRxiv (2025) https://doi.org/10.1101/2025.04.22.25326219.

MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks medRxiv (2025) https://doi.org/10.1101/2025.04.22.25326219

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T13:55:55.813709Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:55:55.813709Z digest=sha256:1f037cac669ab54bbebcf88716a779849829c65e9f00a6093f53a4bb5e4bac68

Observation f6d0eac4-e0cd-498c-bd2c-e20e8168f87f · outbound

This paper cites What Disease does this Patient Have? A Large-scale Open Domain Question Answering Dataset from Medical Exams.

MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks What Disease does this Patient Have? A Large-scale Open Domain Question Answering Dataset from Medical Exams

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T13:55:55.898652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:55:55.898652Z digest=sha256:24a945a38317d4ec237c15d574c609bb0260b2509b06065087a21c6f7868ac0a

Observation c9a08a34-5230-429a-84b7-f606d98e64af · outbound

This paper cites Context Clues: Evaluating Long Context Models for Clinical Prediction Tasks on EHRs.

MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks Context Clues: Evaluating Long Context Models for Clinical Prediction Tasks on EHRs

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T13:55:55.988534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:55:55.988534Z digest=sha256:89a8aaadf769ea204810b8f46c4f5775fe527b6993b1c38a74ebae361d4dc3ac

Observation 4555d654-60d5-4db3-8eee-3727717d765a · outbound

This paper cites Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models.

MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T13:55:56.095187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:55:56.095187Z digest=sha256:c92a97f7c2638223bf40a9211701670a195139dc5d176aba52884ca346226d7d

Observation 39d9ae3e-d96b-486d-8980-1520ef343181 · outbound

This paper cites Large Language Models in the Clinic: A Comprehensive Benchmark.

MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks Large Language Models in the Clinic: A Comprehensive Benchmark

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T13:55:56.187110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:55:56.187110Z digest=sha256:5d6d9a09d8ce274b956ac85d1127ca67515d2b1900335ce7c9dbac62bdb7c1d5

Observation 8be8adeb-4f02-43d0-82ef-6a01a5923047 · outbound

This paper cites Nature Communications15(1), 8384 (2024).

MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks Nature Communications15(1), 8384 (2024)

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:55:58.835599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T13:55:56.253770Z digest=sha256:d05b29728d5d07c62f690d69b02768bc15739315c6b5ebc57951cec2c48dcd61

Observation dedf2ba1-a605-4bae-82a4-68ccf9f18939 · outbound

This paper cites BigBIO: A Framework for Data-Centric Biomedical Natural Language Processing.

MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks BigBIO: A Framework for Data-Centric Biomedical Natural Language Processing

Reference 18

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T13:55:57.676197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T13:55:56.334863Z digest=sha256:39b3ede0ef06f95cffba2556e82d1efbbcfe78d79990aa41b0e4e9cdb4bfeee6

Observation 292e2100-bec1-4525-b1d0-377a389f695c · outbound

This paper cites A Survey on LLM-as-a-Judge.

MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks A Survey on LLM-as-a-Judge

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T13:55:56.416173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:55:56.416173Z digest=sha256:55071d2650414a6f88fdd4faa36ee0812ba18ee169523e29a2de2217acce70a6

Observation 5cfad890-8fb4-4fca-b335-0d262294c47f · outbound

This paper cites Evidently AI.

MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks Evidently AI

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:55:58.629428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T13:55:56.521765Z digest=sha256:8f45da4d07c785f3ddf97010ed775e6d637c72b4c168b58278bf3ea758710a76

Observation bc38713e-f5ab-4a85-86d9-645a25d093ce · outbound

This paper cites Quantifying Variance in Evaluation Benchmarks.

MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks Quantifying Variance in Evaluation Benchmarks

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T13:55:56.617692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:55:56.617692Z digest=sha256:d6b8321eba88ea203f850b4d00192c1e6457b562aa16cc13d312bb5c75849da0

Observation 92e6df23-cbb8-4759-af2d-1c5dc548eb64 · outbound

This paper cites https://arxiv.org/abs/2404.

MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks https://arxiv.org/abs/2404

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:55:58.423463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T13:55:56.705585Z digest=sha256:182219ebebbd30a1d18cd8c978286adcaa64a7a5ae1c5ba6dc9b8abbbfb7c907

Observation 24f9a1af-fbe8-479d-bf54-d98ce4e236d5 · outbound

This paper cites https: //github.com/confident-ai/deepeval.

MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks https: //github.com/confident-ai/deepeval

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:55:58.184629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T13:55:56.770801Z digest=sha256:bdaa6a873138c362733bf02a483fd7fee16b19779ec487599e1b426357946538

Observation 64b6275b-cf4b-4de5-9aed-525cc8e0aa02 · outbound

This paper cites SelfCheckGPT: Zero-Resource Black-Box Hallucination Detection for Generative Large Language Models.

MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks SelfCheckGPT: Zero-Resource Black-Box Hallucination Detection for Generative Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T13:55:56.873804Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:55:56.873804Z digest=sha256:3e4bde62deec7e8958b9863485ae3c7a5a89201a3f8deebad38836945a574cbf

Observation 4e2ebdfd-c285-49c6-9ae9-70387e27c35c · outbound

This paper cites FActScore: Fine-grained Atomic Evaluation of Factual Precision in Long Form Text Generation.

MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks FActScore: Fine-grained Atomic Evaluation of Factual Precision in Long Form Text Generation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T13:55:56.937334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:55:56.937334Z digest=sha256:120b999db8e1cbfad041abaa43fe66de13589d7e1abce0a4f801df3efdbded73

Observation 6d7d5f55-4cf7-4690-8208-b40292fd357d · outbound

This paper cites MPRA Paper No.

MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks MPRA Paper No

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:55:57.961204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T13:55:56.982200Z digest=sha256:16ac61a337e2cdec18d4a929d981c3ffab5ce11e07ea22bce9569bddfc01b30d

Observation d093a4a1-64ad-4481-bf40-2ec07e16a1a7 · outbound

This paper cites Nature Medicine30(4), 1134–1142 (2024) https://doi.org/ 10.1038/s41591-024-02855-5 31.

MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks Nature Medicine30(4), 1134–1142 (2024) https://doi.org/ 10.1038/s41591-024-02855-5 31

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T13:55:57.054222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:55:57.054222Z digest=sha256:27e7ac3ac806d27d5f59711f59d2a0ea4d072de7a594f9b99a340fda50d38a01

Pith citing papers

Observation 45f4b5c5-feec-4703-b60f-3914bc00a949 · inbound

Introducing Answered with Evidence -- a framework for evaluating whether LLM responses to biomedical questions are founded in evidence cites this paper.

Introducing Answered with Evidence -- a framework for evaluating whether LLM responses to biomedical questions are founded in evidence MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:53.792124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:53.792124Z digest=sha256:270f72509f688135548f1a9e5e487e4844b2308ddecc5ea0d7e5236af80c2380

Observation 4cd591a4-3e82-4aee-904a-b4e9ae152850 · inbound

A Novel Evaluation Benchmark for Medical LLMs: Illuminating Safety and Effectiveness in Clinical Domains cites this paper.

A Novel Evaluation Benchmark for Medical LLMs: Illuminating Safety and Effectiveness in Clinical Domains MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T10:46:10.515710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:46:10.515710Z digest=sha256:09f440d2984e66e37949e5262af6fbbcd3f95a5343464a68bfec76352367cc7d

Observation 11e99d4d-feee-4287-ac54-ff0cd2a5ef8f · inbound

A global log for medical AI cites this paper.

A global log for medical AI MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T11:34:03.183686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:34:03.183686Z digest=sha256:23496f6ca05138e8b29342546d6535f323b63e5929bd961445fd07a2512d404b

Observation cc95d819-ab69-4289-91ed-443b94ba985f · inbound

Scalable Stewardship of an LLM-Assisted Clinical Benchmark with Physician Oversight cites this paper.

Scalable Stewardship of an LLM-Assisted Clinical Benchmark with Physician Oversight MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-16T20:23:23.364363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-16T20:21:40.867354Z digest=sha256:e526642e5831f9e01fa948e19509696e39c7c7631b9e4865e0c7d858616fe3d1

Observation fddf0ff6-8b95-4fdb-ae1e-8ecceac2429b · inbound

Automatic Replication of LLM Mistakes in Medical Conversations cites this paper.

Automatic Replication of LLM Mistakes in Medical Conversations MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-16T20:28:24.177713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-16T20:25:24.722562Z digest=sha256:db8e8f8c146902fb9fce8cf4aa9d8af2fda0297b3de2327cd9d4e657c17dc0e7

Observation efeacace-d7dc-428f-b398-55a72c9469c7 · inbound

Using LLM-as-a-Judge/Jury to Advance Scalable, Clinically-Validated Safety Evaluations of Model Responses to Users Demonstrating Psychosis cites this paper.

Using LLM-as-a-Judge/Jury to Advance Scalable, Clinically-Validated Safety Evaluations of Model Responses to Users Demonstrating Psychosis MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-15T09:09:53.374917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T09:06:33.531027Z digest=sha256:11bc60471b6a973313902d59c82417e2e7f5d1aa1f702067615588c9ae3c8502

Observation 176be442-07bd-44c1-a6a2-1d62c26af238 · inbound

How Robust Are Large Language Models for Clinical Numeracy? An Empirical Study on Numerical Reasoning Abilities in Clinical Contexts cites this paper.

How Robust Are Large Language Models for Clinical Numeracy? An Empirical Study on Numerical Reasoning Abilities in Clinical Contexts MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T10:01:03.373551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T15:41:21.704608Z digest=sha256:4275eb79e18d39653c0dc583c4e27453dc2b0eaae1244e6623f05f67582087d6

Observation 327c5508-4d08-437b-88e3-682553db506d · inbound

Green Shielding: A User-Centric Approach Towards Trustworthy AI cites this paper.

Green Shielding: A User-Centric Approach Towards Trustworthy AI MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:56:24.527571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-08T03:43:54.896449Z digest=sha256:2d669fbefab20fbe8fc66c5d4078caddef69f40dc0ccbbd7defc2764d9d5b1fe

Observation 6e0e4844-4077-49ab-a851-21f86558be09 · inbound

SymptomAI: Toward a Conversational AI Agent for Everyday Symptom Assessment cites this paper.

SymptomAI: Toward a Conversational AI Agent for Everyday Symptom Assessment MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:46:53.503600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-07T16:17:39.923337Z digest=sha256:b2c0879e6e5d950a047c46261c0e9f5a01f219f2ac7d8f307fdb91112769c4ed

Observation d561dc02-1a1b-4291-9152-86bdca298f24 · inbound

SymptomAI: Toward a Conversational AI Agent for Everyday Symptom Assessment cites this paper.

SymptomAI: Toward a Conversational AI Agent for Everyday Symptom Assessment MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:41:42.864862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-12T02:21:16.825681Z digest=sha256:1244b50656634abd396c9604294a95cf1a4bd3ac2e2f4dd360f4820e5cff96d7

Observation 1cb41f50-496a-47e3-abb4-b9fbe9bf493a · inbound

Measuring What Matters: Benchmarking Generative, Multimodal, and Agentic AI in Healthcare cites this paper.

Measuring What Matters: Benchmarking Generative, Multimodal, and Agentic AI in Healthcare MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:56:36.287260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-12T01:26:28.419570Z digest=sha256:2c14958bed42ed821b32f87799e02f3ddcf00e70768dafc1625d7961d617af69

Observation e8838776-c841-4e93-b451-6ca2540bee98 · inbound

Event Fields: Learning Latent Event Structure for Waveform Foundation Models cites this paper.

Event Fields: Learning Latent Event Structure for Waveform Foundation Models MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:06:31.984899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-12T01:16:48.039349Z digest=sha256:7194498727e43beb8f784b38b2707a9572231df0951df336e2b5a6d33bf6aa62

Observation aad9683c-c57b-4bdb-b23b-a2ceec53f2c0 · inbound

CLR-voyance: Reinforcing Open-Ended Reasoning for Inpatient Clinical Decision Support with Outcome-Aware Rubrics cites this paper.

CLR-voyance: Reinforcing Open-Ended Reasoning for Inpatient Clinical Decision Support with Outcome-Aware Rubrics MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:11:23.506120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-12T04:32:16.930291Z digest=sha256:a8b2eb8bfeaab598a5483b6e303e3bf03a5af2f814baca6ddd798ca06deee8cf

Observation 5a1ce8a6-9fc1-44fc-a5ce-6d63febeb983 · inbound

CodeClinic: Evaluating Automation of Coding Skills for Clinical Reasoning Agents cites this paper.

CodeClinic: Evaluating Automation of Coding Skills for Clinical Reasoning Agents MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:41:46.226036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-12T02:17:41.650948Z digest=sha256:a6dba644452359ac7e0593b7f2fb4c69ab69a87167823b8ec63f156766789cae

Observation c6de4e66-d417-4eeb-aa60-cd40fb5d06c2 · inbound

WISTERIA: Learning Clinical Representations from Noisy Supervision via Multi-View Consistency in Electronic Health Records cites this paper.

WISTERIA: Learning Clinical Representations from Noisy Supervision via Multi-View Consistency in Electronic Health Records MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:41:45.088746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-12T02:17:54.498339Z digest=sha256:89c68577635b196bdaf3e3a8e63c4ad8233875a160de6cd7f87f206ad455131c

Observation fa76b768-a8be-47a8-b2fb-57d57c8c3083 · inbound

Instructions Shape Production of Language, not Processing cites this paper.

Instructions Shape Production of Language, not Processing MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks

Reference 121

Resolution
verified exact
arxiv_id, observed 2026-05-13T03:12:09.414953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-13T03:09:02.902912Z digest=sha256:c8c557e42bf6b17edaa4526ee1da560f324902f6f7ee7d99f38251604489c517

Observation d81f6e33-e0ee-4a38-82fb-3d17b56c1898 · inbound

Instructions Shape Production of Language, not Processing cites this paper.

Instructions Shape Production of Language, not Processing MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks

Reference 121

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:02:58.194611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-14T21:02:02.135970Z digest=sha256:d010f31f20e8f857561c6c00b575d88e68d06da90e87fbddc480b0379a43614f

Observation 744597a7-4cad-4739-9e57-cd78a33353af · inbound

CHI-Bench: Can AI Agents Automate End-to-End, Long-Horizon, Policy-Rich Healthcare Workflows? cites this paper.

CHI-Bench: Can AI Agents Automate End-to-End, Long-Horizon, Policy-Rich Healthcare Workflows? MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-20T17:48:48.910341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-20T17:45:02.896703Z digest=sha256:53a2e4aeddb02d3d99c6820322d66588c9be490a486ddc75712961bb05898ff1

Observation 191a7621-23b8-420d-ac62-46b885acc01f · inbound

AURORA: Contextual Orthogonalization for Geometric Representation Learning in Healthcare Foundation Models cites this paper.

AURORA: Contextual Orthogonalization for Geometric Representation Learning in Healthcare Foundation Models MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-20T12:08:15.521584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-20T12:06:33.875494Z digest=sha256:873af1a0003098b830f6fee4d46c05e101dab01020ff2c8c88e72cb8ffabe8e9

Observation 09695644-9950-4a37-a589-a75502a21ca0 · inbound

AutoMedBench: Towards Medical AutoResearch with Agentic AI Models cites this paper.

AutoMedBench: Towards Medical AutoResearch with Agentic AI Models MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-07-01T23:06:21.116582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-28T14:38:40.017263Z digest=sha256:56971d24ed1487af6d0823980a4af2c02a7b1b9860819c9aab62859e481fbcfa

Observation 62754d03-7ef2-4445-b55c-ad591e931ff6 · inbound

Beyond English benchmarks: clinical llm evaluation in Brazilian Portuguese cites this paper.

Beyond English benchmarks: clinical llm evaluation in Brazilian Portuguese MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-07-02T19:07:17.939001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-27T21:40:16.052239Z digest=sha256:75d53a71a19d4beefbedad64eb30836ca7e08cc22de102af903ce33d3e6874b0

Observation 26c2190a-afd8-435c-a37b-521b4d7a93bf · inbound

CLExEval: A Human-in-the-Loop Framework for Qualitative Evaluation of LLM Clinical Reasoning cites this paper.

CLExEval: A Human-in-the-Loop Framework for Qualitative Evaluation of LLM Clinical Reasoning MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-07-01T10:25:41.545608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-07-01T05:33:52.027771Z digest=sha256:62d3c24c8ea6bb156fd23e5215ad1b755e75fe0b117399b2cf2b6caae081cb4d

Observation d830393c-29e0-477b-af91-f4f255af879b · inbound

A rubric-based controlled comparison of frontier language models on expert-authored clinical reasoning tasks cites this paper.

A rubric-based controlled comparison of frontier language models on expert-authored clinical reasoning tasks MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T14:08:21.387402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-07-03T14:05:58.062268Z digest=sha256:8932e4ca6bfe19954afb9ccff3907aa27b0fd4b0f3580a589d105bf6d4f38f3f

Observation 4975a1b2-199d-4813-9899-ab15e37aa937 · inbound

Aligning Clinical Needs and AI Capabilities: A Survey on LLMs for Medical Reasoning cites this paper.

Aligning Clinical Needs and AI Capabilities: A Survey on LLMs for Medical Reasoning MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks

Reference 156

Resolution
verified exact
local_arxiv, observed 2026-07-10T18:57:31.509556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-07-10T18:50:22.827472Z digest=sha256:eb19ef027c7838cf19afc2a39da386ef6ce2846f548e559e271b743d0279f21e

Observation c54336c8-775b-4811-a61c-ccb61518a1c0 · inbound

MentalHospital: A Virtual Environment for Evaluating Psychiatric Clinical Encounters cites this paper.

MentalHospital: A Virtual Environment for Evaluating Psychiatric Clinical Encounters MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-07-10T10:37:01.669041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-07-10T10:30:27.256710Z digest=sha256:9c055ed98a09396b8da524908a924ed5dacad5aaa233a862e59725859a47cc4d

Observation da873bf5-463f-41dd-8348-386ac3616068 · inbound

The Path to Self-Evolving Clinical Systems: Scaling Medical Agents from Assistance to Autonomy cites this paper.

The Path to Self-Evolving Clinical Systems: Scaling Medical Agents from Assistance to Autonomy MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-14T06:30:16.612345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T06:30:16.612345Z digest=sha256:70a6530d3156c1fbfe4d59080fbc683a2136b914bdf014c2e4775f3eb7e5a246

Observation b7b69def-859d-41e1-bfc4-b7e23a9c5660 · inbound

MedDDC-Eval: Diagnosis-Decoupled Evaluation of Multi-Turn Medical Consultation Agents cites this paper.

MedDDC-Eval: Diagnosis-Decoupled Evaluation of Multi-Turn Medical Consultation Agents MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-01T13:50:41.494173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T13:50:41.494173Z digest=sha256:4cf50f4a1c1bf8cf9fecce0ec3aa7ea4e0e95ddc358b4887e9433751c0ef293d

Observation 98e3ffa5-fa46-45a7-86ad-a53a6fa0da2a · inbound

MedPRESS: A Multi-turn Benchmark for Patient-Pressure-Induced Medical Sycophancy in LLMs cites this paper.

MedPRESS: A Multi-turn Benchmark for Patient-Pressure-Induced Medical Sycophancy in LLMs MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T05:39:50.996391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T05:39:50.996391Z digest=sha256:41354d7bfbb3ecac86de72520dc485da27bbd6659d8711c470a7a4703d941d3e

Observation 9abbd290-9d57-47a7-b1ab-ed397e779b6f · inbound

CliniCARE-Bench: Clinical Calibrated Audit of Medical Reasoning in EHR cites this paper.

CliniCARE-Bench: Clinical Calibrated Audit of Medical Reasoning in EHR MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T04:18:04.903375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:18:04.903375Z digest=sha256:ede4b3adb5ece36fdbc03ad7cf763e68e955d4fc99f4f95545a9d5a542fff830