Pith. sign in

Paper Citation Record · LEDGER

OpenAIs HealthBench in Action: Evaluating an LLM-Based Medical Assistant on Realistic Clinical Queries

As of 14 August 2026, this Paper Citation Record lists 20 of 20 outbound references and 3 inbound Pith citation observations for arXiv:2509.02594.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.02594 v3

Coverage vector

measured 20 of 20 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T14:21:49.124883Z

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-30T13:13:26.156464Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

20 of 20 outbound references displayed

  • verified exact3
  • verified fuzzy5
  • unresolved9
  • parse uncertain0
  • malformed identifier3
  • metadata mismatch0

External citation measurements

2
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation d991f462-9a44-4780-8137-5e8a7f5e21eb · outbound

This paper cites A systematic review of large language model (LLM) evaluations in clinical medicine.

OpenAIs HealthBench in Action: Evaluating an LLM-Based Medical Assistant on Realistic Clinical Queries A systematic review of large language model (LLM) evaluations in clinical medicine

Reference 1

Resolution
malformed identifier
no resolver link, observed 2026-08-05T14:21:46.712398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:21:46.712398Z digest=sha256:b73dc610f735127c044604e94df2cf1c80a4a6fb2d3bcaeb11409aa1e649e751

Observation b79ed5a8-d09e-40b7-aef8-362eb4c21ff7 · outbound

This paper cites Scientific Evidence for Clinical Text Summarization Using Large Language Models: Scoping Review.

OpenAIs HealthBench in Action: Evaluating an LLM-Based Medical Assistant on Realistic Clinical Queries Scientific Evidence for Clinical Text Summarization Using Large Language Models: Scoping Review

Reference 3

Resolution
malformed identifier
no resolver link, observed 2026-08-05T14:21:46.901628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:21:46.901628Z digest=sha256:7ed4142fce1ba6dbae06315199da218f844672c496451fad30306a5978072700

Observation 290ee266-ba18-49c9-abab-4ea3e5857813 · outbound

This paper cites AI Consult: Real-World Evaluation of Large Language Model-Based Clinical Decision Support in Primary Care.

OpenAIs HealthBench in Action: Evaluating an LLM-Based Medical Assistant on Realistic Clinical Queries AI Consult: Real-World Evaluation of Large Language Model-Based Clinical Decision Support in Primary Care

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:21:52.173666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T14:21:47.097613Z digest=sha256:8de3ad907da7beb9c117adbb8259557a5b06de4fe4a08dfce0c0a7ec53ffa927

Observation 33444acb-a9b2-424e-a8c7-e527cfae02e1 · outbound

This paper cites Capabilities of GPT-4 on Medical Challenge Problems.

OpenAIs HealthBench in Action: Evaluating an LLM-Based Medical Assistant on Realistic Clinical Queries Capabilities of GPT-4 on Medical Challenge Problems

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T14:21:47.221633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:21:47.221633Z digest=sha256:dda7f623b7f1dad42a3abc6394d17756101ae08a9a63c877e7e2932f7e980084

Observation b7b98faa-fdc4-4107-9ed2-3157428ed2d3 · outbound

This paper cites Large Language Models Encode Clinical Knowledge.

OpenAIs HealthBench in Action: Evaluating an LLM-Based Medical Assistant on Realistic Clinical Queries Large Language Models Encode Clinical Knowledge

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T14:21:47.334773Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:21:47.334773Z digest=sha256:76ab869a146aa3e2ad1e72ba1b42221b7279a523b5b28a5484b01be56edc07df

Observation 829959f1-622c-4ff3-9bdb-b8d9d3171c2e · outbound

This paper cites Clinical Reasoning of a Generative Artificial Intelligence Model Compared With Physi- cians.

OpenAIs HealthBench in Action: Evaluating an LLM-Based Medical Assistant on Realistic Clinical Queries Clinical Reasoning of a Generative Artificial Intelligence Model Compared With Physi- cians

Reference 8

Resolution
verified exact
doi, observed 2026-08-05T14:21:50.104772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T14:21:47.434845Z digest=sha256:ee2162ab4a4bda8e3832f0a24b68d668faf8602fd5b825516aac9343c0bb28fa

Observation 8b0c8d4e-4550-48d2-a392-92ea17dd261d · outbound

This paper cites What disease does this patient have? A large-scale open domain question answering dataset from medical exams.

OpenAIs HealthBench in Action: Evaluating an LLM-Based Medical Assistant on Realistic Clinical Queries What disease does this patient have? A large-scale open domain question answering dataset from medical exams

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T14:21:47.634240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:21:47.634240Z digest=sha256:7b6d77aa7fe6f07f5bcf50a3abd28471bb6f8c3ec554f8fd12f403d2a7710e89

Observation 9e983042-da58-4197-b2d2-682e58a4a3ca · outbound

This paper cites MedMCQA: A Large-scale Multi-Subject Multi-Choice Dataset for Medical domain Question Answering.

OpenAIs HealthBench in Action: Evaluating an LLM-Based Medical Assistant on Realistic Clinical Queries MedMCQA: A Large-scale Multi-Subject Multi-Choice Dataset for Medical domain Question Answering

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:21:51.804168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T14:21:47.744755Z digest=sha256:d0f0375602ea6eda2bf9c7abf88fd5d780172f0321d0fa8458cd2c91471650e9

Observation f448757c-4fb3-4412-b70c-cc824f6ded0a · outbound

This paper cites PubMedQA: A Dataset for Biomedical Research Question Answering.

OpenAIs HealthBench in Action: Evaluating an LLM-Based Medical Assistant on Realistic Clinical Queries PubMedQA: A Dataset for Biomedical Research Question Answering

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T14:21:47.847549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:21:47.847549Z digest=sha256:907d7c632d1c7d9b3efeb0a8b192ec12a13dad63c3a5fac01bbdc4e6708570c0

Observation e3df4048-aebb-460e-bbe9-103af12f27e8 · outbound

This paper cites Sequential Diagnosis with Language Models.

OpenAIs HealthBench in Action: Evaluating an LLM-Based Medical Assistant on Realistic Clinical Queries Sequential Diagnosis with Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T14:21:47.914758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:21:47.914758Z digest=sha256:5963edae6f682e269bc1cb3783d351df79875395ead23530b674d6290c1aef2e

Observation 90590ceb-b958-4354-9ca7-36c4b8b5aa1e · outbound

This paper cites HealthBench: Evaluating Large Language Models Towards Improved Human Health.

OpenAIs HealthBench in Action: Evaluating an LLM-Based Medical Assistant on Realistic Clinical Queries HealthBench: Evaluating Large Language Models Towards Improved Human Health

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T14:21:48.024778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:21:48.024778Z digest=sha256:2bf9115b826b02bec0bdb06ba492637094cb57a5f4206149a54cb4d595fc7aa1

Observation b5c23685-9364-4969-ba23-1a77bd1666b3 · outbound

This paper cites Dr.INFO: Agentic Clinical Assistant.

OpenAIs HealthBench in Action: Evaluating an LLM-Based Medical Assistant on Realistic Clinical Queries Dr.INFO: Agentic Clinical Assistant

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:21:51.566636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T14:21:48.166153Z digest=sha256:30d4c217107b5f7b7576aa771f0c3d61d4248d88ea7bd1a6d6bd824cc087e71d

Observation 9c2cb5cc-37b6-4b91-9024-227d766dda89 · outbound

This paper cites Leveraging long context in retrieval augmented language models for medical question answering.

OpenAIs HealthBench in Action: Evaluating an LLM-Based Medical Assistant on Realistic Clinical Queries Leveraging long context in retrieval augmented language models for medical question answering

Reference 15

Resolution
verified exact
doi, observed 2026-08-05T14:21:49.691711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T14:21:48.307072Z digest=sha256:4dae750bbf08f5a21d06f7609d8fb75eb31f49ffbce00d23738e0ee19840b5bd

Observation 45e1f2df-3ae4-4b01-8c72-e8685e9bafae · outbound

This paper cites VISTA: A Rubric-based Visual Task Assessment Instruction-Specific Task Assessments.

OpenAIs HealthBench in Action: Evaluating an LLM-Based Medical Assistant on Realistic Clinical Queries VISTA: A Rubric-based Visual Task Assessment Instruction-Specific Task Assessments

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:21:51.285419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T14:21:48.490522Z digest=sha256:e3b90f9be0602facbf76104803171c57afe593d2d4228e937707216eabb2779c

Observation b1c8718a-f500-42f2-80f9-5492d1609b8e · outbound

This paper cites PaperBench: Evaluating AI's Ability to Replicate AI Research.

OpenAIs HealthBench in Action: Evaluating an LLM-Based Medical Assistant on Realistic Clinical Queries PaperBench: Evaluating AI's Ability to Replicate AI Research

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T14:21:48.654747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:21:48.654747Z digest=sha256:a93777f62cb3b3f582701f542d7ddf413e1002e477e09e5997b3131103a69684

Observation cd8a3431-2352-4886-9242-4b74c792ef2d · outbound

This paper cites WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild.

OpenAIs HealthBench in Action: Evaluating an LLM-Based Medical Assistant on Realistic Clinical Queries WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-05T14:21:48.754749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:21:48.754749Z digest=sha256:71a936688587de5d719decdf6dd438656e4cafc7062323585c4bd10d972d83ae

Observation cf7133c9-94f5-43dc-822f-5eb75b4da0f0 · outbound

This paper cites Autonomous medical evaluation for guideline adherence of large language models.

OpenAIs HealthBench in Action: Evaluating an LLM-Based Medical Assistant on Realistic Clinical Queries Autonomous medical evaluation for guideline adherence of large language models

Reference 20

Resolution
malformed identifier
no resolver link, observed 2026-08-05T14:21:48.834137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:21:48.834137Z digest=sha256:2da3e53226cc7b17f6a843ac4f3ef108a58451f56ea5c61c96b937155dba9f79

Observation 57df649d-5060-47bc-9a32-08469a0b455a · outbound

This paper cites MultiChallenge: A Realistic Multi-Turn Conversation Evaluation Benchmark Challenging to Frontier LLMs.

OpenAIs HealthBench in Action: Evaluating an LLM-Based Medical Assistant on Realistic Clinical Queries MultiChallenge: A Realistic Multi-Turn Conversation Evaluation Benchmark Challenging to Frontier LLMs

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T14:21:48.990246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:21:48.990246Z digest=sha256:5db216f929e737e26a19b463a04815a5610d9f3c4648b9dd07607ed2b09ddc57

Observation 1a6b30a2-6cc9-4ca2-ab04-21ac5fc7e70c · outbound

This paper cites Introducing GPT-5.

OpenAIs HealthBench in Action: Evaluating an LLM-Based Medical Assistant on Realistic Clinical Queries Introducing GPT-5

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:21:51.069150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T14:21:49.124883Z digest=sha256:f42ed275f446c878135aac3329bb97c5cd3e0be3a2ccbb854dcf1efd1f6169ce

Observation 1a42e871-2b06-46fe-be0b-3a73808e7376 · outbound

This paper cites an unresolved cited work.

OpenAIs HealthBench in Action: Evaluating an LLM-Based Medical Assistant on Realistic Clinical Queries Unresolved cited work

Reference 295

Resolution
verified exact
arxiv_id_nonexistent, observed 2026-08-05T14:21:50.817815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T14:21:47.509445Z digest=sha256:22ee28e58a5cc928ebd486e05b8c56d8be2fb29b72be4d5562e288e362ab7bf9

Pith citing papers

Observation e86579b6-042c-4532-831e-1997ffce9ee8 · inbound

DR. INFO at the Point of Care: A Prospective Pilot Study of Physician-Perceived Value of an Agentic AI Clinical Assistant cites this paper.

DR. INFO at the Point of Care: A Prospective Pilot Study of Physician-Perceived Value of an Agentic AI Clinical Assistant OpenAIs HealthBench in Action: Evaluating an LLM-Based Medical Assistant on Realistic Clinical Queries

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-28T03:23:19.351167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T10:33:20.149592Z digest=sha256:60327e4bd54a5a12b25e96717cedb2c4a75b8d57ce8a2159c476ca530af31933

Observation 2123923a-01e5-4e2e-b43c-1c4145388c74 · inbound

MDIA: A Multi-Agent Diagnostic Intelligence Pipeline on HealthBench Professional cites this paper.

MDIA: A Multi-Agent Diagnostic Intelligence Pipeline on HealthBench Professional OpenAIs HealthBench in Action: Evaluating an LLM-Based Medical Assistant on Realistic Clinical Queries

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-07-28T03:23:19.351167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T13:13:26.156464Z digest=sha256:6954ca25cfa871c432a1ad85347d049d2300cbd1bc5d7e5f4d928b3587645a19

Observation 89d6a204-5f3c-47cd-bef8-a5549cc475ba · inbound

Generalistic or Specific Embeddings, Which is Better? An Empirical Study on Search for Clinical Coding in Non-English Languages cites this paper.

Generalistic or Specific Embeddings, Which is Better? An Empirical Study on Search for Clinical Coding in Non-English Languages OpenAIs HealthBench in Action: Evaluating an LLM-Based Medical Assistant on Realistic Clinical Queries

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-07-28T03:23:19.351167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-29T07:27:03.060416Z digest=sha256:847358d8a6addfc106252f81db4d03fd93ea596ab7bf6d8d65ed95bc573a0715