Pith. sign in

Paper Citation Record · LEDGER

MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 23 inbound Pith citation observations for arXiv:2501.14654.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.14654 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 23 of 23 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:52:03.354525Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

4
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 0b8e6533-2940-4938-bc2b-30cd95ddd63f · inbound

Large Language Model Agent: A Survey on Methodology, Applications and Challenges cites this paper.

Large Language Model Agent: A Survey on Methodology, Applications and Challenges MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents

Reference 136

Resolution
metadata mismatch
arxiv_id, observed 2026-05-22T21:52:10.740837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-22T21:51:34.309870Z digest=sha256:faa36a700908ce5d5c295c939f19e5ea3d08e033422972dc931de2e250825124

Observation 3659cc0b-64ed-42c2-8922-6e1b5b36e509 · inbound

MedSentry: Understanding and Mitigating Safety Risks in Medical LLM Multi-Agent Systems cites this paper.

MedSentry: Understanding and Mitigating Safety Risks in Medical LLM Multi-Agent Systems MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T13:52:03.354525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:52:03.354525Z digest=sha256:bd766f433b54ccbabb27f842d0faa0a2dd965b6c9632e52e1e5958d8ecaa3fe1

Observation dd81056e-e4f3-4d72-9174-e2f8aadf642b · inbound

BehaviorSFT: Behavioral Token Conditioning for Clinical Agents Across the Proactivity Spectrum cites this paper.

BehaviorSFT: Behavioral Token Conditioning for Clinical Agents Across the Proactivity Spectrum MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-07T13:28:01.830371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:28:01.830371Z digest=sha256:cbfce13b9f7fb9c187ef2a73d257bae8b49d0e8acd67d5d0acc532a6050a42a6

Observation f4787371-308f-4d4d-86c7-dc604f5bea2f · inbound

Adaptive-VP: A Framework for LLM-Based Virtual Patients that Adapts to Trainees' Dialogue to Facilitate Nurse Communication Training cites this paper.

Adaptive-VP: A Framework for LLM-Based Virtual Patients that Adapts to Trainees' Dialogue to Facilitate Nurse Communication Training MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-07T12:09:03.679079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:09:03.679079Z digest=sha256:8f1af4ba764d66c3402af2e32a8430633b986ff4da450078c7e76292ec9181da

Observation 766ac164-8ae7-487c-9284-875414760586 · inbound

A Comprehensive Survey of Electronic Health Record Modeling: From Deep Learning Approaches to Large Language Models cites this paper.

A Comprehensive Survey of Electronic Health Record Modeling: From Deep Learning Approaches to Large Language Models MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents

Reference 118

Resolution
unresolved
no resolver link, observed 2026-08-06T16:42:29.834880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:42:29.834880Z digest=sha256:9af33e5f736f12d9c674ae6584f2ebb9374423f19a095aa65fd0ffa99d88de1a

Observation e9af4b11-6ba5-424b-98b2-7eec7f0c8735 · inbound

When Agents Look the Same: Quantifying Distillation-Induced Similarity in Tool-Use Behaviors cites this paper.

When Agents Look the Same: Quantifying Distillation-Induced Similarity in Tool-Use Behaviors MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T14:16:18.019393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-09T22:07:58.614654Z digest=sha256:47e9c84926b614be4226cfb857cc5ac4160088c7166f38d3b7c1bb423374c0e9

Observation 5c8bd682-a9b2-4fd6-84c2-078c266d92fd · inbound

BioMedArena: An Open-source Toolkit for Building and Evaluating Biomedical Deep Research Agents cites this paper.

BioMedArena: An Open-source Toolkit for Building and Evaluating Biomedical Deep Research Agents MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T20:06:08.854680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T10:22:34.260242Z digest=sha256:27d1f220f8e4c798026c7dae30070e7a23138b035fa267883ea0d528038a5dc0

Observation f9ee0f25-79ba-4d32-86d2-bf01797963ec · inbound

BioMedArena: An Open-source Toolkit for Building and Evaluating Biomedical Deep Research Agents cites this paper.

BioMedArena: An Open-source Toolkit for Building and Evaluating Biomedical Deep Research Agents MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T23:35:07.124201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T23:33:52.182432Z digest=sha256:459af0301cd77684f46aa44e0dc1b325d8170f743baa852df3adcfaed48c7634

Observation a4f7beb2-ec53-4c51-983b-a2703a7ce921 · inbound

ClinQueryAgent: A Conversational Agent for Population Health Management cites this paper.

ClinQueryAgent: A Conversational Agent for Population Health Management MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents

Reference 148

Resolution
verified exact
arxiv_id, observed 2026-05-21T01:33:55.639019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-21T01:31:07.031424Z digest=sha256:b0f42f7e58775bad17ec336a2cf7c3f8790be75bbbd6011ad8e5aba242ad1d46

Observation a633ad67-e9b3-4c9b-b9af-8addec95764e · inbound

Design and Report Benchmarks for Knowledge Work cites this paper.

Design and Report Benchmarks for Knowledge Work MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents

Reference 127

Resolution
verified exact
arxiv_id, observed 2026-05-25T04:40:23.554227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-25T04:39:14.319133Z digest=sha256:0ea1a7b5a9991ef5e4114861cb3a4b3893063dc6d7a8fe86a5d74a65ecdffd17

Observation 5bb0b44e-83f6-4333-80b3-5cdc13f30f2c · inbound

AutoMedBench: Towards Medical AutoResearch with Agentic AI Models cites this paper.

AutoMedBench: Towards Medical AutoResearch with Agentic AI Models MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-07-01T23:06:21.084048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T14:38:40.017263Z digest=sha256:96a668aab9beed740ce51d9807e8ffe408b0470c0f1629a57deb7e6ce6011010

Observation 49c534cf-00ca-4e6f-a63f-6bdc5a13088f · inbound

UniReason-Med: A Shared Grounded Reasoning Interface for 2D-to-3D Transfer in Medical VQA cites this paper.

UniReason-Med: A Shared Grounded Reasoning Interface for 2D-to-3D Transfer in Medical VQA MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents

Reference 157

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T09:47:59.364770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-27T10:21:12.782864Z digest=sha256:cf671fc6426887dec0ec365f765fd3ba25c85ac394287b54af2b8acfca3d1aac

Observation 72c61c02-8158-4888-868b-fdcef8f96452 · inbound

Agentic Environment Engineering for Large Language Models: A Survey of Environment Modeling, Synthesis, Evaluation, and Application cites this paper.

Agentic Environment Engineering for Large Language Models: A Survey of Environment Modeling, Synthesis, Evaluation, and Application MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents

Reference 47

Resolution
metadata mismatch
arxiv_id, observed 2026-06-27T09:50:48.306082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T09:46:30.702256Z digest=sha256:b561cb3582eb475cc0d078a8a8df72e01f07366b1b4690c2ecca839740aa056c

Observation 1c4768a3-0fc8-408c-9191-6c9e5ae83fe9 · inbound

Measuring Epistemic Resilience of LLMs Under Misleading Medical Context cites this paper.

Measuring Epistemic Resilience of LLMs Under Misleading Medical Context MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-06-27T09:40:47.682854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T09:30:47.726480Z digest=sha256:4ea61ea9e296eb7331f377939a9c2098e14b3b1b65c9f6cc466a9b5be62cc322

Observation 5ba495e1-35b3-4e47-95ff-6e589c2b6bb4 · inbound

Benchmarking AI Agents for Addressing Scientific Challenges Across Scales cites this paper.

Benchmarking AI Agents for Addressing Scientific Challenges Across Scales MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents

Reference 145

Resolution
verified exact
arxiv_id, observed 2026-07-03T11:28:04.327347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T09:34:09.347912Z digest=sha256:1df2ed90182c915c6195344150bfa33d11acdeed27a52bdd641038657ad6ea1d

Observation 0a341d0f-793c-4991-b85d-cf1c49841d6e · inbound

AgentBeats: Agentifying Agent Assessment for Openness, Standardization, and Reproducibility cites this paper.

AgentBeats: Agentifying Agent Assessment for Openness, Standardization, and Reproducibility MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-07-03T15:08:33.136412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T06:41:41.799596Z digest=sha256:37a2231aac13b52739164a1aa3d43b39f32fbf6439dc04a1f3acb91775814d3f

Observation 79b19494-493c-41c3-87a6-b2bf24459395 · inbound

AgentFairBench: Do LLM Agents Discriminate When They Act? cites this paper.

AgentFairBench: Do LLM Agents Discriminate When They Act? MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-07-03T17:38:44.494148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T03:53:38.554457Z digest=sha256:36a5b2f956137e87a7068630c3c2168db43153729e70aff5cbb14785c0a249ea

Observation a206d07b-58fc-4504-8d1b-d6873ae52798 · inbound

How Do Tool-Augmented LLM Agents Perform on Real-World Energy Analytics Tasks? cites this paper.

How Do Tool-Augmented LLM Agents Perform on Real-World Energy Analytics Tasks? MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-07-04T15:29:56.673808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T01:34:10.103638Z digest=sha256:e4febf168117c7cb4d3f055fab2fa269d474303b179ace083306f880587f7ea2

Observation 5035f6d7-57dc-4869-9ca7-57a148805abe · inbound

MedEvoEval: Evaluating Continual Evolution of Doctor Agents through Simulated Clinical Episodes cites this paper.

MedEvoEval: Evaluating Continual Evolution of Doctor Agents through Simulated Clinical Episodes MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-06-30T09:34:34.396230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T09:33:50.358159Z digest=sha256:ae2e791584252e964497dad6821d4e902e56c45d958a1bd542f7965b492622db

Observation 6040f6c6-0641-41fd-a0a6-4cd8d4a1bd1e · inbound

An Empirical Evaluation of Prompt Injection Vulnerabilities in Large Language Models Across Multilingual and Obfuscated Attack Scenarios cites this paper.

An Empirical Evaluation of Prompt Injection Vulnerabilities in Large Language Models Across Multilingual and Obfuscated Attack Scenarios MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents

Reference 28

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T06:54:20.381699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T06:51:43.219551Z digest=sha256:fa638708a0ab09750bc8c2394a16e161aff080f4f28b7a510e11c22bdb3a45ad

Observation b3f8ce0b-193e-4b27-88d2-3383809cc627 · inbound

Cura 1T: Specialized Model for Agentic Healthcare cites this paper.

Cura 1T: Specialized Model for Agentic Healthcare MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-02T02:16:39.860668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:16:39.860668Z digest=sha256:d6d4f6a0ff38475cca949f7db90f41799a2361dfc8c76d4437105074702f024d

Observation 74d61818-c0ee-4ad4-94cd-98cf5826d025 · inbound

MedDDC-Eval: Diagnosis-Decoupled Evaluation of Multi-Turn Medical Consultation Agents cites this paper.

MedDDC-Eval: Diagnosis-Decoupled Evaluation of Multi-Turn Medical Consultation Agents MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T13:50:41.542608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T13:50:41.542608Z digest=sha256:3b8ad6bac317fe566dccdd7b81696211ab6943240ff6490a65c9a85b5e71cc74

Observation d33b646e-a2ec-41a9-8144-9e8dea90379c · inbound

GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks cites this paper.

GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-05T13:09:08.491553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:09:08.491553Z digest=sha256:05ab8d737f3a99651d52fb5ce06bd86ec36886569b641093aed4d165b45e353d