Pith. sign in

Paper Citation Record · LEDGER

MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 23 inbound Pith citation observations for arXiv:2501.14654.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.14654 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 23 of 23 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:52:03.354525Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

4
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 0b8e6533-2940-4938-bc2b-30cd95ddd63f · inbound

Large Language Model Agent: A Survey on Methodology, Applications and Challenges cites this paper.

Large Language Model Agent: A Survey on Methodology, Applications and Challenges MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents

Reference 136

Resolution
metadata mismatch
arxiv_id, observed 2026-05-22T21:52:10.740837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-22T21:51:34.309870Z digest=sha256:f54b3706294f1f080e7aa38dbf24f5bd27fd4b095b4d2eb293ddebc257051221

Observation 3659cc0b-64ed-42c2-8922-6e1b5b36e509 · inbound

MedSentry: Understanding and Mitigating Safety Risks in Medical LLM Multi-Agent Systems cites this paper.

MedSentry: Understanding and Mitigating Safety Risks in Medical LLM Multi-Agent Systems MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T13:52:03.354525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:52:03.354525Z digest=sha256:98abcc3cbdf048d8703e21eb4914a564cedc6956567506fda6ff718db3b65dd8

Observation dd81056e-e4f3-4d72-9174-e2f8aadf642b · inbound

BehaviorSFT: Behavioral Token Conditioning for Clinical Agents Across the Proactivity Spectrum cites this paper.

BehaviorSFT: Behavioral Token Conditioning for Clinical Agents Across the Proactivity Spectrum MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-07T13:28:01.830371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:28:01.830371Z digest=sha256:8a62290bd1d08594db99d47fac3b86c6e30cb886e567087abaeaf017d89872e7

Observation f4787371-308f-4d4d-86c7-dc604f5bea2f · inbound

Adaptive-VP: A Framework for LLM-Based Virtual Patients that Adapts to Trainees' Dialogue to Facilitate Nurse Communication Training cites this paper.

Adaptive-VP: A Framework for LLM-Based Virtual Patients that Adapts to Trainees' Dialogue to Facilitate Nurse Communication Training MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-07T12:09:03.679079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:09:03.679079Z digest=sha256:8f1af4ba764d66c3402af2e32a8430633b986ff4da450078c7e76292ec9181da

Observation 766ac164-8ae7-487c-9284-875414760586 · inbound

A Comprehensive Survey of Electronic Health Record Modeling: From Deep Learning Approaches to Large Language Models cites this paper.

A Comprehensive Survey of Electronic Health Record Modeling: From Deep Learning Approaches to Large Language Models MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents

Reference 118

Resolution
unresolved
no resolver link, observed 2026-08-06T16:42:29.834880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:42:29.834880Z digest=sha256:9af33e5f736f12d9c674ae6584f2ebb9374423f19a095aa65fd0ffa99d88de1a

Observation e9af4b11-6ba5-424b-98b2-7eec7f0c8735 · inbound

When Agents Look the Same: Quantifying Distillation-Induced Similarity in Tool-Use Behaviors cites this paper.

When Agents Look the Same: Quantifying Distillation-Induced Similarity in Tool-Use Behaviors MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T14:16:18.019393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-09T22:07:58.614654Z digest=sha256:e73e170dbe899409060a440b9b61e33bc4485d8286c8aa12ecd29c7d540768b4

Observation 5c8bd682-a9b2-4fd6-84c2-078c266d92fd · inbound

BioMedArena: An Open-source Toolkit for Building and Evaluating Biomedical Deep Research Agents cites this paper.

BioMedArena: An Open-source Toolkit for Building and Evaluating Biomedical Deep Research Agents MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T20:06:08.854680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-08T10:22:34.260242Z digest=sha256:5ddaf3225ff1651df0f6f1148bbe79a9f43f828b80f99b20052fdcb496817186

Observation f9ee0f25-79ba-4d32-86d2-bf01797963ec · inbound

BioMedArena: An Open-source Toolkit for Building and Evaluating Biomedical Deep Research Agents cites this paper.

BioMedArena: An Open-source Toolkit for Building and Evaluating Biomedical Deep Research Agents MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T23:35:07.124201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T23:33:52.182432Z digest=sha256:9475a303e2323a23d12f80f502475688f88e6c1fff8ef2fae2c983f0aa837639

Observation a4f7beb2-ec53-4c51-983b-a2703a7ce921 · inbound

ClinQueryAgent: A Conversational Agent for Population Health Management cites this paper.

ClinQueryAgent: A Conversational Agent for Population Health Management MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents

Reference 148

Resolution
verified exact
arxiv_id, observed 2026-05-21T01:33:55.639019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-21T01:31:07.031424Z digest=sha256:e9aaa62d2067ca46d63c7e7e7c19fe5782d29f97e8223bbbb7df115c8ffcf09b

Observation a633ad67-e9b3-4c9b-b9af-8addec95764e · inbound

Design and Report Benchmarks for Knowledge Work cites this paper.

Design and Report Benchmarks for Knowledge Work MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents

Reference 127

Resolution
verified exact
arxiv_id, observed 2026-05-25T04:40:23.554227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-25T04:39:14.319133Z digest=sha256:7c5ff02af7855ed7ce1bc7baf781d3ae672e7de7cb93d884aac7e78598bd4c81

Observation 5bb0b44e-83f6-4333-80b3-5cdc13f30f2c · inbound

AutoMedBench: Towards Medical AutoResearch with Agentic AI Models cites this paper.

AutoMedBench: Towards Medical AutoResearch with Agentic AI Models MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-07-01T23:06:21.084048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T14:38:40.017263Z digest=sha256:817216d90ded3f44a4c29bc31a0ca40675d5685da685b35b8585f9bda2770169

Observation 49c534cf-00ca-4e6f-a63f-6bdc5a13088f · inbound

UniReason-Med: A Shared Grounded Reasoning Interface for 2D-to-3D Transfer in Medical VQA cites this paper.

UniReason-Med: A Shared Grounded Reasoning Interface for 2D-to-3D Transfer in Medical VQA MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents

Reference 157

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T09:47:59.364770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-27T10:21:12.782864Z digest=sha256:0e1dc846e64ed09d879348790be0f6464e14871720835293e268cfcf3e3ff3c0

Observation 72c61c02-8158-4888-868b-fdcef8f96452 · inbound

Agentic Environment Engineering for Large Language Models: A Survey of Environment Modeling, Synthesis, Evaluation, and Application cites this paper.

Agentic Environment Engineering for Large Language Models: A Survey of Environment Modeling, Synthesis, Evaluation, and Application MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents

Reference 47

Resolution
metadata mismatch
arxiv_id, observed 2026-06-27T09:50:48.306082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T09:46:30.702256Z digest=sha256:c7b7ee37d352ae3be7126bf4eb500be128b8f769ababc6a7520ef31feebbd4f3

Observation 1c4768a3-0fc8-408c-9191-6c9e5ae83fe9 · inbound

Measuring Epistemic Resilience of LLMs Under Misleading Medical Context cites this paper.

Measuring Epistemic Resilience of LLMs Under Misleading Medical Context MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-06-27T09:40:47.682854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T09:30:47.726480Z digest=sha256:46391266f9252b59c698511807cbdd2b5bd5da0aabeb7f43fa9c069a8c0fccc8

Observation 5ba495e1-35b3-4e47-95ff-6e589c2b6bb4 · inbound

Benchmarking AI Agents for Addressing Scientific Challenges Across Scales cites this paper.

Benchmarking AI Agents for Addressing Scientific Challenges Across Scales MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents

Reference 145

Resolution
verified exact
arxiv_id, observed 2026-07-03T11:28:04.327347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T09:34:09.347912Z digest=sha256:8ea1f04240d8ebf5c5fbaf4bb898efc3cdab6b0b27d32608fbedcff2e6633433

Observation 0a341d0f-793c-4991-b85d-cf1c49841d6e · inbound

AgentBeats: Agentifying Agent Assessment for Openness, Standardization, and Reproducibility cites this paper.

AgentBeats: Agentifying Agent Assessment for Openness, Standardization, and Reproducibility MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-07-03T15:08:33.136412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T06:41:41.799596Z digest=sha256:44a503515bb9024b4bab0d75872b99914c8dab739e782b9a2831ff9dc14afb2a

Observation 79b19494-493c-41c3-87a6-b2bf24459395 · inbound

AgentFairBench: Do LLM Agents Discriminate When They Act? cites this paper.

AgentFairBench: Do LLM Agents Discriminate When They Act? MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-07-03T17:38:44.494148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T03:53:38.554457Z digest=sha256:ef96d1568d1fb6ab9ddd68843a28087495c2e95e9e6ac4c86fee8e356e85923f

Observation a206d07b-58fc-4504-8d1b-d6873ae52798 · inbound

How Do Tool-Augmented LLM Agents Perform on Real-World Energy Analytics Tasks? cites this paper.

How Do Tool-Augmented LLM Agents Perform on Real-World Energy Analytics Tasks? MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-07-04T15:29:56.673808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T01:34:10.103638Z digest=sha256:63fc0d75428cb543528d9f1893ee5b576c33c9a876aef356a02768ec02503eff

Observation 5035f6d7-57dc-4869-9ca7-57a148805abe · inbound

MedEvoEval: Evaluating Continual Evolution of Doctor Agents through Simulated Clinical Episodes cites this paper.

MedEvoEval: Evaluating Continual Evolution of Doctor Agents through Simulated Clinical Episodes MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-06-30T09:34:34.396230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T09:33:50.358159Z digest=sha256:24d2f1188362fe44039c65f802bc7104e9a504f72c7e07d0b36389dee6e8465d

Observation 6040f6c6-0641-41fd-a0a6-4cd8d4a1bd1e · inbound

An Empirical Evaluation of Prompt Injection Vulnerabilities in Large Language Models Across Multilingual and Obfuscated Attack Scenarios cites this paper.

An Empirical Evaluation of Prompt Injection Vulnerabilities in Large Language Models Across Multilingual and Obfuscated Attack Scenarios MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents

Reference 28

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T06:54:20.381699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T06:51:43.219551Z digest=sha256:24a1503d7883297b134386c7883c3c79d987237b32de0b127c8b79d375992bd6

Observation b3f8ce0b-193e-4b27-88d2-3383809cc627 · inbound

Cura 1T: Specialized Model for Agentic Healthcare cites this paper.

Cura 1T: Specialized Model for Agentic Healthcare MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-02T02:16:39.860668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:16:39.860668Z digest=sha256:35371ccfec1e7da4af85ba2a733c472dce756dc396514d3849974b98ff6e7177

Observation 74d61818-c0ee-4ad4-94cd-98cf5826d025 · inbound

MedDDC-Eval: Diagnosis-Decoupled Evaluation of Multi-Turn Medical Consultation Agents cites this paper.

MedDDC-Eval: Diagnosis-Decoupled Evaluation of Multi-Turn Medical Consultation Agents MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T13:50:41.542608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T13:50:41.542608Z digest=sha256:3b8ad6bac317fe566dccdd7b81696211ab6943240ff6490a65c9a85b5e71cc74

Observation d33b646e-a2ec-41a9-8144-9e8dea90379c · inbound

GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks cites this paper.

GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-05T13:09:08.491553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:09:08.491553Z digest=sha256:05ab8d737f3a99651d52fb5ce06bd86ec36886569b641093aed4d165b45e353d