Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T10:46:10.522415Z
Paper Citation Record · LEDGER
As of 10 August 2026, this Paper Citation Record lists 32 of 32 outbound references and 5 inbound Pith citation observations for arXiv:2507.23486.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T10:46:10.522415Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-07-13T00:19:33.861692Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-03T11:08:02.948778Z
32 of 32 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation fc25d4d6-c059-44ad-8064-17e3be109249 · outbound
A Novel Evaluation Benchmark for Medical LLMs: Illuminating Safety and Effectiveness in Clinical Domains A., Gui, H., Rezaei, S
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0f85658d-f08c-4a65-8f52-e38b77919984 · outbound
A Novel Evaluation Benchmark for Medical LLMs: Illuminating Safety and Effectiveness in Clinical Domains et al.Towards accurate differential diagnosis with large language models.Nature 642, 451-457, doi:10.1038/s41586-025-08869-4 (2025)
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f97b2fd1-0c6d-4f6f-b5c6-1a79d9b7e55e · outbound
A Novel Evaluation Benchmark for Medical LLMs: Illuminating Safety and Effectiveness in Clinical Domains Unresolved cited work
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e050803-03fe-43b4-989b-1259e5d3d3b4 · outbound
A Novel Evaluation Benchmark for Medical LLMs: Illuminating Safety and Effectiveness in Clinical Domains et al.Foundation models for generalist medical artificial intelligence
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2fd6b8f4-60ef-4d98-8a7d-fa27cd9ffd22 · outbound
A Novel Evaluation Benchmark for Medical LLMs: Illuminating Safety and Effectiveness in Clinical Domains Unresolved cited work
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eb0e64ca-6258-4a4f-83af-9aa82c7683c7 · outbound
A Novel Evaluation Benchmark for Medical LLMs: Illuminating Safety and Effectiveness in Clinical Domains MeDiSumQA: Patient-Oriented Question-Answer Generation from Discharge Letters
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 9d895fdc-55c7-4b1a-b9f9-c3a18ff14d97 · outbound
A Novel Evaluation Benchmark for Medical LLMs: Illuminating Safety and Effectiveness in Clinical Domains et al.Adapted large language models can outperform medical experts in clin- ical text summarization
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8e0589a2-b648-4e5a-a4d7-97debbe431f6 · outbound
A Novel Evaluation Benchmark for Medical LLMs: Illuminating Safety and Effectiveness in Clinical Domains Clean & Clear: Feasibility of Safe LLM Clinical Guidance
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 76ec1c29-fd02-4c19-81e1-04043c52dca7 · outbound
A Novel Evaluation Benchmark for Medical LLMs: Illuminating Safety and Effectiveness in Clinical Domains Unresolved cited work
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f76e3116-4053-46a8-b86d-30f1c82a1d3d · outbound
A Novel Evaluation Benchmark for Medical LLMs: Illuminating Safety and Effectiveness in Clinical Domains Unresolved cited work
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6de4fe4d-4f49-49b0-b41d-1f88f97020ae · outbound
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 1c088b28-fcb0-4393-bba6-5d3efbb4fcf1 · outbound
A Novel Evaluation Benchmark for Medical LLMs: Illuminating Safety and Effectiveness in Clinical Domains MedBench: A Comprehensive, Standardized, and Reliable Benchmarking System for Evaluating Chinese Medical Large Language Models
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f9460823-9764-4395-b2d8-e576d10e5b52 · outbound
A Novel Evaluation Benchmark for Medical LLMs: Illuminating Safety and Effectiveness in Clinical Domains SafeBench: A Safety Evaluation Framework for Multimodal Large Language Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 793b6677-67f9-4e56-9a81-72f5750912d9 · outbound
A Novel Evaluation Benchmark for Medical LLMs: Illuminating Safety and Effectiveness in Clinical Domains Agent-SafetyBench: Evaluating the Safety of LLM Agents
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 36310e5b-7a2c-431f-a77c-6aa87474a2d7 · outbound
A Novel Evaluation Benchmark for Medical LLMs: Illuminating Safety and Effectiveness in Clinical Domains aiXamine: Simplified LLM Safety and Security
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 18da0b2d-f55f-44f3-ae5d-7ce0dcc78126 · outbound
A Novel Evaluation Benchmark for Medical LLMs: Illuminating Safety and Effectiveness in Clinical Domains Unresolved cited work
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cca9673b-cc10-4286-8fbb-641821ba5438 · outbound
A Novel Evaluation Benchmark for Medical LLMs: Illuminating Safety and Effectiveness in Clinical Domains HealthBench: Evaluating Large Language Models Towards Improved Human Health
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c9cb5b69-05e3-4f79-9823-8e1d21eff0d0 · outbound
A Novel Evaluation Benchmark for Medical LLMs: Illuminating Safety and Effectiveness in Clinical Domains Towards Automatic Evaluation for LLMs' Clinical Capabilities: Metric, Data, and Algorithm
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation c1e5acbb-284c-489b-bd4e-424c254d10a2 · outbound
A Novel Evaluation Benchmark for Medical LLMs: Illuminating Safety and Effectiveness in Clinical Domains Towards Expert-Level Medical Question Answering with Large Language Models
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation caf14afb-0b5e-4554-8baa-f677b1b8b321 · outbound
A Novel Evaluation Benchmark for Medical LLMs: Illuminating Safety and Effectiveness in Clinical Domains et al.An evaluation framework for clinical use of large language models in patient interaction tasks
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 8465f969-37a8-4a75-8d95-49740f439ab8 · outbound
A Novel Evaluation Benchmark for Medical LLMs: Illuminating Safety and Effectiveness in Clinical Domains Towards Conversational Diagnostic AI
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d7748517-aa6a-46e7-9df6-e100b55de009 · outbound
A Novel Evaluation Benchmark for Medical LLMs: Illuminating Safety and Effectiveness in Clinical Domains AgentClinic: a multimodal agent benchmark to evaluate AI in simulated clinical environments
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 64a25829-9479-42dd-bc6f-17829ba377d9 · outbound
A Novel Evaluation Benchmark for Medical LLMs: Illuminating Safety and Effectiveness in Clinical Domains An Automatic Evaluation Framework for Multi-turn Medical Consultations Capabilities of Large Language Models
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2edab784-dfb4-414f-9e5b-c691a60f961d · outbound
A Novel Evaluation Benchmark for Medical LLMs: Illuminating Safety and Effectiveness in Clinical Domains LLM-Mini-CEX: Automatic Evaluation of Large Language Model for Diagnostic Conversation
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 99cb9ef8-8cff-4929-a99b-cee4695c1741 · outbound
A Novel Evaluation Benchmark for Medical LLMs: Illuminating Safety and Effectiveness in Clinical Domains Unresolved cited work
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 844963ff-1f99-4258-9bc5-053685c07377 · outbound
A Novel Evaluation Benchmark for Medical LLMs: Illuminating Safety and Effectiveness in Clinical Domains Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 770d0a68-fd0e-4ec0-b153-d764739537eb · outbound
A Novel Evaluation Benchmark for Medical LLMs: Illuminating Safety and Effectiveness in Clinical Domains et al.Automating Evaluation of AI Text Generation in Healthcare with a Large Language Model (LLM)-as-a-Judge
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6a955e62-4ee5-45fb-9a2a-0f821700a341 · outbound
A Novel Evaluation Benchmark for Medical LLMs: Illuminating Safety and Effectiveness in Clinical Domains Unresolved cited work
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ce190359-ddaf-443c-aab8-c7334c01cedb · outbound
A Novel Evaluation Benchmark for Medical LLMs: Illuminating Safety and Effectiveness in Clinical Domains Unresolved cited work
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 4cd591a4-3e82-4aee-904a-b4e9ae152850 · outbound
A Novel Evaluation Benchmark for Medical LLMs: Illuminating Safety and Effectiveness in Clinical Domains MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 314fc882-ee5b-4b39-a5d5-3b617bf6c05a · outbound
A Novel Evaluation Benchmark for Medical LLMs: Illuminating Safety and Effectiveness in Clinical Domains et al.Aligning Large Language Models with Humans: A Comprehensive Survey of ChatGPT's Aptitude in Pharmacology
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 8bef7af2-1cd9-43f7-b755-df775440a5c4 · outbound
A Novel Evaluation Benchmark for Medical LLMs: Illuminating Safety and Effectiveness in Clinical Domains T., Bardak, A
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 65b203d6-9762-4c96-a5ff-584bbaaaac63 · inbound
IatroBench: Pre-Registered Evidence of Iatrogenic Harm from AI Safety Measures A Novel Evaluation Benchmark for Medical LLMs: Illuminating Safety and Effectiveness in Clinical Domains
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 0444d725-a30e-4366-a6b4-685e8958ff48 · inbound
IatroBench: Pre-Registered Evidence of Iatrogenic Harm from AI Safety Measures A Novel Evaluation Benchmark for Medical LLMs: Illuminating Safety and Effectiveness in Clinical Domains
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ad62a974-59d9-428c-b042-6fd8c90d11cd · inbound
MedVIGIL: Evaluating Trustworthy Medical VLMs Under Broken Visual Evidence A Novel Evaluation Benchmark for Medical LLMs: Illuminating Safety and Effectiveness in Clinical Domains
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 88a75dbd-b112-4ed7-ab12-6276381603a5 · inbound
MedVIGIL: Evaluating Trustworthy Medical VLMs Under Broken Visual Evidence A Novel Evaluation Benchmark for Medical LLMs: Illuminating Safety and Effectiveness in Clinical Domains
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 8018e963-9bd1-4870-a88d-8634dea74c4c · inbound
Deployment-Centered Evaluation: Predicting Query-Level Rejection Risk in a Clinical LLM System A Novel Evaluation Benchmark for Medical LLMs: Illuminating Safety and Effectiveness in Clinical Domains
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.