Pith. sign in

Paper Citation Record · LEDGER

A Novel Evaluation Benchmark for Medical LLMs: Illuminating Safety and Effectiveness in Clinical Domains

As of 10 August 2026, this Paper Citation Record lists 32 of 32 outbound references and 5 inbound Pith citation observations for arXiv:2507.23486.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.23486 v3

Coverage vector

measured 32 of 32 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T10:46:10.522415Z

measured 37 of 37 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-13T00:19:33.861692Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T11:08:02.948778Z

Reference resolution

32 of 32 outbound references displayed

  • verified exact7
  • verified fuzzy0
  • unresolved23
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation fc25d4d6-c059-44ad-8064-17e3be109249 · outbound

This paper cites A., Gui, H., Rezaei, S.

A Novel Evaluation Benchmark for Medical LLMs: Illuminating Safety and Effectiveness in Clinical Domains A., Gui, H., Rezaei, S

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T10:46:10.406277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:46:10.406277Z digest=sha256:1ebaae3dd781037a025bc9c7ab022de696fd4a86855d526a3007e5704f5a76fd

Observation 0f85658d-f08c-4a65-8f52-e38b77919984 · outbound

This paper cites et al.Towards accurate differential diagnosis with large language models.Nature 642, 451-457, doi:10.1038/s41586-025-08869-4 (2025).

A Novel Evaluation Benchmark for Medical LLMs: Illuminating Safety and Effectiveness in Clinical Domains et al.Towards accurate differential diagnosis with large language models.Nature 642, 451-457, doi:10.1038/s41586-025-08869-4 (2025)

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T10:46:10.411592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:46:10.411592Z digest=sha256:9144e493bc554b8c121894655a0e9bacd61f60ec82e3744f29d67fc929970baf

Observation f97b2fd1-0c6d-4f6f-b5c6-1a79d9b7e55e · outbound

This paper cites an unresolved cited work.

A Novel Evaluation Benchmark for Medical LLMs: Illuminating Safety and Effectiveness in Clinical Domains Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T10:46:10.416849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:46:10.416849Z digest=sha256:7c19be50b2d88db54c3218ebb83d325fe9f7b410a25e74beee84820b6a951166

Observation 9e050803-03fe-43b4-989b-1259e5d3d3b4 · outbound

This paper cites et al.Foundation models for generalist medical artificial intelligence.

A Novel Evaluation Benchmark for Medical LLMs: Illuminating Safety and Effectiveness in Clinical Domains et al.Foundation models for generalist medical artificial intelligence

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T10:46:10.421342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:46:10.421342Z digest=sha256:8d3edb86f28692cee4e3e3f02578903f6fbfbb0211cb8573c599c90495959c98

Observation 2fd6b8f4-60ef-4d98-8a7d-fa27cd9ffd22 · outbound

This paper cites an unresolved cited work.

A Novel Evaluation Benchmark for Medical LLMs: Illuminating Safety and Effectiveness in Clinical Domains Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T10:46:10.425603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:46:10.425603Z digest=sha256:6505abdaae51e3ee98f06f94c700ad9dee7c6a42112901713b91dfa6c40399f4

Observation eb0e64ca-6258-4a4f-83af-9aa82c7683c7 · outbound

This paper cites MeDiSumQA: Patient-Oriented Question-Answer Generation from Discharge Letters.

A Novel Evaluation Benchmark for Medical LLMs: Illuminating Safety and Effectiveness in Clinical Domains MeDiSumQA: Patient-Oriented Question-Answer Generation from Discharge Letters

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-08-06T10:46:10.663501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T10:46:10.429615Z digest=sha256:9fb25054f3a8bd9d900ab7426d006d1631e7772f78a8cda2ee8c01875fbceb23

Observation 9d895fdc-55c7-4b1a-b9f9-c3a18ff14d97 · outbound

This paper cites et al.Adapted large language models can outperform medical experts in clin- ical text summarization.

A Novel Evaluation Benchmark for Medical LLMs: Illuminating Safety and Effectiveness in Clinical Domains et al.Adapted large language models can outperform medical experts in clin- ical text summarization

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T10:46:10.434980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:46:10.434980Z digest=sha256:f0e20e4653bb0b6c2933d7636ea70b983179e0adb898c17c43aba298a4e18526

Observation 8e0589a2-b648-4e5a-a4d7-97debbe431f6 · outbound

This paper cites Clean & Clear: Feasibility of Safe LLM Clinical Guidance.

A Novel Evaluation Benchmark for Medical LLMs: Illuminating Safety and Effectiveness in Clinical Domains Clean & Clear: Feasibility of Safe LLM Clinical Guidance

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-08-06T10:46:11.003075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T10:46:10.438729Z digest=sha256:26d698c59f2f4e33f95ce6afc89cafbff738507795b710f064f2b3eeb00201d6

Observation 76ec1c29-fd02-4c19-81e1-04043c52dca7 · outbound

This paper cites an unresolved cited work.

A Novel Evaluation Benchmark for Medical LLMs: Illuminating Safety and Effectiveness in Clinical Domains Unresolved cited work

Reference 9

Resolution
malformed identifier
no resolver link, observed 2026-08-06T10:46:10.442312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:46:10.442312Z digest=sha256:45bfc60124a0f12e0e932c52e1962d9340d25fec290351f89d452362e12e2c99

Observation f76e3116-4053-46a8-b86d-30f1c82a1d3d · outbound

This paper cites an unresolved cited work.

A Novel Evaluation Benchmark for Medical LLMs: Illuminating Safety and Effectiveness in Clinical Domains Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T10:46:10.445844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:46:10.445844Z digest=sha256:9d7205532377a8339e36ef77d096f932a21405728dc9be5b23a20fb9dcb91057

Observation 6de4fe4d-4f49-49b0-b41d-1f88f97020ae · outbound

This paper cites & Cho, B.

A Novel Evaluation Benchmark for Medical LLMs: Illuminating Safety and Effectiveness in Clinical Domains & Cho, B

Reference 11

Resolution
verified exact
doi, observed 2026-08-06T10:46:10.620912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T10:46:10.448927Z digest=sha256:d50060a2cb4066fb9e811880d883665dbdb76784644b4fa6f89d864bb28e32e4

Observation 1c088b28-fcb0-4393-bba6-5d3efbb4fcf1 · outbound

This paper cites MedBench: A Comprehensive, Standardized, and Reliable Benchmarking System for Evaluating Chinese Medical Large Language Models.

A Novel Evaluation Benchmark for Medical LLMs: Illuminating Safety and Effectiveness in Clinical Domains MedBench: A Comprehensive, Standardized, and Reliable Benchmarking System for Evaluating Chinese Medical Large Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T10:46:10.452156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:46:10.452156Z digest=sha256:afa26932ac9e27d6f9e9bbf7f5d2d3321a21988ad57beeeb229982aefe93bbcb

Observation f9460823-9764-4395-b2d8-e576d10e5b52 · outbound

This paper cites SafeBench: A Safety Evaluation Framework for Multimodal Large Language Models.

A Novel Evaluation Benchmark for Medical LLMs: Illuminating Safety and Effectiveness in Clinical Domains SafeBench: A Safety Evaluation Framework for Multimodal Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T10:46:10.455850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:46:10.455850Z digest=sha256:18cad15e2554a9f8dc37d97ebdc27fd5391992af8f61704fe7b6bdbfab458b99

Observation 793b6677-67f9-4e56-9a81-72f5750912d9 · outbound

This paper cites Agent-SafetyBench: Evaluating the Safety of LLM Agents.

A Novel Evaluation Benchmark for Medical LLMs: Illuminating Safety and Effectiveness in Clinical Domains Agent-SafetyBench: Evaluating the Safety of LLM Agents

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T10:46:10.459472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:46:10.459472Z digest=sha256:5a91cf020b65db36ea5d557d735cd9b30c430b567907404bb598f695f51b050b

Observation 36310e5b-7a2c-431f-a77c-6aa87474a2d7 · outbound

This paper cites aiXamine: Simplified LLM Safety and Security.

A Novel Evaluation Benchmark for Medical LLMs: Illuminating Safety and Effectiveness in Clinical Domains aiXamine: Simplified LLM Safety and Security

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T10:46:10.462713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:46:10.462713Z digest=sha256:538aed61e26a93ffa823f5a1abf19428c52b3c282cc49c66a2d8662de8beaa8e

Observation 18da0b2d-f55f-44f3-ae5d-7ce0dcc78126 · outbound

This paper cites an unresolved cited work.

A Novel Evaluation Benchmark for Medical LLMs: Illuminating Safety and Effectiveness in Clinical Domains Unresolved cited work

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T10:46:10.466118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:46:10.466118Z digest=sha256:a85797273c7a307a47d5b101108f86ee7358328ee963b8fd45df65e66bc41977

Observation cca9673b-cc10-4286-8fbb-641821ba5438 · outbound

This paper cites HealthBench: Evaluating Large Language Models Towards Improved Human Health.

A Novel Evaluation Benchmark for Medical LLMs: Illuminating Safety and Effectiveness in Clinical Domains HealthBench: Evaluating Large Language Models Towards Improved Human Health

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T10:46:10.469198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:46:10.469198Z digest=sha256:3dfcd4e8bd09cde8fdadddef687275d47743e7f2c049adc3fe30892aef87f415

Observation c9cb5b69-05e3-4f79-9823-8e1d21eff0d0 · outbound

This paper cites Towards Automatic Evaluation for LLMs' Clinical Capabilities: Metric, Data, and Algorithm.

A Novel Evaluation Benchmark for Medical LLMs: Illuminating Safety and Effectiveness in Clinical Domains Towards Automatic Evaluation for LLMs' Clinical Capabilities: Metric, Data, and Algorithm

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-08-06T10:46:10.904582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T10:46:10.472579Z digest=sha256:5f8b99c4ab3485391bb5e56fd92757a9cebb5f95b048a40bfac5fae00ac878b8

Observation c1e5acbb-284c-489b-bd4e-424c254d10a2 · outbound

This paper cites Towards Expert-Level Medical Question Answering with Large Language Models.

A Novel Evaluation Benchmark for Medical LLMs: Illuminating Safety and Effectiveness in Clinical Domains Towards Expert-Level Medical Question Answering with Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T10:46:10.475973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:46:10.475973Z digest=sha256:c7bef224441c70d952bcaf1d6a543cb2d29cbd73dbfb0039cd1dff24967c803d

Observation caf14afb-0b5e-4554-8baa-f677b1b8b321 · outbound

This paper cites et al.An evaluation framework for clinical use of large language models in patient interaction tasks.

A Novel Evaluation Benchmark for Medical LLMs: Illuminating Safety and Effectiveness in Clinical Domains et al.An evaluation framework for clinical use of large language models in patient interaction tasks

Reference 20

Resolution
verified exact
doi, observed 2026-08-06T10:46:10.597414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T10:46:10.479077Z digest=sha256:bfd9c36a768a09895d704f79c56a65877826e09bbaa3ef264e390f3a5cfca653

Observation 8465f969-37a8-4a75-8d95-49740f439ab8 · outbound

This paper cites Towards Conversational Diagnostic AI.

A Novel Evaluation Benchmark for Medical LLMs: Illuminating Safety and Effectiveness in Clinical Domains Towards Conversational Diagnostic AI

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T10:46:10.482561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:46:10.482561Z digest=sha256:1369a9dac916fc8750c5c9123ec62f7c37adc8677f051a792b927f0c85801c2a

Observation d7748517-aa6a-46e7-9df6-e100b55de009 · outbound

This paper cites AgentClinic: a multimodal agent benchmark to evaluate AI in simulated clinical environments.

A Novel Evaluation Benchmark for Medical LLMs: Illuminating Safety and Effectiveness in Clinical Domains AgentClinic: a multimodal agent benchmark to evaluate AI in simulated clinical environments

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T10:46:10.485794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:46:10.485794Z digest=sha256:f61ee681811a9aa0d17922b217ca9d721e6a80d03e6274a9c694346e30b23807

Observation 64a25829-9479-42dd-bc6f-17829ba377d9 · outbound

This paper cites An Automatic Evaluation Framework for Multi-turn Medical Consultations Capabilities of Large Language Models.

A Novel Evaluation Benchmark for Medical LLMs: Illuminating Safety and Effectiveness in Clinical Domains An Automatic Evaluation Framework for Multi-turn Medical Consultations Capabilities of Large Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T10:46:10.489111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:46:10.489111Z digest=sha256:1cac30ace6d7b22e0d091d89d225bc7f16d5b44df55fb8305e60bc63a0592358

Observation 2edab784-dfb4-414f-9e5b-c691a60f961d · outbound

This paper cites LLM-Mini-CEX: Automatic Evaluation of Large Language Model for Diagnostic Conversation.

A Novel Evaluation Benchmark for Medical LLMs: Illuminating Safety and Effectiveness in Clinical Domains LLM-Mini-CEX: Automatic Evaluation of Large Language Model for Diagnostic Conversation

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-08-06T10:46:10.842469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T10:46:10.492712Z digest=sha256:f61ff58a2b7d97397d4ab994427e57da67bbd78007a9346d762e69cb21eb5fb4

Observation 99cb9ef8-8cff-4929-a99b-cee4695c1741 · outbound

This paper cites an unresolved cited work.

A Novel Evaluation Benchmark for Medical LLMs: Illuminating Safety and Effectiveness in Clinical Domains Unresolved cited work

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T10:46:10.496554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:46:10.496554Z digest=sha256:32d8ed9b3fb3d6e90c961d28bd2fe3bd95b52a27f8e851ee9671d0648d5074f9

Observation 844963ff-1f99-4258-9bc5-053685c07377 · outbound

This paper cites Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena.

A Novel Evaluation Benchmark for Medical LLMs: Illuminating Safety and Effectiveness in Clinical Domains Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T10:46:10.500484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:46:10.500484Z digest=sha256:ddf99139d233a3988429fd50ac280546a13ca6f2a5be1cd4fcc6128d8cb983f2

Observation 770d0a68-fd0e-4ec0-b153-d764739537eb · outbound

This paper cites et al.Automating Evaluation of AI Text Generation in Healthcare with a Large Language Model (LLM)-as-a-Judge.

A Novel Evaluation Benchmark for Medical LLMs: Illuminating Safety and Effectiveness in Clinical Domains et al.Automating Evaluation of AI Text Generation in Healthcare with a Large Language Model (LLM)-as-a-Judge

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T10:46:10.504332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:46:10.504332Z digest=sha256:22c25a4411a0cd6df9a68de9ce5c8c629d0fea12d5743fe9a4d8fdc584a2afe2

Observation 6a955e62-4ee5-45fb-9a2a-0f821700a341 · outbound

This paper cites an unresolved cited work.

A Novel Evaluation Benchmark for Medical LLMs: Illuminating Safety and Effectiveness in Clinical Domains Unresolved cited work

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T10:46:10.508681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:46:10.508681Z digest=sha256:40416b0262d24027e2c6cb3e8abb915d3d4c0a8db7ca4873e9bc37b4dd35fa3f

Observation ce190359-ddaf-443c-aab8-c7334c01cedb · outbound

This paper cites an unresolved cited work.

A Novel Evaluation Benchmark for Medical LLMs: Illuminating Safety and Effectiveness in Clinical Domains Unresolved cited work

Reference 29

Resolution
metadata mismatch
raw_fallback, observed 2026-08-06T10:46:10.815851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T10:46:10.512605Z digest=sha256:7d88ed6b262adb44ad7419d4ce5c3a1ec60f1c830efad7036fb2ce7e79b15e17

Observation 4cd591a4-3e82-4aee-904a-b4e9ae152850 · outbound

This paper cites MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks.

A Novel Evaluation Benchmark for Medical LLMs: Illuminating Safety and Effectiveness in Clinical Domains MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T10:46:10.515710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:46:10.515710Z digest=sha256:0c2f05233661e8cc1681aaf987baa9da206eb6634b4ba922e0e9e8d1fd4a54b0

Observation 314fc882-ee5b-4b39-a5d5-3b617bf6c05a · outbound

This paper cites et al.Aligning Large Language Models with Humans: A Comprehensive Survey of ChatGPT's Aptitude in Pharmacology.

A Novel Evaluation Benchmark for Medical LLMs: Illuminating Safety and Effectiveness in Clinical Domains et al.Aligning Large Language Models with Humans: A Comprehensive Survey of ChatGPT's Aptitude in Pharmacology

Reference 31

Resolution
verified exact
doi, observed 2026-08-06T10:46:10.562934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T10:46:10.519347Z digest=sha256:815b53143e7ff808e18bf863fb30b4472abc8101010f843203ad708b5af93e64

Observation 8bef7af2-1cd9-43f7-b755-df775440a5c4 · outbound

This paper cites T., Bardak, A.

A Novel Evaluation Benchmark for Medical LLMs: Illuminating Safety and Effectiveness in Clinical Domains T., Bardak, A

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T10:46:10.522415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:46:10.522415Z digest=sha256:3901cf7d61a54c942df624f09fcb6700f384b453fa3685b179f7569055daad23

Pith citing papers

Observation 65b203d6-9762-4c96-a5ff-584bbaaaac63 · inbound

IatroBench: Pre-Registered Evidence of Iatrogenic Harm from AI Safety Measures cites this paper.

IatroBench: Pre-Registered Evidence of Iatrogenic Harm from AI Safety Measures A Novel Evaluation Benchmark for Medical LLMs: Illuminating Safety and Effectiveness in Clinical Domains

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:35:52.677796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T18:25:53.037936Z digest=sha256:ea211e502bdba92757e9e996a854d60c3a567c5be2ee9d0648fba1d1f3ee75fa

Observation 0444d725-a30e-4366-a6b4-685e8958ff48 · inbound

IatroBench: Pre-Registered Evidence of Iatrogenic Harm from AI Safety Measures cites this paper.

IatroBench: Pre-Registered Evidence of Iatrogenic Harm from AI Safety Measures A Novel Evaluation Benchmark for Medical LLMs: Illuminating Safety and Effectiveness in Clinical Domains

Reference 34

Resolution
unresolved
no resolver link, observed 2026-07-13T00:19:33.861692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T00:19:33.861692Z digest=sha256:b58413b0cc303fee43bd08d15d508b4038b8b2424ebc2a383641f343dd7ad377

Observation ad62a974-59d9-428c-b042-6fd8c90d11cd · inbound

MedVIGIL: Evaluating Trustworthy Medical VLMs Under Broken Visual Evidence cites this paper.

MedVIGIL: Evaluating Trustworthy Medical VLMs Under Broken Visual Evidence A Novel Evaluation Benchmark for Medical LLMs: Illuminating Safety and Effectiveness in Clinical Domains

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:45:55.214681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T03:45:42.703353Z digest=sha256:8615501d5a489ced3c7547d94a6a2cb5d4673bae0df95ea753a622e67ca4b3ff

Observation 88a75dbd-b112-4ed7-ab12-6276381603a5 · inbound

MedVIGIL: Evaluating Trustworthy Medical VLMs Under Broken Visual Evidence cites this paper.

MedVIGIL: Evaluating Trustworthy Medical VLMs Under Broken Visual Evidence A Novel Evaluation Benchmark for Medical LLMs: Illuminating Safety and Effectiveness in Clinical Domains

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-25T06:20:24.309424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-25T06:19:11.597882Z digest=sha256:f14063b04b49b15ff749b4176b87046a392ac2d939d74a22a26ead762e1c7057

Observation 8018e963-9bd1-4870-a88d-8634dea74c4c · inbound

Deployment-Centered Evaluation: Predicting Query-Level Rejection Risk in a Clinical LLM System cites this paper.

Deployment-Centered Evaluation: Predicting Query-Level Rejection Risk in a Clinical LLM System A Novel Evaluation Benchmark for Medical LLMs: Illuminating Safety and Effectiveness in Clinical Domains

Reference 37

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T11:08:02.950590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T09:43:12.633039Z digest=sha256:40b2c1d352b8d079c4b74059884f8f7e88a4ae112b8595fef06fd8775532664e