Pith. sign in

Paper Citation Record · LEDGER

Reasoning in Real World Clinical Care: Why Large Language Models Are Not Yet Safe for Autonomous Clinical Decision Support

As of 9 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 0 inbound Pith citation observations for arXiv:2607.28677.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.28677 v1

Coverage vector

measured 29 of 29 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-03T00:47:08.786613Z

measured 29 of 29 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

29 of 29 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved29
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation de96c79e-4a40-40a5-8179-5e64eac28911 · outbound

This paper cites Performance of a large language model on the reasoning tasks of a physician.

Reasoning in Real World Clinical Care: Why Large Language Models Are Not Yet Safe for Autonomous Clinical Decision Support Performance of a large language model on the reasoning tasks of a physician

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T00:47:05.752339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:47:05.752339Z digest=sha256:5447d312d03020ea8fcd4c5bd5b0c9c6b071eb1818c2043333be4d3c2d048433

Observation b74099ce-fab3-4452-b685-1b63fff2d913 · outbound

This paper cites Reliability of LLMs as medical assistants for the general public: a randomized preregistered study.

Reasoning in Real World Clinical Care: Why Large Language Models Are Not Yet Safe for Autonomous Clinical Decision Support Reliability of LLMs as medical assistants for the general public: a randomized preregistered study

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-03T00:47:05.840374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:47:05.840374Z digest=sha256:17f31fefea5cb217ae577f0dec3de83b440069b82ec75efdc8ef0d8a2c481f16

Observation 7b945728-8bc2-4755-9509-9281d59801e5 · outbound

This paper cites Measuring what Matters: Construct Validity in Large Language Model Benchmarks2025 November 3, 2025.

Reasoning in Real World Clinical Care: Why Large Language Models Are Not Yet Safe for Autonomous Clinical Decision Support Measuring what Matters: Construct Validity in Large Language Model Benchmarks2025 November 3, 2025

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-03T00:47:05.945108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:47:05.945108Z digest=sha256:fade59c707c06ad89183fe772d9c93cc45136fd64aa8a1a73a8c9a16eee68ddf

Observation fce241fd-572f-4cc1-9c1d-660504b4d9fd · outbound

This paper cites On the robustness of medical term representations in locally deployable language models.

Reasoning in Real World Clinical Care: Why Large Language Models Are Not Yet Safe for Autonomous Clinical Decision Support On the robustness of medical term representations in locally deployable language models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T00:47:06.111582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:47:06.111582Z digest=sha256:ae26e48a040fd231f502b8c48f6cc2f12fb2bf1ba1eaac4af6faad162b27a505

Observation f2784f09-c73c-48ff-b9bf-53c9ba25b0d8 · outbound

This paper cites Health AI needs meaningful human involvement: lessons from war.

Reasoning in Real World Clinical Care: Why Large Language Models Are Not Yet Safe for Autonomous Clinical Decision Support Health AI needs meaningful human involvement: lessons from war

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T00:47:06.220549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:47:06.220549Z digest=sha256:c424c5ad7cc94b13ff965357d5872524e89760d816aa7ec4b6e3e470c2976f2e

Observation ecee4c67-07d8-45fd-bd8a-8b4fe65eb86f · outbound

This paper cites From Concept to Clinic: Real World Evidence for Autonomous AI Deployment in Primary Care Telemedicine.

Reasoning in Real World Clinical Care: Why Large Language Models Are Not Yet Safe for Autonomous Clinical Decision Support From Concept to Clinic: Real World Evidence for Autonomous AI Deployment in Primary Care Telemedicine

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T00:47:06.389839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:47:06.389839Z digest=sha256:eeeb84e38e13d6ffa70fd02c7e257d432280a46e7dcbef180a0a0f4d78492cb1

Observation d5498374-9fa5-4b33-a0d7-26f93bedc42d · outbound

This paper cites Towards autonomous medical artificial intelligence agents.

Reasoning in Real World Clinical Care: Why Large Language Models Are Not Yet Safe for Autonomous Clinical Decision Support Towards autonomous medical artificial intelligence agents

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T00:47:06.486690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:47:06.486690Z digest=sha256:f1acd53d87def0b21196b881b072f77ae510d5d89d5a97fea818c5c9d5c338d1

Observation bfb0457c-5416-4141-822b-4a2cc3be5c05 · outbound

This paper cites A prospective clinical feasibility study of a conversational diagnostic AI in an ambulatory primary care clinic.

Reasoning in Real World Clinical Care: Why Large Language Models Are Not Yet Safe for Autonomous Clinical Decision Support A prospective clinical feasibility study of a conversational diagnostic AI in an ambulatory primary care clinic

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-03T00:47:06.542200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:47:06.542200Z digest=sha256:ce7b2c405b03b4b95e908ce2ced533922e019b0cfcd60e1fcf29d9b933075367

Observation a2dccb91-62a7-4e8d-932a-6e10c5fb2df4 · outbound

This paper cites Safety of a large language model-based clinical decision support system in African primary healthcare.

Reasoning in Real World Clinical Care: Why Large Language Models Are Not Yet Safe for Autonomous Clinical Decision Support Safety of a large language model-based clinical decision support system in African primary healthcare

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T00:47:06.666268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:47:06.666268Z digest=sha256:29fab3e90bfffd332b7bedd8bdd392cfc1d78cda0c42798301333f238ae3b96e

Observation 20518bf9-356d-4738-a5ad-8f65ee5cae15 · outbound

This paper cites AI-based Clinical Decision Support for Primary Care: A Real-World Study.

Reasoning in Real World Clinical Care: Why Large Language Models Are Not Yet Safe for Autonomous Clinical Decision Support AI-based Clinical Decision Support for Primary Care: A Real-World Study

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-03T00:47:06.754803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:47:06.754803Z digest=sha256:a6289399ec75259c0cba86075c52e764602e7b264aafb37bec2e8a918009678b

Observation 035c0676-8880-4aaf-bec3-b24b63b55bd4 · outbound

This paper cites Medical errors in large language models revealed using 1,000 synthetic clinical transcripts.

Reasoning in Real World Clinical Care: Why Large Language Models Are Not Yet Safe for Autonomous Clinical Decision Support Medical errors in large language models revealed using 1,000 synthetic clinical transcripts

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T00:47:06.851847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:47:06.851847Z digest=sha256:7885827032f9bc27646c911236b28b743d77dcbe17adfa38902c8a6bf4230c26

Observation c00148ae-0141-4595-8dba-67898da83f87 · outbound

This paper cites Large Language Models lack essential metacognition for reliable medical reasoning.

Reasoning in Real World Clinical Care: Why Large Language Models Are Not Yet Safe for Autonomous Clinical Decision Support Large Language Models lack essential metacognition for reliable medical reasoning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T00:47:06.907210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:47:06.907210Z digest=sha256:67199181058cc582d39bc5f581fe1ae5ed8774b0c9330637367ecb1ea5127622

Observation d85b0cde-3c7c-4c81-b0ec-817a1e640eee · outbound

This paper cites Towards conversational diagnostic artificial intelligence.

Reasoning in Real World Clinical Care: Why Large Language Models Are Not Yet Safe for Autonomous Clinical Decision Support Towards conversational diagnostic artificial intelligence

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T00:47:07.069220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:47:07.069220Z digest=sha256:b3bf7d61f5697f6e3a95592581a2545b11ef990ff4d70445ffce32567093ff2d

Observation 8dbb94b9-03f0-4e4f-a837-c683fd606076 · outbound

This paper cites Towards Conversational AI for Disease Management.

Reasoning in Real World Clinical Care: Why Large Language Models Are Not Yet Safe for Autonomous Clinical Decision Support Towards Conversational AI for Disease Management

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-03T00:47:07.234978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:47:07.234978Z digest=sha256:07a3aa68412464e85826042ce60a5f006df46aa2f268a290cc254b0eac1e0952

Observation 0b016d29-21a5-4973-a476-dc71e2ed4388 · outbound

This paper cites Testing and Evalua- tion of Health Care Applications of Large Language Models: A Systematic Review.

Reasoning in Real World Clinical Care: Why Large Language Models Are Not Yet Safe for Autonomous Clinical Decision Support Testing and Evalua- tion of Health Care Applications of Large Language Models: A Systematic Review

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-03T00:47:07.401396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:47:07.401396Z digest=sha256:d86179131ef17afc80aa7ce058c53c7b713587ac38a0bb758fd57f9ba3c83763

Observation b74fbffc-c8cd-4bce-a7db-44304f367166 · outbound

This paper cites LLM-assisted systematic review of large language models in clinical medicine.

Reasoning in Real World Clinical Care: Why Large Language Models Are Not Yet Safe for Autonomous Clinical Decision Support LLM-assisted systematic review of large language models in clinical medicine

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-03T00:47:07.513681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:47:07.513681Z digest=sha256:1c76d8cef016ade53bbcf2235f3c82c442a825d0cd19a573620ed2862a186794

Observation 0a5bcce6-3244-4658-b3c3-83e1f25ad143 · outbound

This paper cites Multi-model assurance analysis showing large language models are highly vulnerable to adversarial hallucination attacks during clinical decision support.

Reasoning in Real World Clinical Care: Why Large Language Models Are Not Yet Safe for Autonomous Clinical Decision Support Multi-model assurance analysis showing large language models are highly vulnerable to adversarial hallucination attacks during clinical decision support

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T00:47:07.674935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:47:07.674935Z digest=sha256:3d0b47e9493d2a64ae2f32705339a0209164ed6e1465b2c9eac47942da19053c

Observation 471de8cb-5145-4435-8013-9643d29439e8 · outbound

This paper cites MedMisBench: Measuring Epistemic Resilience of LLMs Under Misleading Medical Context.

Reasoning in Real World Clinical Care: Why Large Language Models Are Not Yet Safe for Autonomous Clinical Decision Support MedMisBench: Measuring Epistemic Resilience of LLMs Under Misleading Medical Context

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T00:47:07.840702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:47:07.840702Z digest=sha256:b6ab3cc1eaa809e2796c09025ecc6d739a60fbc52fa516ce418dbf4f3d246a2d

Observation b5515b52-0bcc-4319-bd09-816bbdb412ef · outbound

This paper cites Training language models to be warm can reduce accuracy and increase sycophancy.

Reasoning in Real World Clinical Care: Why Large Language Models Are Not Yet Safe for Autonomous Clinical Decision Support Training language models to be warm can reduce accuracy and increase sycophancy

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-03T00:47:07.900627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:47:07.900627Z digest=sha256:02fc9815468ef77f16609f2c5f61caa6e68fe31786364752dc0a3519a68e2885

Observation a4b24a9d-51dc-4923-96b4-ea1e1d9911cb · outbound

This paper cites Competing Biases underlie Overconfidence and Underconfidence in LLMs.

Reasoning in Real World Clinical Care: Why Large Language Models Are Not Yet Safe for Autonomous Clinical Decision Support Competing Biases underlie Overconfidence and Underconfidence in LLMs

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-03T00:47:08.008719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:47:08.008719Z digest=sha256:ac9fbc68a1b57d8ef9d59edcee6433d4f3192a7e359bc48f934f0293eb411744

Observation 4a28b52e-1033-4391-bfe9-df546d31a1dc · outbound

This paper cites Assessment of Large Language Models in Clinical Reasoning: A Novel Benchmarking Study.

Reasoning in Real World Clinical Care: Why Large Language Models Are Not Yet Safe for Autonomous Clinical Decision Support Assessment of Large Language Models in Clinical Reasoning: A Novel Benchmarking Study

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-03T00:47:08.039934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:47:08.039934Z digest=sha256:c882e63f89a38b46a1bdc7f1eb09be8f0d8068bdaab06fadd110669821eba94f

Observation 511ac63f-afcc-442a-a47c-4b86c6f4099d · outbound

This paper cites First, do NOHARM: towards clinically safe large language models.

Reasoning in Real World Clinical Care: Why Large Language Models Are Not Yet Safe for Autonomous Clinical Decision Support First, do NOHARM: towards clinically safe large language models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-03T00:47:08.117061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:47:08.117061Z digest=sha256:a78c6b6c7b47d6f1bc01f8340898a547f7bfa18acc9c0c033036feaf8111ea14

Observation 14f78564-fe69-4c59-bb6f-5d2f97640ba6 · outbound

This paper cites MediQ: Question-Asking LLMs and a Benchmark for Reliable Interactive Clinical Reasoning.

Reasoning in Real World Clinical Care: Why Large Language Models Are Not Yet Safe for Autonomous Clinical Decision Support MediQ: Question-Asking LLMs and a Benchmark for Reliable Interactive Clinical Reasoning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-03T00:47:08.177007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:47:08.177007Z digest=sha256:ad8a39bffb5ef69ac44eba6ea3b8418707855a41d217a0b6cad9529bd50b7b55

Observation 827f14c3-925e-4fe8-af3f-f3090208de68 · outbound

This paper cites SymptomAI: Toward a Con- versational AI Agent for Everyday Symptom Assessment.

Reasoning in Real World Clinical Care: Why Large Language Models Are Not Yet Safe for Autonomous Clinical Decision Support SymptomAI: Toward a Con- versational AI Agent for Everyday Symptom Assessment

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-03T00:47:08.285001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:47:08.285001Z digest=sha256:5081a050f06e8cf99006161fb4b1de97e48549fbc2e002179101cc817d7ee588

Observation 91c4f6c5-71d2-4782-9134-7b3028b411a7 · outbound

This paper cites A POMDP Formulation of Preference Elicitation Problems.

Reasoning in Real World Clinical Care: Why Large Language Models Are Not Yet Safe for Autonomous Clinical Decision Support A POMDP Formulation of Preference Elicitation Problems

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-03T00:47:08.396535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:47:08.396535Z digest=sha256:8ca89ce9ccfe699149790639c48dd7c61384bbdaf3706546d506cd470f651c02

Observation c7d11616-466c-4f77-8618-b75311bf09eb · outbound

This paper cites A clinical environment simulator for dynamic AI evaluation.

Reasoning in Real World Clinical Care: Why Large Language Models Are Not Yet Safe for Autonomous Clinical Decision Support A clinical environment simulator for dynamic AI evaluation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-03T00:47:08.504867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:47:08.504867Z digest=sha256:a185768fa5845a1e53d4d7475f36a60c8cd5cc6947389f11e7db0cb45ea3c187

Observation 993d9874-d12c-4727-b5dd-74d99e6f52a2 · outbound

This paper cites LLMs Don't Know Their Own Decision Boundaries: The Unreliability of Self-Generated Counterfactual Explanations.

Reasoning in Real World Clinical Care: Why Large Language Models Are Not Yet Safe for Autonomous Clinical Decision Support LLMs Don't Know Their Own Decision Boundaries: The Unreliability of Self-Generated Counterfactual Explanations

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-03T00:47:08.564508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:47:08.564508Z digest=sha256:c85be0b90d5e0b30b6c1f94f5759d3930be8bfd1b6e2d375660b5445df3ad321

Observation d50a2074-7081-42a9-9b88-1b298f12bbc8 · outbound

This paper cites AI, Health, and Health Care Today and Tomorrow: The JAMA Summit Report on Artificial Intelligence.

Reasoning in Real World Clinical Care: Why Large Language Models Are Not Yet Safe for Autonomous Clinical Decision Support AI, Health, and Health Care Today and Tomorrow: The JAMA Summit Report on Artificial Intelligence

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-03T00:47:08.624687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:47:08.624687Z digest=sha256:4e2c33d87b75a6d36b5333b0de006afd6fa3b90ddd5c20dbbe0ff86653cee14f

Observation abf90368-e3e8-47f5-b310-f25a8c54446e · outbound

This paper cites Nature Medicine.

Reasoning in Real World Clinical Care: Why Large Language Models Are Not Yet Safe for Autonomous Clinical Decision Support Nature Medicine

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-03T00:47:08.786613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:47:08.786613Z digest=sha256:c767f2f7508799c2e98bc6c0ace21e526573fae59924e69962800343e405c1fd

Pith citing papers

No inbound Pith citation observations are available.