Pith. sign in

Paper Citation Record · LEDGER

MEDIC: Comprehensive Evaluation of Leading Indicators for LLM Safety and Utility in Clinical Applications

As of 19 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 12 inbound Pith citation observations for arXiv:2409.07314.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2409.07314 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 12 of 12 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 12 of 12 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T10:28:11.073411Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

7
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 9a12a93e-e86c-4160-86de-5e0312108afa · inbound

Bridging Language Barriers in Healthcare: A Study on Arabic LLMs cites this paper.

Bridging Language Barriers in Healthcare: A Study on Arabic LLMs MEDIC: Comprehensive Evaluation of Leading Indicators for LLM Safety and Utility in Clinical Applications

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T19:41:02.187844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:41:02.187844Z digest=sha256:1adcf32180eaea0729bc70d2863a3299594f4be2e4bb761d752131cdb31e9b3a

Observation 3d3cc3b0-46af-4e9a-9014-4905e879e5b6 · inbound

Automatic Evaluation of Healthcare LLMs Beyond Question-Answering cites this paper.

Automatic Evaluation of Healthcare LLMs Beyond Question-Answering MEDIC: Comprehensive Evaluation of Leading Indicators for LLM Safety and Utility in Clinical Applications

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-08T14:46:29.636932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:46:29.636932Z digest=sha256:aa9249fd4d78befa1bdc75afccc81d0d37ddd4c2824c4c49e18353b9a2d43be9

Observation 48b23760-3ce1-4fc0-8875-80697cb08a23 · inbound

Stabilizing Reasoning in Medical LLMs with Continued Pretraining and Reasoning Preference Optimization cites this paper.

Stabilizing Reasoning in Medical LLMs with Continued Pretraining and Reasoning Preference Optimization MEDIC: Comprehensive Evaluation of Leading Indicators for LLM Safety and Utility in Clinical Applications

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-16T10:28:11.073411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:28:11.073411Z digest=sha256:9079c30c97c70dc84189447206f96d22b5f97108c89e2390ec024c5b726385e0

Observation 309f45aa-3e2b-4f55-9588-c6017336445b · inbound

MedArabiQ: Benchmarking Large Language Models on Arabic Medical Tasks cites this paper.

MedArabiQ: Benchmarking Large Language Models on Arabic Medical Tasks MEDIC: Comprehensive Evaluation of Leading Indicators for LLM Safety and Utility in Clinical Applications

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T23:55:11.836006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:55:11.836006Z digest=sha256:2875bb698740e830d2940216851568f74ec3ddccbf77102bb4800b9fe52fe89f

Observation eb689811-c640-4b1f-bf1d-1e35e99e3f29 · inbound

The Aloe Family Recipe for Open and Specialized Healthcare LLMs cites this paper.

The Aloe Family Recipe for Open and Specialized Healthcare LLMs MEDIC: Comprehensive Evaluation of Leading Indicators for LLM Safety and Utility in Clinical Applications

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-15T23:36:30.024034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:36:30.024034Z digest=sha256:5e1095c76fc81a6573b5b75773d8113e08a6f07fb53bb20124c9b8125414b92e

Observation 46b88ea4-36b7-4e71-a9ad-d96715b6329a · inbound

Can LLMs Reason Structurally? Benchmarking via the Lens of Data Structures cites this paper.

Can LLMs Reason Structurally? Benchmarking via the Lens of Data Structures MEDIC: Comprehensive Evaluation of Leading Indicators for LLM Safety and Utility in Clinical Applications

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T12:42:39.012433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:42:39.012433Z digest=sha256:fad248318f1001169a73232fb4f2458082eacd0d995c703409c32ec0ffa87b13

Observation 09b62148-7810-4c8f-9bbd-69eb80beb5c4 · inbound

ClinBench-HPB: A Clinical Benchmark for Evaluating LLMs in Hepato-Pancreato-Biliary Diseases cites this paper.

ClinBench-HPB: A Clinical Benchmark for Evaluating LLMs in Hepato-Pancreato-Biliary Diseases MEDIC: Comprehensive Evaluation of Leading Indicators for LLM Safety and Utility in Clinical Applications

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T12:28:01.247662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:28:01.247662Z digest=sha256:36b050a5c7c35e8cdb422966fb24cfee55bb0370d6b3a233f43ec30bb93a4f1e

Observation d9c327ef-6b38-42ba-8014-72870728a54a · inbound

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making cites this paper.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making MEDIC: Comprehensive Evaluation of Leading Indicators for LLM Safety and Utility in Clinical Applications

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.919409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.919409Z digest=sha256:0fb9f8a64a99c57da151663a851ef7bb27fec3f7706fe8453a755cfbccdfa8d8

Observation 194980ef-8d66-47ce-885c-05672c89046a · inbound

HIVMedQA: Benchmarking large language models for HIV medical decision support cites this paper.

HIVMedQA: Benchmarking large language models for HIV medical decision support MEDIC: Comprehensive Evaluation of Leading Indicators for LLM Safety and Utility in Clinical Applications

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T14:42:18.438834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:42:18.438834Z digest=sha256:2395774d35d453ff35082dae380c88c50b9c98f65d9d0fd0a2598c8dc52b30c4

Observation 950e3daa-e671-410a-b583-02f898882b4b · inbound

Development and Preliminary Evaluation of a Domain-Specific Large Language Model for Tuberculosis Care in South Africa cites this paper.

Development and Preliminary Evaluation of a Domain-Specific Large Language Model for Tuberculosis Care in South Africa MEDIC: Comprehensive Evaluation of Leading Indicators for LLM Safety and Utility in Clinical Applications

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-07-22T03:21:54.600556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-14T22:48:41.826046Z digest=sha256:0b23a1129add03fef3e3a73c3f100ce0f58dc5b1dc176d073795a89a07502fa2

Observation 5fb30f76-7f66-4807-9636-3942a95ae5f0 · inbound

Evaluating Patient Safety Risks in Generative AI: Development and Validation of a FMECA Framework for Generated Clinical Content cites this paper.

Evaluating Patient Safety Risks in Generative AI: Development and Validation of a FMECA Framework for Generated Clinical Content MEDIC: Comprehensive Evaluation of Leading Indicators for LLM Safety and Utility in Clinical Applications

Reference 34

Resolution
metadata mismatch
arxiv_id, observed 2026-07-22T03:21:54.600556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-09T20:54:12.601057Z digest=sha256:75bb1018b35b80183f883dd51a55d7511aea13a23b0e1531e0f93f551175ba45

Observation 913a0366-970a-461d-a17a-14865e2d0ea9 · inbound

PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents cites this paper.

PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents MEDIC: Comprehensive Evaluation of Leading Indicators for LLM Safety and Utility in Clinical Applications

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-01T02:21:21.301494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:21:21.301494Z digest=sha256:0c95012d1afc8ca0fc2a2ce9221525a26189c3fb7b7e257dfc37aab2f1b4d92f