Pith. sign in

Paper Citation Record · LEDGER

MedFuzz: Exploring the Robustness of Large Language Models in Medical Question Answering

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 5 inbound Pith citation observations for arXiv:2406.06573.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.06573 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 5 of 5 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:52:02.715263Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T19:07:17.626802Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 7f88f812-50ac-424a-a43f-d8e1f92f5859 · inbound

Ask Patients with Patience: Enabling LLMs for Human-Centric Medical Dialogue with Grounded Reasoning cites this paper.

Ask Patients with Patience: Enabling LLMs for Human-Centric Medical Dialogue with Grounded Reasoning MedFuzz: Exploring the Robustness of Large Language Models in Medical Question Answering

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-23T03:37:27.818864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-23T03:35:56.594095Z digest=sha256:26226cd530075319196a09cc056904b6178bdb945f086d122b825597d31b5b00

Observation fe13c786-3e20-4461-bf06-bc8efaca8a00 · inbound

MedSentry: Understanding and Mitigating Safety Risks in Medical LLM Multi-Agent Systems cites this paper.

MedSentry: Understanding and Mitigating Safety Risks in Medical LLM Multi-Agent Systems MedFuzz: Exploring the Robustness of Large Language Models in Medical Question Answering

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T13:52:02.715263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:52:02.715263Z digest=sha256:2a466d812f59d15a3f0f367ad68bf8728ebf0726024ba19379397bc09969cc55

Observation d2966f35-8df5-4396-8403-1ab3f0feb77c · inbound

Testing for LLM response differences: the case of a composite null consisting of semantically irrelevant query perturbations cites this paper.

Testing for LLM response differences: the case of a composite null consisting of semantically irrelevant query perturbations MedFuzz: Exploring the Robustness of Large Language Models in Medical Question Answering

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T17:26:38.350247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T17:26:38.350247Z digest=sha256:7fcf0e314faf71def4872f60e6790a22f6aa1e67110d8f0f731f332c4cdfa655

Observation 30da5c93-d000-4640-9d40-0f9ded6dda8c · inbound

MedDialBench: Benchmarking LLM Diagnostic Robustness under Parametric Adversarial Patient Behaviors cites this paper.

MedDialBench: Benchmarking LLM Diagnostic Robustness under Parametric Adversarial Patient Behaviors MedFuzz: Exploring the Robustness of Large Language Models in Medical Question Answering

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T23:50:56.650048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T18:49:51.261741Z digest=sha256:eb6a43dff3efee9e174570d7221dc349ccea8ad3e0174492ad1e467202d9d221

Observation 652670ec-ea8a-412d-a967-82db346d339c · inbound

When Large Language Models Fail in Healthcare: Evaluating Sensitivity to Prompt Variations cites this paper.

When Large Language Models Fail in Healthcare: Evaluating Sensitivity to Prompt Variations MedFuzz: Exploring the Robustness of Large Language Models in Medical Question Answering

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-02T19:07:17.628109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T21:41:20.354463Z digest=sha256:9409cb69d1c885d7b6a44713996e30211a39d143938577fde6131e78a3c5a28b