Pith. sign in

Paper Citation Record · LEDGER

Large Language Models in the Clinic: A Comprehensive Benchmark

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 6 inbound Pith citation observations for arXiv:2405.00716.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2405.00716 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 6 of 6 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:55:56.187110Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T21:38:58.254119Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 39d9ae3e-d96b-486d-8980-1520ef343181 · inbound

MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks cites this paper.

MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks Large Language Models in the Clinic: A Comprehensive Benchmark

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T13:55:56.187110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:55:56.187110Z digest=sha256:da3af50c9a64bf732bb0f29631007a0f1bd30efcbc6251580a1c7da4ad86c79a

Observation aa193aca-eb56-4e33-9a8d-a88fb74a919b · inbound

ImmunoFOMO: Are Language Models missing what oncologists see? cites this paper.

ImmunoFOMO: Are Language Models missing what oncologists see? Large Language Models in the Clinic: A Comprehensive Benchmark

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T04:07:42.889123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:07:42.889123Z digest=sha256:3cacd9fdaa92dd3813073b8863ed484b6f45239f5d374928fbd4406750a67de2

Observation 82c7e22c-8afc-46f9-b938-7edba28cadab · inbound

Towards the Next Frontier of LLMs, Training on Private Data: A Cross-Domain Benchmark for Federated Fine-Tuning cites this paper.

Towards the Next Frontier of LLMs, Training on Private Data: A Cross-Domain Benchmark for Federated Fine-Tuning Large Language Models in the Clinic: A Comprehensive Benchmark

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-15T04:59:45.909200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T04:55:24.482203Z digest=sha256:243f922dfbb16241732f476006d4e297c13ae589ec038038f4388bee1b8cf6a9

Observation 0836966d-fb21-497c-b2b3-791991e9362d · inbound

EHRBench: An Automated and Reliable EHR-based Benchmark for Clinical Decision Making with LLMs cites this paper.

EHRBench: An Automated and Reliable EHR-based Benchmark for Clinical Decision Making with LLMs Large Language Models in the Clinic: A Comprehensive Benchmark

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-06-29T06:43:10.396778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-29T06:41:06.828814Z digest=sha256:e03b223d4360344c71681dc15ad5d3691ee497b28b6b1c5aecc5e0bb98575d04

Observation dcc158fe-9db7-46db-8f09-3a105a042014 · inbound

Possible or Definite? A Benchmark for Evaluating Diagnostic Uncertainty Preservation in Clinical Text cites this paper.

Possible or Definite? A Benchmark for Evaluating Diagnostic Uncertainty Preservation in Clinical Text Large Language Models in the Clinic: A Comprehensive Benchmark

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-07-03T21:38:58.255535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-27T00:23:55.654447Z digest=sha256:6714a3147782ffe2c36e73284aecce534caef54eb68983a4559b08f973eb9669

Observation 183bae52-dff1-4a53-abfe-032c7dd9e5f8 · inbound

When Medical Safety Alignment Fails: A Benchmark for Evaluating LLMs on High-Risk Medical Queries cites this paper.

When Medical Safety Alignment Fails: A Benchmark for Evaluating LLMs on High-Risk Medical Queries Large Language Models in the Clinic: A Comprehensive Benchmark

Reference 56

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T11:04:37.463197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-30T11:03:52.230759Z digest=sha256:052145ee2fa4182187968d585b21063bafc7ad61e9a231b607d14f6b1389b1c1