Pith. sign in

Paper Citation Record · LEDGER

CliBench: A Multifaceted and Multigranular Evaluation of Large Language Models for Clinical Decision Making

As of 21 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 6 inbound Pith citation observations for arXiv:2406.09923.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.09923 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 6 of 6 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T16:49:58.033699Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T05:56:40.768689Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation ee9db274-6580-4991-a46e-72e880079c8c · inbound

Unlocking Historical Clinical Trial Data with ALIGN: A Compositional Large Language Model System for Medical Coding cites this paper.

Unlocking Historical Clinical Trial Data with ALIGN: A Compositional Large Language Model System for Medical Coding CliBench: A Multifaceted and Multigranular Evaluation of Large Language Models for Clinical Decision Making

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T16:49:58.033699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:49:58.033699Z digest=sha256:2965cf2deaa1dcab6ca1211d4e4308ca1b3bbabdeafb0021df6ad7edcb4e4ea6

Observation 499ed902-b58c-481e-8342-e8dc74807f84 · inbound

DiaLLMs: EHR Enhanced Clinical Conversational System for Clinical Test Recommendation and Diagnosis Prediction cites this paper.

DiaLLMs: EHR Enhanced Clinical Conversational System for Clinical Test Recommendation and Diagnosis Prediction CliBench: A Multifaceted and Multigranular Evaluation of Large Language Models for Clinical Decision Making

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T23:03:28.619956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:03:28.619956Z digest=sha256:ab814597a33723ab310639e11df45d9d79cc6f8dad9f2a5d625638b07d052451

Observation 57b2a74b-1710-4fa6-923d-5fa37199ebb2 · inbound

RealICU: Do LLM Agents Understand Long-Context ICU Data? A Benchmark Beyond Behavior Imitation cites this paper.

RealICU: Do LLM Agents Understand Long-Context ICU Data? A Benchmark Beyond Behavior Imitation CliBench: A Multifaceted and Multigranular Evaluation of Large Language Models for Clinical Decision Making

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-14T18:49:23.623196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-14T18:47:46.239810Z digest=sha256:5f18fd7fcf1beb5f5d3529ba3bb482acbf18c1aea19af7b21a1d5cb2d723f574

Observation 9d86b0ff-14f0-4611-bc9f-039c9f9b2163 · inbound

RxEval: A Prescription-Level Benchmark for Evaluating LLM Medication Recommendation cites this paper.

RxEval: A Prescription-Level Benchmark for Evaluating LLM Medication Recommendation CliBench: A Multifaceted and Multigranular Evaluation of Large Language Models for Clinical Decision Making

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-15T02:09:38.826201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-15T02:08:53.361260Z digest=sha256:69a48c8661e84265287aa3819af8068c1eca3128cce2820d4446cba6142ebca2

Observation 3498cba2-8389-4c36-adac-68607ee9c49e · inbound

Med-HEAL: Analyzing and Mitigating Hallucinations in Medical LLMs with Hallucination-Aware In-Context Learning cites this paper.

Med-HEAL: Analyzing and Mitigating Hallucinations in Medical LLMs with Hallucination-Aware In-Context Learning CliBench: A Multifaceted and Multigranular Evaluation of Large Language Models for Clinical Decision Making

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-06-28T17:22:25.052883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-28T17:18:07.068055Z digest=sha256:8bc061c5835e254192b4681433294a91498b701a88087335ee81d291642d46e2

Observation 30442862-ac91-4096-8744-a23a501a6d80 · inbound

D2MDT: Department-aware Multidisciplinary Team Consultation with Deliberation for Efficient Clinical Prediction cites this paper.

D2MDT: Department-aware Multidisciplinary Team Consultation with Deliberation for Efficient Clinical Prediction CliBench: A Multifaceted and Multigranular Evaluation of Large Language Models for Clinical Decision Making

Reference 98

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T05:56:40.770259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-06-28T07:52:47.812215Z digest=sha256:ac4ffb5cd6f28296b22bc8b8fff0c6f988d1d701e03c0b51596e02c43392fd77