Pith. sign in

Paper Citation Record · LEDGER

MedHallu: A Comprehensive Benchmark for Detecting Medical Hallucinations in Large Language Models

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 13 inbound Pith citation observations for arXiv:2502.14302.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.14302 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 13 of 13 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:50:39.181332Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-28T17:22:25.043169Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 616bc472-ce89-4b3b-84ed-06d0e390e60c · inbound

Teaching with Lies: Curriculum DPO on Synthetic Negatives for Hallucination Detection cites this paper.

Teaching with Lies: Curriculum DPO on Synthetic Negatives for Hallucination Detection MedHallu: A Comprehensive Benchmark for Detecting Medical Hallucinations in Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T14:50:39.181332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:50:39.181332Z digest=sha256:78a83915920f75048501fec52de422f799e069fdd15a303f26e9076d9c93ba69

Observation 0fc89b08-7783-49e2-8fbb-8724936e0e93 · inbound

MTCMB: A Multi-Task Benchmark Framework for Evaluating LLMs on Knowledge, Reasoning, and Safety in Traditional Chinese Medicine cites this paper.

MTCMB: A Multi-Task Benchmark Framework for Evaluating LLMs on Knowledge, Reasoning, and Safety in Traditional Chinese Medicine MedHallu: A Comprehensive Benchmark for Detecting Medical Hallucinations in Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T11:50:54.319986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:50:54.319986Z digest=sha256:7ade08dbdd0e7955c67bde51ce5b6e36b6c24c6c49f51fc70e967a4a7d67c294

Observation dee28d99-0f05-47a3-ab70-c6ddcfb66ee1 · inbound

MIRIAD: Augmenting LLMs with millions of medical query-response pairs cites this paper.

MIRIAD: Augmenting LLMs with millions of medical query-response pairs MedHallu: A Comprehensive Benchmark for Detecting Medical Hallucinations in Large Language Models

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T06:04:43.698234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:04:43.698234Z digest=sha256:955f10b564bad2420cd171c69dc4e4f07a11dae399370100cf084b46a4dfa66b

Observation bce7b73c-af7a-49e4-a45d-2a4b95922a44 · inbound

A Comprehensive Survey of Electronic Health Record Modeling: From Deep Learning Approaches to Large Language Models cites this paper.

A Comprehensive Survey of Electronic Health Record Modeling: From Deep Learning Approaches to Large Language Models MedHallu: A Comprehensive Benchmark for Detecting Medical Hallucinations in Large Language Models

Reference 222

Resolution
unresolved
no resolver link, observed 2026-08-06T16:42:40.509104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:42:40.509104Z digest=sha256:b00f66f289b38cd0fce3faf698f35fa5b6052043a50dc94aae111f498b24484a

Observation 43c86410-77e4-4d81-931d-bf04c7ef53fc · inbound

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming cites this paper.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming MedHallu: A Comprehensive Benchmark for Detecting Medical Hallucinations in Large Language Models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:41.098401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:41.098401Z digest=sha256:e9a6356a85ad6fb6bd11d4c66a9a1542dfc023e2485d1618c8b0cd589e0f97c5

Observation e18ba6a1-0946-4f75-8de4-93c2923e1f23 · inbound

A comprehensive taxonomy of hallucinations in Large Language Models cites this paper.

A comprehensive taxonomy of hallucinations in Large Language Models MedHallu: A Comprehensive Benchmark for Detecting Medical Hallucinations in Large Language Models

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-06T05:29:17.220851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:29:17.220851Z digest=sha256:31d3a1698873635003ed294711b67458f469890bd13872fef77e2a7a9ea1fe5f

Observation 3edd2e50-c5a4-41df-b9db-8008de7a93d5 · inbound

ADRD-Bench: A Preliminary LLM Benchmark for Alzheimer's Disease and Related Dementias cites this paper.

ADRD-Bench: A Preliminary LLM Benchmark for Alzheimer's Disease and Related Dementias MedHallu: A Comprehensive Benchmark for Detecting Medical Hallucinations in Large Language Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-03T00:11:51.228826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:11:51.228826Z digest=sha256:bd54ea5acaec6eb089b31ff1057edf03649f68e6e688011576f06486d195de10

Observation 5fc6f648-32c3-4f31-a0dd-df2e90a4cab6 · inbound

A Multi-Stage Validation Framework for Trustworthy Large-scale Clinical Information Extraction using Large Language Models cites this paper.

A Multi-Stage Validation Framework for Trustworthy Large-scale Clinical Information Extraction using Large Language Models MedHallu: A Comprehensive Benchmark for Detecting Medical Hallucinations in Large Language Models

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:25:51.499750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T19:10:24.243459Z digest=sha256:2c6c62aaa0414584d8776cc74769db6c9f97dafe3512394c387efb86ccc32499

Observation 98c356bb-ac41-4c5b-b730-dfa05d2da826 · inbound

MedFabric and EtHER: A Data-Centric Framework for Word-Level Fabrication Generation and Detection in Medical LLMs cites this paper.

MedFabric and EtHER: A Data-Centric Framework for Word-Level Fabrication Generation and Detection in Medical LLMs MedHallu: A Comprehensive Benchmark for Detecting Medical Hallucinations in Large Language Models

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:16:09.093897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T17:43:52.309129Z digest=sha256:c4cb762c469ee57126ead7d75a574050c009e5dc31d1cf1690ce90b86bf5857d

Observation 06b7bafa-e995-49de-8378-9fb777edb307 · inbound

Hallucination Detection via Activations of Open-Weight Proxy Analyzers cites this paper.

Hallucination Detection via Activations of Open-Weight Proxy Analyzers MedHallu: A Comprehensive Benchmark for Detecting Medical Hallucinations in Large Language Models

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-11T02:45:58.153199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-11T02:43:07.105363Z digest=sha256:854cacb8c2f36af03179d017a7a3f6534c07c7e20bd808bb69d273d9c5bd63c0

Observation 9cf5cea6-78d0-4478-98d7-9747fad40877 · inbound

Graph Alignment Topology as an Inductive Bias for Grounding Detection cites this paper.

Graph Alignment Topology as an Inductive Bias for Grounding Detection MedHallu: A Comprehensive Benchmark for Detecting Medical Hallucinations in Large Language Models

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-25T05:55:24.062128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-25T05:54:55.236735Z digest=sha256:a50d6b81c95ea5bb7fcac941c026f4d475d5f846d9ac5c97eb03838dfb139d9b

Observation 9a86bc4e-9802-4ae4-b2e0-474de2a7107d · inbound

Med-HEAL: Analyzing and Mitigating Hallucinations in Medical LLMs with Hallucination-Aware In-Context Learning cites this paper.

Med-HEAL: Analyzing and Mitigating Hallucinations in Medical LLMs with Hallucination-Aware In-Context Learning MedHallu: A Comprehensive Benchmark for Detecting Medical Hallucinations in Large Language Models

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-06-28T17:22:25.044721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T17:18:07.068055Z digest=sha256:39518030f3f3962671a1a3e427403aa928946d352d2cb45ffff6c0c7fbff9590

Observation dd963348-d297-476c-ba69-ee27f02649c5 · inbound

CLIR-Bench: Benchmarking Multimodal Question Answering over Irregular Clinical Time Series cites this paper.

CLIR-Bench: Benchmarking Multimodal Question Answering over Irregular Clinical Time Series MedHallu: A Comprehensive Benchmark for Detecting Medical Hallucinations in Large Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-07-14T14:52:42.813309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:52:42.813309Z digest=sha256:39ee1c65f1dfd6c5977a6b83655b2180e9d3ba17d44fa85e3e8ef87d2bdf7719