Pith. sign in

Paper Citation Record · LEDGER

The Hallucinations Leaderboard -- An Open Effort to Measure Hallucinations in Large Language Models

As of 21 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 8 inbound Pith citation observations for arXiv:2404.05904.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2404.05904 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 8 of 8 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T19:08:01.731694Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

6
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 65d19ca8-5988-448b-939d-d1d27a8416f8 · inbound

ViBe: A Text-to-Video Benchmark for Evaluating Hallucination in Large Multimodal Models cites this paper.

ViBe: A Text-to-Video Benchmark for Evaluating Hallucination in Large Multimodal Models The Hallucinations Leaderboard -- An Open Effort to Measure Hallucinations in Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T19:17:10.550198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:17:10.550198Z digest=sha256:50d30d985595c2b6135d93a45d50d9ba12e76dd460e77512730bc3f1a7faaee6

Observation 6a3e447b-20c5-4766-8e9d-17884e552a98 · inbound

Self-Training Large Language Models for Tool-Use Without Demonstrations cites this paper.

Self-Training Large Language Models for Tool-Use Without Demonstrations The Hallucinations Leaderboard -- An Open Effort to Measure Hallucinations in Large Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-08T17:40:56.231844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:40:56.231844Z digest=sha256:5cb36b6d2539f947118bb5f7d5e93068b559a0bed272f13627aa5a25bcca96e6

Observation 74ec6aa7-33f2-41cd-95be-175a30dc88df · inbound

Expect the Unexpected: FailSafe Long Context QA for Finance cites this paper.

Expect the Unexpected: FailSafe Long Context QA for Finance The Hallucinations Leaderboard -- An Open Effort to Measure Hallucinations in Large Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-08T15:54:43.130182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:54:43.130182Z digest=sha256:ac1fc64458594c695b781ae6b92458c1327005f75736d843a5c30fa3d805feb5

Observation 9ae3ded5-0dbb-4794-accb-8e714e4015cf · inbound

The Science of Evaluating Foundation Models cites this paper.

The Science of Evaluating Foundation Models The Hallucinations Leaderboard -- An Open Effort to Measure Hallucinations in Large Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.698050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.698050Z digest=sha256:8981dc456901e62c5529d0fd6935290b44c158ac05424462fbbd00a5a8ddae5b

Observation fa21df7c-74da-4bf7-9a59-e14f053de7b8 · inbound

XtraGPT: Context-Aware and Controllable Academic Paper Revision via Human-AI Collaboration cites this paper.

XtraGPT: Context-Aware and Controllable Academic Paper Revision via Human-AI Collaboration The Hallucinations Leaderboard -- An Open Effort to Measure Hallucinations in Large Language Models

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-22T14:44:54.462418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-22T14:43:19.221814Z digest=sha256:4485c6bdeebdffb3e848be7bc352632e673c7250496b3601bc478607f34d4142

Observation df185e1e-b0b9-4f9f-8787-5edaf4ed2544 · inbound

HIDE and Seek: Detecting Hallucinations in Language Models via Decoupled Representations cites this paper.

HIDE and Seek: Detecting Hallucinations in Language Models via Decoupled Representations The Hallucinations Leaderboard -- An Open Effort to Measure Hallucinations in Large Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T19:08:01.731694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:08:01.731694Z digest=sha256:6df0c7fe78e12aee2367e1857032f83cccad089c2156b61e99325159f770ef00

Observation f8f2e2ec-1786-4ee5-bd9e-b853c47256f5 · inbound

Rethinking Evaluation for LLM Hallucination Detection: A Desiderata, A New RAG-based Benchmark, New Insights cites this paper.

Rethinking Evaluation for LLM Hallucination Detection: A Desiderata, A New RAG-based Benchmark, New Insights The Hallucinations Leaderboard -- An Open Effort to Measure Hallucinations in Large Language Models

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T01:37:03.587462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-13T01:34:43.464899Z digest=sha256:f716332f459dd00dac3a15e7c4cce752689895798200e21d08f110641d3127a5

Observation deecb24e-55f2-4a87-92df-258daa9a45e8 · inbound

REALISTA: Realistic Latent Adversarial Attacks that Elicit LLM Hallucinations cites this paper.

REALISTA: Realistic Latent Adversarial Attacks that Elicit LLM Hallucinations The Hallucinations Leaderboard -- An Open Effort to Measure Hallucinations in Large Language Models

Reference 173

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:17:56.398975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-14T20:13:10.814899Z digest=sha256:36dfef1c320bd4277411b2ccde38cbf44c995922e892ad3ded9b1dd687401b9d