Pith. sign in

Paper Citation Record · LEDGER

The Hallucinations Leaderboard -- An Open Effort to Measure Hallucinations in Large Language Models

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 6 inbound Pith citation observations for arXiv:2404.05904.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2404.05904 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 6 of 6 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T17:40:56.231844Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

6
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 6a3e447b-20c5-4766-8e9d-17884e552a98 · inbound

Self-Training Large Language Models for Tool-Use Without Demonstrations cites this paper.

Self-Training Large Language Models for Tool-Use Without Demonstrations The Hallucinations Leaderboard -- An Open Effort to Measure Hallucinations in Large Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-08T17:40:56.231844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:40:56.231844Z digest=sha256:eb6fe333b5c1419e8ce64e570076c8e283194747eb7282eca99b69274e1fc825

Observation 74ec6aa7-33f2-41cd-95be-175a30dc88df · inbound

Expect the Unexpected: FailSafe Long Context QA for Finance cites this paper.

Expect the Unexpected: FailSafe Long Context QA for Finance The Hallucinations Leaderboard -- An Open Effort to Measure Hallucinations in Large Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-08T15:54:43.130182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:54:43.130182Z digest=sha256:9ea8d66ef360c1cf3e9a07a959f57934645bb063c26ea54737422329f74b65f6

Observation 9ae3ded5-0dbb-4794-accb-8e714e4015cf · inbound

The Science of Evaluating Foundation Models cites this paper.

The Science of Evaluating Foundation Models The Hallucinations Leaderboard -- An Open Effort to Measure Hallucinations in Large Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.698050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.698050Z digest=sha256:1ecfdcb277aa092495f819b97fae542d72729df58090acfd6bfdb0ad302a2362

Observation fa21df7c-74da-4bf7-9a59-e14f053de7b8 · inbound

XtraGPT: Context-Aware and Controllable Academic Paper Revision via Human-AI Collaboration cites this paper.

XtraGPT: Context-Aware and Controllable Academic Paper Revision via Human-AI Collaboration The Hallucinations Leaderboard -- An Open Effort to Measure Hallucinations in Large Language Models

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-22T14:44:54.462418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-22T14:43:19.221814Z digest=sha256:1806ce5bf80165f316d8c842f7f55face117dd9acf107730fbcd32720ba3bc66

Observation f8f2e2ec-1786-4ee5-bd9e-b853c47256f5 · inbound

Rethinking Evaluation for LLM Hallucination Detection: A Desiderata, A New RAG-based Benchmark, New Insights cites this paper.

Rethinking Evaluation for LLM Hallucination Detection: A Desiderata, A New RAG-based Benchmark, New Insights The Hallucinations Leaderboard -- An Open Effort to Measure Hallucinations in Large Language Models

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T01:37:03.587462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T01:34:43.464899Z digest=sha256:69705482fbdb0395be3adc26fcda293fe153bcdf42f30c367382d885b48d1258

Observation deecb24e-55f2-4a87-92df-258daa9a45e8 · inbound

REALISTA: Realistic Latent Adversarial Attacks that Elicit LLM Hallucinations cites this paper.

REALISTA: Realistic Latent Adversarial Attacks that Elicit LLM Hallucinations The Hallucinations Leaderboard -- An Open Effort to Measure Hallucinations in Large Language Models

Reference 173

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:17:56.398975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-14T20:13:10.814899Z digest=sha256:c5c754158803d426cbd35be9bc6a3656234cafa0a3ad2b3f276791fa552070c7