Pith. sign in

Paper Citation Record · LEDGER

HalluLens: LLM Hallucination Benchmark

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 25 inbound Pith citation observations for arXiv:2504.17550.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.17550 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 25 of 25 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 25 of 25 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:44:30.304843Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 6c6d2e74-fa0b-4a8b-bb16-c419cb0eefa0 · inbound

The Hallucination Tax of Reinforcement Finetuning cites this paper.

The Hallucination Tax of Reinforcement Finetuning HalluLens: LLM Hallucination Benchmark

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T15:44:30.304843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:44:30.304843Z digest=sha256:1b79d5a14d47b108a114a515c663d6f1755bd2f2d0bb4550e3bbf355eb8b3be5

Observation 710392c0-8593-4199-9562-ece45783b8a7 · inbound

AbstentionBench: Reasoning LLMs Fail on Unanswerable Questions cites this paper.

AbstentionBench: Reasoning LLMs Fail on Unanswerable Questions HalluLens: LLM Hallucination Benchmark

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:08.125965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:08.125965Z digest=sha256:d559cd94358298e90eff9f485299a78aa61d593d6b6ad99ae84018febf9f933c

Observation 77e35e4d-2d50-473b-92a1-f8497ba24f36 · inbound

Machine Mirages: Defining the Undefined cites this paper.

Machine Mirages: Defining the Undefined HalluLens: LLM Hallucination Benchmark

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T11:24:29.959566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:24:29.959566Z digest=sha256:b473fb5832011f68fea0f030073debcc8076cb39d52474ead533afc6df855125

Observation dd8252b2-fc3d-4546-81ec-605e437dbb68 · inbound

Embodied AI Agents: Modeling the World cites this paper.

Embodied AI Agents: Modeling the World HalluLens: LLM Hallucination Benchmark

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T22:10:25.436137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:10:25.436137Z digest=sha256:fdafefd458d38856cf82b3dbc432e5429bb78b9f9a9af3fea61b5d11c2578987

Observation 37e98791-aba2-4f3b-9a78-ecf26545fd0c · inbound

Introducing the Swiss Food Knowledge Graph: AI for Context-Aware Nutrition Recommendation cites this paper.

Introducing the Swiss Food Knowledge Graph: AI for Context-Aware Nutrition Recommendation HalluLens: LLM Hallucination Benchmark

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T17:43:11.362519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:43:11.362519Z digest=sha256:7bcaf53af5925a202a1fa3d4abd1ee2f1a8b701d3c170b34d47b9f29197106e8

Observation c3bc9cd2-4707-482c-8568-42007d7edf06 · inbound

MIRAGE-Bench: LLM Agent is Hallucinating and Where to Find Them cites this paper.

MIRAGE-Bench: LLM Agent is Hallucinating and Where to Find Them HalluLens: LLM Hallucination Benchmark

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T13:06:36.599018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:06:36.599018Z digest=sha256:f533d383f20bdebfad77ef2ba3dfde045f830d30c6d8eb189be2ce42190c21cb

Observation 5831c5ed-3040-4a23-a8c2-3a38c4a34cce · inbound

A comprehensive taxonomy of hallucinations in Large Language Models cites this paper.

A comprehensive taxonomy of hallucinations in Large Language Models HalluLens: LLM Hallucination Benchmark

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T05:29:09.780706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:29:09.780706Z digest=sha256:4c463e937a91093c519eb2452e8f7b4b7d624851dcbf22d5d8d02d58e36abc51

Observation 08f453e8-d49c-4bdc-8a33-614d4f0addf4 · inbound

ReasoningTrack: Chain-of-Thought Reasoning for Long-term Vision-Language Tracking cites this paper.

ReasoningTrack: Chain-of-Thought Reasoning for Long-term Vision-Language Tracking HalluLens: LLM Hallucination Benchmark

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T23:33:19.965610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:33:19.965610Z digest=sha256:6621bbb4d2facaa523e7e8bd7f0d106edc15170ed450f05dc178046538f9329d

Observation 1caab12d-ef18-4070-a689-d66bcbeb9b4a · inbound

GOSU: Retrieval-Augmented Generation with Global-Level Optimized Semantic Unit-Centric Framework cites this paper.

GOSU: Retrieval-Augmented Generation with Global-Level Optimized Semantic Unit-Centric Framework HalluLens: LLM Hallucination Benchmark

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T13:36:35.132035Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:36:35.132035Z digest=sha256:eb01ed70e4a475a1cbfe7f13e8c05498cc221da673b7edb37b73dc3bfa05cab5

Observation 1e25fd92-76c3-4f24-9d4d-bd7323c38648 · inbound

ReFACT: A Benchmark for Scientific Confabulation Detection with Positional Error Annotations cites this paper.

ReFACT: A Benchmark for Scientific Confabulation Detection with Positional Error Annotations HalluLens: LLM Hallucination Benchmark

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-18T12:56:24.497247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-18T12:54:01.717015Z digest=sha256:114945a7cd33e2caffbca41106bb932c57f3f4671b3989ffcfa7832f6f553fbc

Observation eb0f60c0-8233-4d36-b187-45b6acb266dc · inbound

When Do Hallucinations Arise? A Graph Perspective on the Evolution of Path Reuse and Path Compression cites this paper.

When Do Hallucinations Arise? A Graph Perspective on the Evolution of Path Reuse and Path Compression HalluLens: LLM Hallucination Benchmark

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-13T13:07:00.928770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T13:07:00.928770Z digest=sha256:7e1fc2242e0e98367143dbbf33cc3946ada3413382dd69133119e45c2922b7e8

Observation 5cb32dba-21c3-4776-b09b-ea6890e3981c · inbound

From Binary Groundedness to Support Relations: Towards a Reader-Centred Taxonomy for Comprehension of AI Output cites this paper.

From Binary Groundedness to Support Relations: Towards a Reader-Centred Taxonomy for Comprehension of AI Output HalluLens: LLM Hallucination Benchmark

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:30:58.657451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T17:37:28.129376Z digest=sha256:a2ddbd905a93a8e6330dfd7294f0254103fa79b6e86600014a7cfe7de0b413b9

Observation 6be78a1e-743e-44e0-8a4b-e759714bc979 · inbound

HalluClear: Diagnosing, Evaluating and Mitigating Hallucinations in GUI Agents cites this paper.

HalluClear: Diagnosing, Evaluating and Mitigating Hallucinations in GUI Agents HalluLens: LLM Hallucination Benchmark

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-10T06:31:30.913700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T06:27:19.717895Z digest=sha256:bac1e1cdef22507ceb02536b08a230809daf90d5c18c6237e7ce878b3fac21b7

Observation fde6b7d0-6551-476f-9e37-e6c3b7aa2dd7 · inbound

Contrastive Attribution in the Wild: An Interpretability Analysis of LLM Failures on Realistic Benchmarks cites this paper.

Contrastive Attribution in the Wild: An Interpretability Analysis of LLM Failures on Realistic Benchmarks HalluLens: LLM Hallucination Benchmark

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T09:38:42.843693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T05:12:31.218055Z digest=sha256:fc5da52d540fb3f90018c0dfe00f59eac4c6670f9442349700080d238f3e9e97

Observation 58b08582-fad0-494c-bc2f-8345ec78c00c · inbound

Benchmarking Source-Sensitive Reasoning in Turkish: Humans and LLMs under Evidential Trust Manipulation cites this paper.

Benchmarking Source-Sensitive Reasoning in Turkish: Humans and LLMs under Evidential Trust Manipulation HalluLens: LLM Hallucination Benchmark

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:56:34.399829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-08T03:42:06.876417Z digest=sha256:c0b030eab862fe62d8aff6970ffbdda622c637d6f3f4b899f0781397b80641ad

Observation 6428576b-9a66-41e8-92e0-ce1a174edf42 · inbound

Hallucination as an Anomaly: Dynamic Intervention via Probabilistic Circuits cites this paper.

Hallucination as an Anomaly: Dynamic Intervention via Probabilistic Circuits HalluLens: LLM Hallucination Benchmark

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:51:11.274747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-08T10:52:14.666333Z digest=sha256:47caea7144976689ad6b81a06a103a158e1cf7f6001fbd7ece1f8af7489f8d37

Observation b96bea36-ddba-4ac7-a459-694aa94f0410 · inbound

StereoTales: A Multilingual Framework for Open-Ended Stereotype Discovery in LLMs cites this paper.

StereoTales: A Multilingual Framework for Open-Ended Stereotype Discovery in LLMs HalluLens: LLM Hallucination Benchmark

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:51:27.293122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T04:50:17.399580Z digest=sha256:c9a1d99f2f14890b800109c96723c47d14cc7f57a169ac70b455cbac84ad700a

Observation ce33085a-6e5d-4c06-9941-2928e6d6c9d7 · inbound

StereoTales: A Multilingual Framework for Open-Ended Stereotype Discovery in LLMs cites this paper.

StereoTales: A Multilingual Framework for Open-Ended Stereotype Discovery in LLMs HalluLens: LLM Hallucination Benchmark

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:22:28.981194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T07:20:32.494840Z digest=sha256:14b69ba91d39fdc4cd06d9fe938b5f5f053917c79a1ba0447d04c19c171ef541

Observation bd559f56-718a-48b3-aefb-91e6c6cd3c4c · inbound

REALISTA: Realistic Latent Adversarial Attacks that Elicit LLM Hallucinations cites this paper.

REALISTA: Realistic Latent Adversarial Attacks that Elicit LLM Hallucinations HalluLens: LLM Hallucination Benchmark

Reference 105

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:17:56.743792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-14T20:13:10.814899Z digest=sha256:28af6cbfe0d0ba065c1d0614f561c54d04bc488c93ff1852923b3980792d5fc5

Observation a6760fd2-3410-4673-83a9-401894c769b2 · inbound

Do No Harm? Hallucination and Actor-Level Abuse in Web-Deployed Medical Large Language Models cites this paper.

Do No Harm? Hallucination and Actor-Level Abuse in Web-Deployed Medical Large Language Models HalluLens: LLM Hallucination Benchmark

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-21T05:53:59.416150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-21T05:49:55.738870Z digest=sha256:cab9eba12e786935ad48913281fbe0ee96015c2670974ef5744f270781bc93b0

Observation ccf7b7c9-e403-42c9-b9d7-0bdaec1b7542 · inbound

K-FinHallu: A Hallucination Detection Benchmark for Multi-Turn RAG in Korean Finance cites this paper.

K-FinHallu: A Hallucination Detection Benchmark for Multi-Turn RAG in Korean Finance HalluLens: LLM Hallucination Benchmark

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-06-29T09:03:16.148279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-29T08:54:48.807164Z digest=sha256:696181303d33ed9b0207befcf88a355fd3d45a5fd02e5868a81bcef7930a588f

Observation b82edf81-756c-4f85-939b-097951a4d9e7 · inbound

Latent Performance Profiling of Large Language Models cites this paper.

Latent Performance Profiling of Large Language Models HalluLens: LLM Hallucination Benchmark

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-06-29T07:33:13.394631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-29T07:31:02.595386Z digest=sha256:9a67c805dc27b340d799f06ec956bf24433eabc9560204571b53722599ded726

Observation f3365005-c053-4dbb-9a77-89f8f2fcb14d · inbound

Disentangling Visual and Factual Correctness in LVLMs' Visualization Literacy cites this paper.

Disentangling Visual and Factual Correctness in LVLMs' Visualization Literacy HalluLens: LLM Hallucination Benchmark

Reference 41

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T02:06:26.198841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T11:19:51.893740Z digest=sha256:17e3bebea36170d4bf797d2c51b6354bdb912ab34ecdb229706f97bc555d1521

Observation 8026be1b-17af-41bf-9403-7f7421cd59c0 · inbound

What Do People Actually Want From AI? Mapping Preference Plurality cites this paper.

What Do People Actually Want From AI? Mapping Preference Plurality HalluLens: LLM Hallucination Benchmark

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-06-28T01:41:29.860815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T01:32:04.660400Z digest=sha256:d9e2712b3f83574f5c233401af6738d60d41fed99f324c8b8d851c7bbad59a80

Observation 00ed01a9-0708-42ef-9e46-ad803e5d813a · inbound

Contextualized Evaluation of Vision Language Models through Dynamic, Multi-turn Interactions cites this paper.

Contextualized Evaluation of Vision Language Models through Dynamic, Multi-turn Interactions HalluLens: LLM Hallucination Benchmark

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-02T01:57:59.043987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:57:59.043987Z digest=sha256:bd653a6d97d26e07330ab68169787a565f7aabe69b8618e224cd77674490435e