Pith. sign in

Paper Citation Record · LEDGER

HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 53 inbound Pith citation observations for arXiv:2305.11747.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2305.11747 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 53 of 53 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 53 of 53 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T16:40:09.360884Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

33
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation f785cd9a-dd14-4950-938f-38ac44edc2ae · inbound

A Survey of Hallucination in Large Foundation Models cites this paper.

A Survey of Hallucination in Large Foundation Models HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 129

Resolution
verified exact
arxiv_id, observed 2026-05-16T15:21:00.927160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-16T15:21:00.778049Z digest=sha256:8d17993faab73bc2f782c3a1ce932f300b4d93be42893452eb168bcf453ac104

Observation 3ad50d03-cfc5-42ae-8f07-8c17fd598d2e · inbound

Ragas: Automated Evaluation of Retrieval Augmented Generation cites this paper.

Ragas: Automated Evaluation of Retrieval Augmented Generation HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T21:37:40.921250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T21:37:40.907162Z digest=sha256:3fbbbec0e945fe30fa5d5788400f5f3fcf88c4be11dea96422d9e5a1939ceb92

Observation d3b46f2b-f045-40d8-9e11-be3075ccf31c · inbound

Measuring short-form factuality in large language models cites this paper.

Measuring short-form factuality in large language models HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:45:50.262718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-15T06:45:50.219157Z digest=sha256:2dc31a675eb6bc295df01744f416c26844a9a5408efd56fbc250198727ec1900

Observation 92a2703a-e9dc-4386-87c4-8f8c1c178919 · inbound

OnionEval: An Unified Evaluation of Fact-conflicting Hallucination for Small-Large Language Models cites this paper.

OnionEval: An Unified Evaluation of Fact-conflicting Hallucination for Small-Large Language Models HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 2000

Resolution
unresolved
no resolver link, observed 2026-08-10T16:40:09.360884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:40:09.360884Z digest=sha256:fac2943734e43293bbb70126e5e44be929ec2c20670920e5060e9e63b959bdd3

Observation ba0a48c1-541e-4cba-ba66-c9c0066dbb80 · inbound

LLMs to Support a Domain Specific Knowledge Assistant cites this paper.

LLMs to Support a Domain Specific Knowledge Assistant HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-08T23:35:38.467302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T23:35:38.467302Z digest=sha256:d70a3a700f0c1c08b3a79b7c2755f4924836db8029ff57d7ce7732d61384d93b

Observation ba465a51-296c-4227-881a-eef3ca3c7eed · inbound

TruthFlow: Truthful LLM Generation via Representation Flow Correction cites this paper.

TruthFlow: Truthful LLM Generation via Representation Flow Correction HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-08T22:23:28.190876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:23:28.190876Z digest=sha256:74fc51e89674db8b7e0c0cb65e672165d6589ec006e621b21e90a08e1f6d6945

Observation aa4d4467-2c3e-4e9d-851e-8e2b8c873dd6 · inbound

Learning Conformal Abstention Policies for Adaptive Risk Management in Large Language and Vision-Language Models cites this paper.

Learning Conformal Abstention Policies for Adaptive Risk Management in Large Language and Vision-Language Models HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-08T18:23:05.799001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:23:05.799001Z digest=sha256:1d3eb999bf1581d0ef0ba4ff9b08857f36769c6007f026734ad43fdc680ccd05

Observation 9e01f88b-2345-4b57-9cc3-a5da29ba029f · inbound

MIH-TCCT: Mitigating Inconsistent Hallucinations in LLMs via Event-Driven Text-Code Cyclic Training cites this paper.

MIH-TCCT: Mitigating Inconsistent Hallucinations in LLMs via Event-Driven Text-Code Cyclic Training HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T23:19:36.717618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:19:36.717618Z digest=sha256:a2ac646fd664ea30d7b7419dd28b8d4db08ed1ac49ea0baf9a24282c4d25ea28

Observation 90b11f8a-52fb-4b25-9266-8b7ae0f47d75 · inbound

The Science of Evaluating Foundation Models cites this paper.

The Science of Evaluating Foundation Models HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.740185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.740185Z digest=sha256:1bd2218739cdae5c8bf68fb90f8266162636cb8f2e0da9dbb0f088eda01a8894

Observation 335fa9a4-8a13-4ac9-9e57-ff735608b911 · inbound

Exposing the Ghost in the Transformer: Abnormal Detection for Large Language Models via Hidden State Forensics cites this paper.

Exposing the Ghost in the Transformer: Abnormal Detection for Large Language Models via Hidden State Forensics HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-05-22T22:32:12.959996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T22:27:18.533162Z digest=sha256:cf5ea0c6a0a653326399dd0e88a6a663f0bf88907671476c94afb493ed5fd827

Observation 64283d9c-668a-4ea5-a299-eefde29d0344 · inbound

Do You Keep an Eye on What I Ask? Mitigating Multimodal Hallucination via Attention-Guided Ensemble Decoding cites this paper.

Do You Keep an Eye on What I Ask? Mitigating Multimodal Hallucination via Attention-Guided Ensemble Decoding HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T14:49:59.159168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:49:59.159168Z digest=sha256:d083e2c741e0f026a61946eafaf067d67efb0c249cafed5ee2c2a5870b307c6b

Observation eef52ccb-362f-41b1-8a75-1afe822787b0 · inbound

Teaching with Lies: Curriculum DPO on Synthetic Negatives for Hallucination Detection cites this paper.

Teaching with Lies: Curriculum DPO on Synthetic Negatives for Hallucination Detection HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T14:50:38.384166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:50:38.384166Z digest=sha256:26279c2c14ede0ee3958ad4aa750abd4794f73eb9e8e6e3a89932cf309849da5

Observation d1caff94-d1e9-41e8-b735-60124e8e6d9c · inbound

HACo-Det: A Study Towards Fine-Grained Machine-Generated Text Detection under Human-AI Coauthoring cites this paper.

HACo-Det: A Study Towards Fine-Grained Machine-Generated Text Detection under Human-AI Coauthoring HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T11:16:25.683786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:16:25.683786Z digest=sha256:330bbf300d98a9456ad6e2f657ff6adf7cd7c9ee426b2eec9dbe64bba6cdc260

Observation c0594533-16f9-4a33-a6a3-fb39606e34e9 · inbound

Beyond Facts: Evaluating Intent Hallucination in Large Language Models cites this paper.

Beyond Facts: Evaluating Intent Hallucination in Large Language Models HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T05:59:43.915918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:59:43.915918Z digest=sha256:0ac42031d8c70bd97ce0094e5f9c907d3a4e3f091e5ebba8ca6f347614260ed3

Observation ac6ba528-14f1-40f3-8587-189ec7a0cb6d · inbound

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs cites this paper.

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:26.745240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:26.745240Z digest=sha256:32ae2d7e149247607a37eae7622d86d0811340ea418b09d48021870ee459452c

Observation 917d0631-5793-4d33-90d9-0b2db20f746f · inbound

RealFactBench: A Benchmark for Evaluating Large Language Models in Real-World Fact-Checking cites this paper.

RealFactBench: A Benchmark for Evaluating Large Language Models in Real-World Fact-Checking HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T00:51:47.094995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:51:47.094995Z digest=sha256:e83cd5bb763ef7bb1cac8efd966564332dd8245e8bdabef1c92df9748ccc415e

Observation 0616a75f-4787-4f89-97ea-6cddf9a4deda · inbound

Loki's Dance of Illusions: A Comprehensive Survey of Hallucination in Large Language Models cites this paper.

Loki's Dance of Illusions: A Comprehensive Survey of Hallucination in Large Language Models HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:46.515677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:46.515677Z digest=sha256:dea86463a2ebe62dd67083ad9b802d86032e8a6153daa4aa9787d8ddd53c0323

Observation 00711774-f145-48f1-b97b-07caf98b913c · inbound

Mitigating Geospatial Knowledge Hallucination in Large Language Models: Benchmarking and Dynamic Factuality Aligning cites this paper.

Mitigating Geospatial Knowledge Hallucination in Large Language Models: Benchmarking and Dynamic Factuality Aligning HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T14:20:39.469184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:20:39.469184Z digest=sha256:e5be6dd782f69bf21a8324756a28830c7f6fd43df20171f61fca69eaaec4e922

Observation 4d035fcd-6748-45df-b424-add5c791887f · inbound

FECT: Factuality Evaluation of Interpretive AI-Generated Claims in Contact Center Conversation Transcripts cites this paper.

FECT: Factuality Evaluation of Interpretive AI-Generated Claims in Contact Center Conversation Transcripts HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T13:56:44.067208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:56:44.067208Z digest=sha256:3b6c7c850af21e0ed2f5066c88e23f945cd6a908d6d910414bdebd01d9089f39

Observation c4296152-11e7-4b14-811f-77ec64f5c1bb · inbound

ExploreGS: Explorable 3D Scene Reconstruction with Virtual Camera Samplings and Diffusion Priors cites this paper.

ExploreGS: Explorable 3D Scene Reconstruction with Virtual Camera Samplings and Diffusion Priors HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-05T23:01:47.451336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:01:47.451336Z digest=sha256:e3660ab65c482bd8ac8dc1f5cbea13f72b83cd5b5f75ceac6981bf936867b532

Observation 6c3719b8-cc2a-418b-83d7-cf561052a6a8 · inbound

Sealing The Backdoor: Unlearning Adversarial Text Triggers In Diffusion Models Using Knowledge Distillation cites this paper.

Sealing The Backdoor: Unlearning Adversarial Text Triggers In Diffusion Models Using Knowledge Distillation HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-05T18:47:45.233361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T18:47:45.233361Z digest=sha256:f7b5083d9da613264ecf02ef1825dcefd4fd59fe94353b32f1c8bc0c1ab07f17

Observation 18ecd53b-ea21-44a8-b9cb-5d1231cafb1b · inbound

Decoding Memories: An Efficient Pipeline for Self-Consistency Hallucination Detection cites this paper.

Decoding Memories: An Efficient Pipeline for Self-Consistency Hallucination Detection HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T14:33:18.845937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:33:18.845937Z digest=sha256:2b7e39db9b5dc33fa8a18ba546c1da12cd02b1a344329eea356be1b7fb6a49bc

Observation c4a89746-0f22-4228-98c3-d07851f6bbbe · inbound

Beyond ROUGE: N-Gram Subspace Features for LLM Hallucination Detection cites this paper.

Beyond ROUGE: N-Gram Subspace Features for LLM Hallucination Detection HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-05T10:50:31.830147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:50:31.830147Z digest=sha256:674f01da27377392bf5285ba4515f82273ab05756eff91080426abbcd1642c6e

Observation 074076d0-24f0-47bd-84f2-ba6105a69771 · inbound

HumanAgencyBench: Scalable Evaluation of Human Agency Support in AI Assistants cites this paper.

HumanAgencyBench: Scalable Evaluation of Human Agency Support in AI Assistants HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-04T20:35:55.014746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:35:55.014746Z digest=sha256:f41095470f39feb34c062979e5aaaa87f1964876724d9b383dbefe85d60e7fb3

Observation 33d4ea83-1633-4ada-866a-7f811392355a · inbound

Investigating Symbolic Triggers of Hallucination in Gemma Models Across HaluEval and TruthfulQA cites this paper.

Investigating Symbolic Triggers of Hallucination in Gemma Models Across HaluEval and TruthfulQA HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T22:18:44.151647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:18:44.151647Z digest=sha256:1b21f977eb4c9071e1116681b9f178a00c97c5382ec1aa5c52f0d3481d9637fa

Observation e57d29f1-5bdd-4c49-b3cb-1a233824cbe4 · inbound

Red-Bandit: Test-Time Adaptation for LLM Red-Teaming via Bandit-Guided LoRA Experts cites this paper.

Red-Bandit: Test-Time Adaptation for LLM Red-Teaming via Bandit-Guided LoRA Experts HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-21T20:54:21.574843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T20:53:58.198974Z digest=sha256:833da330ace309e69a68d42fa9ceadd9ea1cdb95aaf65dc0b6fe146cc2c0bb96

Observation c14198a1-4884-413e-8e8d-e25c7366b5f9 · inbound

Not All Needles Are Found: How Fact Distribution and Don't Make It Up Prompts Shape Retrieval, Reasoning, and Hallucination in Long-Context LLMs cites this paper.

Not All Needles Are Found: How Fact Distribution and Don't Make It Up Prompts Shape Retrieval, Reasoning, and Hallucination in Long-Context LLMs HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-03T12:42:35.085906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:42:35.085906Z digest=sha256:de4ababe1fcfd7f58b68ff62e166a5d4f8342e18da6ad505503423a747dc424d

Observation 131a42ec-7df5-4cf0-8d20-7d23abb9f237 · inbound

When Numbers Start Talking: Implicit Numerical Coordination Among LLM-Based Agents cites this paper.

When Numbers Start Talking: Implicit Numerical Coordination Among LLM-Based Agents HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-16T16:41:06.170142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T16:39:34.129361Z digest=sha256:474bd1edb62e7d642734681f228ceaa7202cac4520d678f3b7fd407ade60462a

Observation 9707cb7f-61cd-4609-97eb-f4797891b0df · inbound

GhostCite: A Large-Scale Analysis of Citation Validity in the Age of Large Language Models cites this paper.

GhostCite: A Large-Scale Analysis of Citation Validity in the Age of Large Language Models HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 28

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T07:00:42.679785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T07:00:41.337663Z digest=sha256:c5e3ef72ff82dbf8229c253b3559f410488690801cdb0f3366040811ed50dd9b

Observation 596840b4-65ce-4ef8-ae82-b1ae0a9a0564 · inbound

Beyond Benchmark Islands: Toward Representative Trustworthiness Evaluation for Agentic AI cites this paper.

Beyond Benchmark Islands: Toward Representative Trustworthiness Evaluation for Agentic AI HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-22T10:21:23.218723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T10:19:56.003219Z digest=sha256:ab41b910652c22923f4f0f47089a9e8008f5789caac2ddef4b2b7debea0adcec

Observation 0057c348-a62d-46cb-8728-958a9dbd383a · inbound

Hallucination Basins: A Dynamic Framework for Understanding and Controlling LLM Hallucinations cites this paper.

Hallucination Basins: A Dynamic Framework for Understanding and Controlling LLM Hallucinations HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T23:40:52.241366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T18:57:00.087829Z digest=sha256:ddfe9c6380e80711ab8b452f3bdeb687d8aef81564a97f3c2bc821558ccc6e33

Observation 2b179962-ea43-420d-85d1-d7be6a086cd1 · inbound

Cram Less to Fit More: Training Data Pruning Improves Memorization of Facts cites this paper.

Cram Less to Fit More: Training Data Pruning Improves Memorization of Facts HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 50

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T06:15:59.084616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T17:42:31.465077Z digest=sha256:d461af42a28bb41ed64c5e2bd3123424a6c6cabaea2ff5689d31625b742934fc

Observation 09b24ac1-4f0f-44e7-b8a7-061460ee66e8 · inbound

Too Nice to Tell the Truth: Quantifying Agreeableness-Driven Sycophancy in Role-Playing Language Models cites this paper.

Too Nice to Tell the Truth: Quantifying Agreeableness-Driven Sycophancy in Role-Playing Language Models HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:21:00.939630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T16:05:09.033412Z digest=sha256:d45074698df575daa635adab127b6df289747c93ed27c180c30a62b6cfaaa016

Observation 4bee8821-3fb9-423f-acb1-331238bfe827 · inbound

Self-Correcting RAG: Enhancing Faithfulness via MMKP Context Selection and NLI-Guided MCTS cites this paper.

Self-Correcting RAG: Enhancing Faithfulness via MMKP Context Selection and NLI-Guided MCTS HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T09:31:05.190349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T15:58:15.150613Z digest=sha256:bec42f2f7f4fe30663789399e1aae4f3b19a2c75112ee632686d6a809901ecd0

Observation cbb8b087-7bba-4429-99cb-bdd18d0757e3 · inbound

RAGognizer: Hallucination-Aware Fine-Tuning via Detection Head Integration cites this paper.

RAGognizer: Hallucination-Aware Fine-Tuning via Detection Head Integration HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 28

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T08:43:02.040437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T08:38:42.029762Z digest=sha256:92ae37dc99f15f05b8f0c8659b1d716739d54885d4109aea4e4bdcdc0387847c

Observation 657a417e-3833-43d4-8854-a9151f0d8e85 · inbound

HalluScan: A Systematic Benchmark for Detecting and Mitigating Hallucinations in Instruction-Following LLMs cites this paper.

HalluScan: A Systematic Benchmark for Detecting and Mitigating Hallucinations in Instruction-Following LLMs HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-25T06:50:28.037822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-25T06:49:16.755597Z digest=sha256:ce89e8dce5f5b95bbe37630c1f3bfbebe7f91b88dfa52d88d2960409de367b9e

Observation bc60a18a-ddd8-497b-a0b9-d156d92e05d5 · inbound

CXR-ContraBench: Benchmarking Negated-Option Attraction in Medical VLMs cites this paper.

CXR-ContraBench: Benchmarking Negated-Option Attraction in Medical VLMs HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:41:09.490126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T14:49:53.357083Z digest=sha256:6686b44f2e0788ad7e933169570024ea3e6fbe5c2809575061691b35088f1f8d

Observation e244d198-3317-4186-b063-28b3a535b664 · inbound

HalluScore: Large Language Model Hallucination Question Answering Benchmark cites this paper.

HalluScore: Large Language Model Hallucination Question Answering Benchmark HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T20:32:45.413600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T20:31:20.017866Z digest=sha256:c4f993ee0f40c354fb7d215f91d24087f6a0753f705ed7400f57b2d96b98f68f

Observation 4c646df3-dd00-4251-aaf9-ff4d5a7d5a5c · inbound

Design and Report Benchmarks for Knowledge Work cites this paper.

Design and Report Benchmarks for Knowledge Work HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 61

Resolution
metadata mismatch
arxiv_id, observed 2026-05-25T04:40:23.520503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-25T04:39:14.319133Z digest=sha256:99d3d4df148e8b5ffa64e4e40ce8a799503b1e3b505a40b03ea2aa20445fe4db

Observation 4f0a9440-709a-4989-89ec-5a48a97d5838 · inbound

MultiHaluDet: Multilingual Hallucination Detection via LLM Hidden State Probing cites this paper.

MultiHaluDet: Multilingual Hallucination Detection via LLM Hidden State Probing HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-06-30T12:34:38.644555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T12:29:59.165791Z digest=sha256:c1e85dc09fcda22b88c0994ce692b3c421fe7cbf0a128e94a48e6ad4400d6b8c

Observation 3c55d3a7-6110-48c8-8451-e0ba570521a9 · inbound

Can LLMs Use Linguistic Uncertainty Markers to Reliably Reflect Intrinsic Confidence? cites this paper.

Can LLMs Use Linguistic Uncertainty Markers to Reliably Reflect Intrinsic Confidence? HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 49

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T12:23:24.355085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T12:18:36.854164Z digest=sha256:c946c0aef13d96b4d3aa13c3e05e65004d96113f3f7cef381a00c84d020e29be

Observation 42d5f45e-2169-47e8-b63d-3bc7f987349b · inbound

K-FinHallu: A Hallucination Detection Benchmark for Multi-Turn RAG in Korean Finance cites this paper.

K-FinHallu: A Hallucination Detection Benchmark for Multi-Turn RAG in Korean Finance HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-06-29T09:03:16.135655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-29T08:54:48.807164Z digest=sha256:575ca8c114741c4009d81a85ee28112bf2e0841996bf84e624994bdf92e2608e

Observation 4219f27f-bf66-4e6e-a05d-9a30eb551044 · inbound

Cascading Hallucination in Agentic RAG: The CHARM Framework for Detection and Mitigation cites this paper.

Cascading Hallucination in Agentic RAG: The CHARM Framework for Detection and Mitigation HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 38

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T07:56:47.451220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T06:33:11.701246Z digest=sha256:a30d248746dcda1f68e9b0d09ee39ebe099ca4f8eae24a7106bbc98b56aac552

Observation 42723d12-efeb-47cf-8361-a050b16ffdc1 · inbound

Analyzing the Correlation Between Hallucinations and Knowledge Conflicts in Large Language Models cites this paper.

Analyzing the Correlation Between Hallucinations and Knowledge Conflicts in Large Language Models HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-02T22:37:26.443815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T18:43:58.988270Z digest=sha256:67302b969c5bf9c14e213d1c9d51278030bfc41937d38230bfe0043501fa95d4

Observation fafa27c4-dae0-4fea-8701-446acae67341 · inbound

LLM-as-an-Investigator: Evidence-First Reasoning for Robust Interactive Problem Diagnosis cites this paper.

LLM-as-an-Investigator: Evidence-First Reasoning for Robust Interactive Problem Diagnosis HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-07-03T14:38:29.256488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T06:58:42.823851Z digest=sha256:d694578aa3ed1f729715cb8ba8e393e40b81d2a528f6ba3a4651dfa9eaa9954a

Observation 00832ebc-def9-41d3-96d5-1bcd38389bca · inbound

Reinforcement Learning with Metacognitive Feedback Elicits Faithful Uncertainty Expression in LLMs cites this paper.

Reinforcement Learning with Metacognitive Feedback Elicits Faithful Uncertainty Expression in LLMs HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 59

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T10:35:42.138233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-01T05:22:38.232552Z digest=sha256:e542f6fd522ee1ab064e0bdee743da7d2ca93ca61e89b83008a28071e219f3d1

Observation 83edaeed-4c33-4fa0-be17-a7fd669a0976 · inbound

Confidently Wrong: Detecting Hallucinations in Financial Question Answering from LLM Internal States cites this paper.

Confidently Wrong: Detecting Hallucinations in Financial Question Answering from LLM Internal States HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-14T05:45:08.896651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T05:45:08.896651Z digest=sha256:dfba3e0c81bd0f38f558de633f830d1c0c96a8f20902b1bcbe40cf4561360e6b

Observation da4e7f99-9b59-4835-b4b3-4818e0f03a67 · inbound

The Test Oracle Problem in Synthetic LLM-as-Judge Corpora: Disappearance, Distortion and a Validation Protocol cites this paper.

The Test Oracle Problem in Synthetic LLM-as-Judge Corpora: Disappearance, Distortion and a Validation Protocol HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-02T04:30:14.920973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:30:14.920973Z digest=sha256:b22a2425c43aed7cbb8466a415ba036d8ff1d61670121b8da2c003b9afa12b20

Observation fc2a8aa5-5343-4316-9cc9-3159358ed8aa · inbound

PROBE: Benchmarking Code Generation in Large Language Models cites this paper.

PROBE: Benchmarking Code Generation in Large Language Models HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-02T03:41:02.990486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T03:41:02.990486Z digest=sha256:83a807b35183a0b3639a0231b73bcc65f3f89471df43fd6d51e5f28c6d89ea67

Observation 035824ea-f668-4ebb-a588-122af89fabc2 · inbound

Reliability Scales Inversely: Hallucinations Snowball Faster in Bigger Language Models cites this paper.

Reliability Scales Inversely: Hallucinations Snowball Faster in Bigger Language Models HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-02T09:19:44.989763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T09:19:44.989763Z digest=sha256:af83d82bfe0e9e8063c54351e1c57e037b7050482e617e0c8290c28cae545111

Observation 76f21532-3274-4787-8d51-d5a7ca7eb161 · inbound

Reasoning Error from Known Fact: Step-Level Self-Consistency Group Relative Policy Optimization for LLM cites this paper.

Reasoning Error from Known Fact: Step-Level Self-Consistency Group Relative Policy Optimization for LLM HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-01T14:01:26.820493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T14:01:26.820493Z digest=sha256:9027082e555448093178c742c1015eb3b8569daa8c2c3bf338a592ffef8508b0

Observation ccb68686-ef83-4ab6-90c1-e65570b663bd · inbound

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment cites this paper.

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 106

Resolution
unresolved
no resolver link, observed 2026-08-01T07:12:17.577089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:12:17.577089Z digest=sha256:78a4cad95ddf7cf4e0a119003b512acdb031dbce6babc488793851ab859825fc

Observation 5612b22d-9735-422c-954e-09e0dd4a580c · inbound

Decomposed Entailment for Factuality Checking and Hallucination Detection cites this paper.

Decomposed Entailment for Factuality Checking and Hallucination Detection HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T22:53:05.080981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T22:53:05.080981Z digest=sha256:2cd342dc088741d1a014cae0a4dd3730e04e0aa634f5caaa86c282defe70cb76