Pith. sign in

Paper Citation Record · LEDGER

HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 52 inbound Pith citation observations for arXiv:2305.11747.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2305.11747 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 52 of 52 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 52 of 52 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T23:35:38.467302Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

33
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation f785cd9a-dd14-4950-938f-38ac44edc2ae · inbound

A Survey of Hallucination in Large Foundation Models cites this paper.

A Survey of Hallucination in Large Foundation Models HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 129

Resolution
verified exact
arxiv_id, observed 2026-05-16T15:21:00.927160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-16T15:21:00.778049Z digest=sha256:2cfd885b15f5ab3e97a90d692cfdd38ed291808c3ce15a440caec194cb0ae4f6

Observation 3ad50d03-cfc5-42ae-8f07-8c17fd598d2e · inbound

Ragas: Automated Evaluation of Retrieval Augmented Generation cites this paper.

Ragas: Automated Evaluation of Retrieval Augmented Generation HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T21:37:40.921250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T21:37:40.907162Z digest=sha256:95b1bc4945b444fb74e9fe2b3036ec79d938be7a78e39c4a6dab88ae0823c5c8

Observation d3b46f2b-f045-40d8-9e11-be3075ccf31c · inbound

Measuring short-form factuality in large language models cites this paper.

Measuring short-form factuality in large language models HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:45:50.262718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-15T06:45:50.219157Z digest=sha256:2380a90bdaf8e9b6c3ca75f56ee4862c9344675fe8d5fa2d3d3f48484acf2193

Observation ba0a48c1-541e-4cba-ba66-c9c0066dbb80 · inbound

LLMs to Support a Domain Specific Knowledge Assistant cites this paper.

LLMs to Support a Domain Specific Knowledge Assistant HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-08T23:35:38.467302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T23:35:38.467302Z digest=sha256:540da350673c001a14034281c2bd18ed662378d6c44f7221a93e78f174525433

Observation ba465a51-296c-4227-881a-eef3ca3c7eed · inbound

TruthFlow: Truthful LLM Generation via Representation Flow Correction cites this paper.

TruthFlow: Truthful LLM Generation via Representation Flow Correction HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-08T22:23:28.190876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:23:28.190876Z digest=sha256:cdf9da908569f8252aa326b617cc81e5d1d684bf6a8ab6ce7c051b9d7875c35a

Observation aa4d4467-2c3e-4e9d-851e-8e2b8c873dd6 · inbound

Learning Conformal Abstention Policies for Adaptive Risk Management in Large Language and Vision-Language Models cites this paper.

Learning Conformal Abstention Policies for Adaptive Risk Management in Large Language and Vision-Language Models HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-08T18:23:05.799001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:23:05.799001Z digest=sha256:026cbe4df998319acfee04834c8161a184194c2a0ceeef08b14d06715acecfff

Observation 9e01f88b-2345-4b57-9cc3-a5da29ba029f · inbound

MIH-TCCT: Mitigating Inconsistent Hallucinations in LLMs via Event-Driven Text-Code Cyclic Training cites this paper.

MIH-TCCT: Mitigating Inconsistent Hallucinations in LLMs via Event-Driven Text-Code Cyclic Training HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T23:19:36.717618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:19:36.717618Z digest=sha256:6addb7e356d90e651c6272aaf90c61422b1cbfaa7036bd2b66e8c0a03ad219a6

Observation 90b11f8a-52fb-4b25-9266-8b7ae0f47d75 · inbound

The Science of Evaluating Foundation Models cites this paper.

The Science of Evaluating Foundation Models HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.740185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.740185Z digest=sha256:0fb5357067e86716768bbb33bb8283b50405a21e56115fde92a55c1a6e4e28b3

Observation 335fa9a4-8a13-4ac9-9e57-ff735608b911 · inbound

Exposing the Ghost in the Transformer: Abnormal Detection for Large Language Models via Hidden State Forensics cites this paper.

Exposing the Ghost in the Transformer: Abnormal Detection for Large Language Models via Hidden State Forensics HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-05-22T22:32:12.959996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-22T22:27:18.533162Z digest=sha256:be049af93c37f26a399dcf827dc63b331e6837c90937df8f102be663b785f810

Observation 64283d9c-668a-4ea5-a299-eefde29d0344 · inbound

Do You Keep an Eye on What I Ask? Mitigating Multimodal Hallucination via Attention-Guided Ensemble Decoding cites this paper.

Do You Keep an Eye on What I Ask? Mitigating Multimodal Hallucination via Attention-Guided Ensemble Decoding HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T14:49:59.159168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:49:59.159168Z digest=sha256:146c5d2c12c9e6b69bb424585b2bc123059310720ec6135be4061e0411a5ee33

Observation eef52ccb-362f-41b1-8a75-1afe822787b0 · inbound

Teaching with Lies: Curriculum DPO on Synthetic Negatives for Hallucination Detection cites this paper.

Teaching with Lies: Curriculum DPO on Synthetic Negatives for Hallucination Detection HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T14:50:38.384166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:50:38.384166Z digest=sha256:4c58596e98afcb710ac848a252534bb2ee546ae301b5e7931f823d56a5d42ddd

Observation d1caff94-d1e9-41e8-b735-60124e8e6d9c · inbound

HACo-Det: A Study Towards Fine-Grained Machine-Generated Text Detection under Human-AI Coauthoring cites this paper.

HACo-Det: A Study Towards Fine-Grained Machine-Generated Text Detection under Human-AI Coauthoring HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T11:16:25.683786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:16:25.683786Z digest=sha256:368439ea6e33ab3c78ca18a562f176cd17eb8a8fd9330f7f7b07eb76606751ce

Observation c0594533-16f9-4a33-a6a3-fb39606e34e9 · inbound

Beyond Facts: Evaluating Intent Hallucination in Large Language Models cites this paper.

Beyond Facts: Evaluating Intent Hallucination in Large Language Models HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T05:59:43.915918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:59:43.915918Z digest=sha256:10734fba5e07b13e2d3c6e65f54c6368a122b5c8c528db89ce68506414562658

Observation ac6ba528-14f1-40f3-8587-189ec7a0cb6d · inbound

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs cites this paper.

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:26.745240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:26.745240Z digest=sha256:0e843dc0d3f1d05c47f5fc30e18dc41a27a067dd336abdddd3b858537170d6f5

Observation 917d0631-5793-4d33-90d9-0b2db20f746f · inbound

RealFactBench: A Benchmark for Evaluating Large Language Models in Real-World Fact-Checking cites this paper.

RealFactBench: A Benchmark for Evaluating Large Language Models in Real-World Fact-Checking HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T00:51:47.094995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:51:47.094995Z digest=sha256:e580c7a8ccc28f93db9b86f4201222d0f741d3fa1c9d450f46da59f391d2aa08

Observation 0616a75f-4787-4f89-97ea-6cddf9a4deda · inbound

Loki's Dance of Illusions: A Comprehensive Survey of Hallucination in Large Language Models cites this paper.

Loki's Dance of Illusions: A Comprehensive Survey of Hallucination in Large Language Models HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:46.515677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:46.515677Z digest=sha256:445ec16f0880571086dda006c2489af2bae194582811fbb964f2bbcf6d63c782

Observation 00711774-f145-48f1-b97b-07caf98b913c · inbound

Mitigating Geospatial Knowledge Hallucination in Large Language Models: Benchmarking and Dynamic Factuality Aligning cites this paper.

Mitigating Geospatial Knowledge Hallucination in Large Language Models: Benchmarking and Dynamic Factuality Aligning HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T14:20:39.469184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:20:39.469184Z digest=sha256:d80cd35fbf599533c6c1ecf4688ed03b386e79f53ed61f8f04cf0f5296b956c1

Observation 4d035fcd-6748-45df-b424-add5c791887f · inbound

FECT: Factuality Evaluation of Interpretive AI-Generated Claims in Contact Center Conversation Transcripts cites this paper.

FECT: Factuality Evaluation of Interpretive AI-Generated Claims in Contact Center Conversation Transcripts HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T13:56:44.067208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:56:44.067208Z digest=sha256:2132fdb1cb17e179f3d95115a9c67fc9728f12105e3bad06d274ec1ced296cd7

Observation c4296152-11e7-4b14-811f-77ec64f5c1bb · inbound

ExploreGS: Explorable 3D Scene Reconstruction with Virtual Camera Samplings and Diffusion Priors cites this paper.

ExploreGS: Explorable 3D Scene Reconstruction with Virtual Camera Samplings and Diffusion Priors HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-05T23:01:47.451336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:01:47.451336Z digest=sha256:d7b88a17bbfaf92d6fbce0d9b44db3cdca0041204c07e35c488de76d585bd4ec

Observation 6c3719b8-cc2a-418b-83d7-cf561052a6a8 · inbound

Sealing The Backdoor: Unlearning Adversarial Text Triggers In Diffusion Models Using Knowledge Distillation cites this paper.

Sealing The Backdoor: Unlearning Adversarial Text Triggers In Diffusion Models Using Knowledge Distillation HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-05T18:47:45.233361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T18:47:45.233361Z digest=sha256:a4842f3bd35b85a86c8ab1a76bee52fdf294418253e5ad6bc07b91b8d7b0c41f

Observation 18ecd53b-ea21-44a8-b9cb-5d1231cafb1b · inbound

Decoding Memories: An Efficient Pipeline for Self-Consistency Hallucination Detection cites this paper.

Decoding Memories: An Efficient Pipeline for Self-Consistency Hallucination Detection HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T14:33:18.845937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:33:18.845937Z digest=sha256:4c9cb9183ea5f6adc3927416f7180706f1e8dd75cab44c4f19fbb405b3f9032c

Observation c4a89746-0f22-4228-98c3-d07851f6bbbe · inbound

Beyond ROUGE: N-Gram Subspace Features for LLM Hallucination Detection cites this paper.

Beyond ROUGE: N-Gram Subspace Features for LLM Hallucination Detection HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-05T10:50:31.830147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:50:31.830147Z digest=sha256:cff5695e078e7e2461882b3d61e0a9ae5a1eae0c39b50d628b62af91324f445c

Observation 074076d0-24f0-47bd-84f2-ba6105a69771 · inbound

HumanAgencyBench: Scalable Evaluation of Human Agency Support in AI Assistants cites this paper.

HumanAgencyBench: Scalable Evaluation of Human Agency Support in AI Assistants HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-04T20:35:55.014746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:35:55.014746Z digest=sha256:aa29d2b997a7792091c4a1e0dc5cf3c48bb1d6365538e8d2936db0024e68b326

Observation 33d4ea83-1633-4ada-866a-7f811392355a · inbound

Investigating Symbolic Triggers of Hallucination in Gemma Models Across HaluEval and TruthfulQA cites this paper.

Investigating Symbolic Triggers of Hallucination in Gemma Models Across HaluEval and TruthfulQA HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T22:18:44.151647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:18:44.151647Z digest=sha256:e5b2be85e78a45e092e063dab4839db5b4286eaa9bfadb61a4d07de6af0393ce

Observation e57d29f1-5bdd-4c49-b3cb-1a233824cbe4 · inbound

Red-Bandit: Test-Time Adaptation for LLM Red-Teaming via Bandit-Guided LoRA Experts cites this paper.

Red-Bandit: Test-Time Adaptation for LLM Red-Teaming via Bandit-Guided LoRA Experts HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-21T20:54:21.574843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-21T20:53:58.198974Z digest=sha256:71c4a56d663368c1b033cb844d29b5fa05b88a9f5a29319aa91da73a3b14bb9b

Observation c14198a1-4884-413e-8e8d-e25c7366b5f9 · inbound

Not All Needles Are Found: How Fact Distribution and Don't Make It Up Prompts Shape Retrieval, Reasoning, and Hallucination in Long-Context LLMs cites this paper.

Not All Needles Are Found: How Fact Distribution and Don't Make It Up Prompts Shape Retrieval, Reasoning, and Hallucination in Long-Context LLMs HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-03T12:42:35.085906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:42:35.085906Z digest=sha256:58c550d6d1d2638bd296c0cbb1b684d310d0a325a192f47c7f5a08411f119a1e

Observation 131a42ec-7df5-4cf0-8d20-7d23abb9f237 · inbound

When Numbers Start Talking: Implicit Numerical Coordination Among LLM-Based Agents cites this paper.

When Numbers Start Talking: Implicit Numerical Coordination Among LLM-Based Agents HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-16T16:41:06.170142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T16:39:34.129361Z digest=sha256:50a33e8c0723120a8a31ff3353135ce3e45f763006fc776f6ce5137766969446

Observation 9707cb7f-61cd-4609-97eb-f4797891b0df · inbound

GhostCite: A Large-Scale Analysis of Citation Validity in the Age of Large Language Models cites this paper.

GhostCite: A Large-Scale Analysis of Citation Validity in the Age of Large Language Models HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 28

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T07:00:42.679785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T07:00:41.337663Z digest=sha256:eee0ed2b243b305db067f8bf8b569911ccc233fe3700d556cd632239382c639e

Observation 596840b4-65ce-4ef8-ae82-b1ae0a9a0564 · inbound

Beyond Benchmark Islands: Toward Representative Trustworthiness Evaluation for Agentic AI cites this paper.

Beyond Benchmark Islands: Toward Representative Trustworthiness Evaluation for Agentic AI HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-22T10:21:23.218723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-22T10:19:56.003219Z digest=sha256:119a312af7f587ebf03460e9ab5f3a6ccffe75a02506df673aea4f8ad2cc46fc

Observation 0057c348-a62d-46cb-8728-958a9dbd383a · inbound

Hallucination Basins: A Dynamic Framework for Understanding and Controlling LLM Hallucinations cites this paper.

Hallucination Basins: A Dynamic Framework for Understanding and Controlling LLM Hallucinations HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T23:40:52.241366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-10T18:57:00.087829Z digest=sha256:6dd5d40560e80b52844b27d09c973fdb1a1701187af882ec91deee9dd41faa31

Observation 2b179962-ea43-420d-85d1-d7be6a086cd1 · inbound

Cram Less to Fit More: Training Data Pruning Improves Memorization of Facts cites this paper.

Cram Less to Fit More: Training Data Pruning Improves Memorization of Facts HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 50

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T06:15:59.084616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-10T17:42:31.465077Z digest=sha256:8491662b2e80c897d64838e4c8e7b066082ab5ee27e4aaa12bacb4906290955a

Observation 09b24ac1-4f0f-44e7-b8a7-061460ee66e8 · inbound

Too Nice to Tell the Truth: Quantifying Agreeableness-Driven Sycophancy in Role-Playing Language Models cites this paper.

Too Nice to Tell the Truth: Quantifying Agreeableness-Driven Sycophancy in Role-Playing Language Models HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:21:00.939630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-10T16:05:09.033412Z digest=sha256:39f4696d124139d54815442da621e6075e73476cfe424d6cfb8a719d1c0b66d5

Observation 4bee8821-3fb9-423f-acb1-331238bfe827 · inbound

Self-Correcting RAG: Enhancing Faithfulness via MMKP Context Selection and NLI-Guided MCTS cites this paper.

Self-Correcting RAG: Enhancing Faithfulness via MMKP Context Selection and NLI-Guided MCTS HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T09:31:05.190349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T15:58:15.150613Z digest=sha256:2d875c9c714a6cfa8b047c133b0cbde332f327799862a9ac31ab25adf6059f73

Observation cbb8b087-7bba-4429-99cb-bdd18d0757e3 · inbound

RAGognizer: Hallucination-Aware Fine-Tuning via Detection Head Integration cites this paper.

RAGognizer: Hallucination-Aware Fine-Tuning via Detection Head Integration HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 28

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T08:43:02.040437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-10T08:38:42.029762Z digest=sha256:e2fdaaf1942c989cdfc51f716d01319d8f9bf12e79e79e36cb9dfc331011bcfd

Observation 657a417e-3833-43d4-8854-a9151f0d8e85 · inbound

HalluScan: A Systematic Benchmark for Detecting and Mitigating Hallucinations in Instruction-Following LLMs cites this paper.

HalluScan: A Systematic Benchmark for Detecting and Mitigating Hallucinations in Instruction-Following LLMs HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-25T06:50:28.037822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-25T06:49:16.755597Z digest=sha256:c77c88a59d147192588826357e6d98b5e47cb02945191cf8405d9b68d20dcd08

Observation bc60a18a-ddd8-497b-a0b9-d156d92e05d5 · inbound

CXR-ContraBench: Benchmarking Negated-Option Attraction in Medical VLMs cites this paper.

CXR-ContraBench: Benchmarking Negated-Option Attraction in Medical VLMs HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:41:09.490126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T14:49:53.357083Z digest=sha256:e77231dd381b478869dcb3d9b26736afe45b41378fe7d268c24b4d57b53c64a1

Observation e244d198-3317-4186-b063-28b3a535b664 · inbound

HalluScore: Large Language Model Hallucination Question Answering Benchmark cites this paper.

HalluScore: Large Language Model Hallucination Question Answering Benchmark HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T20:32:45.413600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T20:31:20.017866Z digest=sha256:f38b22ec38a0525003ba1bb65a4b498c44c45cb7a7fe0018fbd03c1712c2ef9f

Observation 4c646df3-dd00-4251-aaf9-ff4d5a7d5a5c · inbound

Design and Report Benchmarks for Knowledge Work cites this paper.

Design and Report Benchmarks for Knowledge Work HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 61

Resolution
metadata mismatch
arxiv_id, observed 2026-05-25T04:40:23.520503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-25T04:39:14.319133Z digest=sha256:1029ccfea3a13f343e60cfc5668e6cf77eca0585b60208e4a069b2110f7d1adf

Observation 4f0a9440-709a-4989-89ec-5a48a97d5838 · inbound

MultiHaluDet: Multilingual Hallucination Detection via LLM Hidden State Probing cites this paper.

MultiHaluDet: Multilingual Hallucination Detection via LLM Hidden State Probing HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-06-30T12:34:38.644555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T12:29:59.165791Z digest=sha256:b37c163c2b1a93730f611e2f8c4fde83479a51de11cd098e0f3f82bca7824d8d

Observation 3c55d3a7-6110-48c8-8451-e0ba570521a9 · inbound

Can LLMs Use Linguistic Uncertainty Markers to Reliably Reflect Intrinsic Confidence? cites this paper.

Can LLMs Use Linguistic Uncertainty Markers to Reliably Reflect Intrinsic Confidence? HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 49

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T12:23:24.355085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-29T12:18:36.854164Z digest=sha256:755b296488d0ba15151b8a08af6fe4d8572e387694f04b24e06095cabf8905c2

Observation 42d5f45e-2169-47e8-b63d-3bc7f987349b · inbound

K-FinHallu: A Hallucination Detection Benchmark for Multi-Turn RAG in Korean Finance cites this paper.

K-FinHallu: A Hallucination Detection Benchmark for Multi-Turn RAG in Korean Finance HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-06-29T09:03:16.135655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-29T08:54:48.807164Z digest=sha256:9255036e7aae6e536f2fce0d33e9b33fa3d2ff62c1f7915520523646db6466f7

Observation 4219f27f-bf66-4e6e-a05d-9a30eb551044 · inbound

Cascading Hallucination in Agentic RAG: The CHARM Framework for Detection and Mitigation cites this paper.

Cascading Hallucination in Agentic RAG: The CHARM Framework for Detection and Mitigation HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 38

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T07:56:47.451220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T06:33:11.701246Z digest=sha256:7f04a7be2cdd0b7e7613367406643974b7ca2c053c3e76a532fae0f502e19e23

Observation 42723d12-efeb-47cf-8361-a050b16ffdc1 · inbound

Analyzing the Correlation Between Hallucinations and Knowledge Conflicts in Large Language Models cites this paper.

Analyzing the Correlation Between Hallucinations and Knowledge Conflicts in Large Language Models HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-02T22:37:26.443815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T18:43:58.988270Z digest=sha256:8202cf92204673b3c61808a8f50d7b834121b5f489938b776bea9b58bc91d5b0

Observation fafa27c4-dae0-4fea-8701-446acae67341 · inbound

LLM-as-an-Investigator: Evidence-First Reasoning for Robust Interactive Problem Diagnosis cites this paper.

LLM-as-an-Investigator: Evidence-First Reasoning for Robust Interactive Problem Diagnosis HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-07-03T14:38:29.256488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T06:58:42.823851Z digest=sha256:14aa47db750d921acd1eb607d8fca30d906d78c16fd1b15426aba5bccddb2025

Observation 00832ebc-def9-41d3-96d5-1bcd38389bca · inbound

Reinforcement Learning with Metacognitive Feedback Elicits Faithful Uncertainty Expression in LLMs cites this paper.

Reinforcement Learning with Metacognitive Feedback Elicits Faithful Uncertainty Expression in LLMs HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 59

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T10:35:42.138233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-01T05:22:38.232552Z digest=sha256:15bb4cb332b4de54ed1e32f57eb3eec72b7d0e9dc8c37aefdbfa884db31e5d80

Observation 83edaeed-4c33-4fa0-be17-a7fd669a0976 · inbound

Confidently Wrong: Detecting Hallucinations in Financial Question Answering from LLM Internal States cites this paper.

Confidently Wrong: Detecting Hallucinations in Financial Question Answering from LLM Internal States HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-14T05:45:08.896651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T05:45:08.896651Z digest=sha256:41459fdac211997f155e455cf3f47d647e51e37d93dfc20bbde3ad193c850673

Observation da4e7f99-9b59-4835-b4b3-4818e0f03a67 · inbound

The Test Oracle Problem in Synthetic LLM-as-Judge Corpora: Disappearance, Distortion and a Validation Protocol cites this paper.

The Test Oracle Problem in Synthetic LLM-as-Judge Corpora: Disappearance, Distortion and a Validation Protocol HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-02T04:30:14.920973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:30:14.920973Z digest=sha256:f1b35fde073a946a32fbc891991a4d6510bf945cbfd3fc5ba62ffd5316070cf8

Observation fc2a8aa5-5343-4316-9cc9-3159358ed8aa · inbound

PROBE: Benchmarking Code Generation in Large Language Models cites this paper.

PROBE: Benchmarking Code Generation in Large Language Models HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-02T03:41:02.990486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T03:41:02.990486Z digest=sha256:7a29a2dfa09b580302fef2c2548c4d8021a79f2658c589160e018746e83962e2

Observation 035824ea-f668-4ebb-a588-122af89fabc2 · inbound

Reliability Scales Inversely: Hallucinations Snowball Faster in Bigger Language Models cites this paper.

Reliability Scales Inversely: Hallucinations Snowball Faster in Bigger Language Models HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-02T09:19:44.989763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T09:19:44.989763Z digest=sha256:24a11c882018ca77d466f5e3c06b7a0d2b92d83113b631ad544832765ed5a2fb

Observation 76f21532-3274-4787-8d51-d5a7ca7eb161 · inbound

Reasoning Error from Known Fact: Step-Level Self-Consistency Group Relative Policy Optimization for LLM cites this paper.

Reasoning Error from Known Fact: Step-Level Self-Consistency Group Relative Policy Optimization for LLM HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-01T14:01:26.820493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T14:01:26.820493Z digest=sha256:e4573b817dfa9c6f427459f5b1a86dfe2482fca6c17b50ef7f29b53d64dad58b

Observation ccb68686-ef83-4ab6-90c1-e65570b663bd · inbound

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment cites this paper.

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 106

Resolution
unresolved
no resolver link, observed 2026-08-01T07:12:17.577089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:12:17.577089Z digest=sha256:44a46ff62c7079ae1be13144c83caa579942177c74667372284158595276e642

Observation 5612b22d-9735-422c-954e-09e0dd4a580c · inbound

Decomposed Entailment for Factuality Checking and Hallucination Detection cites this paper.

Decomposed Entailment for Factuality Checking and Hallucination Detection HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T22:53:05.080981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T22:53:05.080981Z digest=sha256:9468ba102fa3c9c9acf8af43b4a059816beb8a7312f77b1fb198cc884bc94f5c