Pith. sign in

REVIEW 4 cited by

HALoGEN: Fantastic LLM Hallucinations and Where to Find Them

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2501.08292 v1 pith:4VUTL42T submitted 2025-01-14 cs.CL cs.AI

classification cs.CLcs.AI
keywords modelshallucinationserrorsgenerationsgenerativeknowledgelanguagetype
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Despite their impressive ability to generate high-quality and fluent text, generative large language models (LLMs) also produce hallucinations: statements that are misaligned with established world knowledge or provided input context. However, measuring hallucination can be challenging, as having humans verify model generations on-the-fly is both expensive and time-consuming. In this work, we release HALoGEN, a comprehensive hallucination benchmark consisting of: (1) 10,923 prompts for generative models spanning nine domains including programming, scientific attribution, and summarization, and (2) automatic high-precision verifiers for each use case that decompose LLM generations into atomic units, and verify each unit against a high-quality knowledge source. We use this framework to evaluate ~150,000 generations from 14 language models, finding that even the best-performing models are riddled with hallucinations (sometimes up to 86% of generated atomic facts depending on the domain). We further define a novel error classification for LLM hallucinations based on whether they likely stem from incorrect recollection of training data (Type A errors), or incorrect knowledge in training data (Type B errors), or are fabrication (Type C errors). We hope our framework provides a foundation to enable the principled study of why generative models hallucinate, and advances the development of trustworthy large language models.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. HALLMARK: Diagnosing Three Failure Modes in LLM Citation Verifiers

    cs.CR 2026-07 conditional novelty 6.0 of 10

    HALLMARK shows that citation verifiers' false-positive rate, not recall, is what determines whether their flags are mostly real catches or mostly noise at realistic hallucination rates.

  2. Auditable Context-Aware HFMD Forecasting with Structured LLM Agents

    cs.LG 2025-11 conditional novelty 6.0 of 10

    An LLM-based two-agent system can forecast hand-foot-mouth disease cases with accuracy comparable to top numerical models while generating human-readable risk explanations.

  3. MIRAGE-Bench: LLM Agent is Hallucinating and Where to Find Them

    cs.AI 2025-07 conditional novelty 6.0 of 10

    A new benchmark with a three-way hallucination taxonomy, snapshot-based test cases, and an LLM judge shows LLM agents hallucinate at over 30% of risky decision points, with open and closed models closer than expected.

  4. A Survey on LLM-based News Recommender Systems

    cs.IR 2025-02 conditional novelty 5.0 of 10

    A survey that categorizes LLM-based news recommender systems and reports benchmark comparisons on MIND and Adressa.

Pith tools