Pith. sign in

REVIEW 6 cited by

Calibrated Language Models Must Hallucinate

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2311.14648 v3 pith:DNDMRFLN submitted 2023-11-24 cs.CL cs.AI

classification cs.CLcs.AI
keywords factsmodelsdatahallucinationslanguagetrainingoncestatistical
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recent language models generate false but plausible-sounding text with surprising frequency. Such "hallucinations" are an obstacle to the usability of language-based AI systems and can harm people who rely upon their outputs. This work shows that there is an inherent statistical lower-bound on the rate that pretrained language models hallucinate certain types of facts, having nothing to do with the transformer LM architecture or data quality. For "arbitrary" facts whose veracity cannot be determined from the training data, we show that hallucinations must occur at a certain rate for language models that satisfy a statistical calibration condition appropriate for generative language models. Specifically, if the maximum probability of any fact is bounded, we show that the probability of generating a hallucination is close to the fraction of facts that occur exactly once in the training data (a "Good-Turing" estimate), even assuming ideal training data without errors. One conclusion is that models pretrained to be sufficiently good predictors (i.e., calibrated) may require post-training to mitigate hallucinations on the type of arbitrary facts that tend to appear once in the training set. However, our analysis also suggests that there is no statistical reason that pretraining will lead to hallucination on facts that tend to appear more than once in the training data (like references to publications such as articles and books, whose hallucinations have been particularly notable and problematic) or on systematic facts (like arithmetic calculations). Therefore, different architectures and learning algorithms may mitigate these latter types of hallucinations.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Do LLM Recommenders Know When They're Hallucinating? Auditing Confidence Calibration in Catalog Faithfulness

    cs.IR 2026-08 conditional novelty 7.0 of 10

    Verbalized confidence from four zero-shot LLM recommenders is systematically under-confident and cannot separate correct items from catalog hallucinations, so confidence-gated abstention barely reduces hallucination.

  2. Predictable Compression Failures: Order Sensitivity and Information Budgeting for Evidence-Grounded Binary Adjudication

    stat.ML 2025-09 conditional novelty 6.0 of 10

    Order-induced prediction error in binary question answering grows logarithmically with evidence length, and a pre-specified information-sufficiency gate abstains on uncertain items to hold hallucination near zero.

  3. BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment

    cs.CL 2024-11 conditional novelty 5.0 of 10

    BPO balances knowledge breadth and depth in preference data by compressing prompts and dynamically augmenting the number of response pairs per prompt using gradient-based clustering, achieving stronger alignment with ...

  4. Faster AI, Uneven Frontier: Rapid Crossings, a Jagged Frontier, and the Repositioning of Human Judgment

    cs.HC 2026-07 accept novelty 4.0 of 10

    AI has rapidly crossed expert baselines on bounded tasks on a still-jagged frontier, so human work must shift from production to specification, verification, and oversight.

  5. Inteligencia Artificial jur\'idica y el desaf\'io de la veracidad: an\'alisis de alucinaciones, optimizaci\'on de RAG y principios para una integraci\'on responsable

    cs.AI 2025-09 conditional novelty 4.0 of 10

    Legal AI hallucination persists in commercial RAG tools (17-34%+ of queries), so the report argues the fix is consultative, source-citing system design plus mandatory human oversight, not better generative models.

  6. Decoding Knowledge in Large Language Models: A Framework for Categorization and Comprehension

    cs.CL 2025-01 reject novelty 4.0 of 10

    A sampling-based framework classifies LLM knowledge into six correctness-confidence categories and applies them to measure how chain-of-thought prompting, instruction tuning, and layer depth reshape model knowledge.

Pith tools