Pith. sign in

REVIEW 5 cited by

Fact-Checking the Output of Large Language Models via Token-Level Uncertainty Quantification

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.04696 v2 pith:LEKBR5LZ submitted 2024-03-07 cs.CL cs.AIcs.LG

classification cs.CLcs.AIcs.LG
keywords uncertaintyoutputfact-checkingquantificationclaimllmstoken-levelclaims
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large language models (LLMs) are notorious for hallucinating, i.e., producing erroneous claims in their output. Such hallucinations can be dangerous, as occasional factual inaccuracies in the generated text might be obscured by the rest of the output being generally factually correct, making it extremely hard for the users to spot them. Current services that leverage LLMs usually do not provide any means for detecting unreliable generations. Here, we aim to bridge this gap. In particular, we propose a novel fact-checking and hallucination detection pipeline based on token-level uncertainty quantification. Uncertainty scores leverage information encapsulated in the output of a neural network or its layers to detect unreliable predictions, and we show that they can be used to fact-check the atomic claims in the LLM output. Moreover, we present a novel token-level uncertainty quantification method that removes the impact of uncertainty about what claim to generate on the current step and what surface form to use. Our method Claim Conditioned Probability (CCP) measures only the uncertainty of a particular claim value expressed by the model. Experiments on the task of biography generation demonstrate strong improvements for CCP compared to the baselines for seven LLMs and four languages. Human evaluation reveals that the fact-checking pipeline based on uncertainty quantification is competitive with a fact-checking tool that leverages external knowledge.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. VecFontLLM: Anchor-Guided Direct Synthesis of Chinese Vector Fonts

    cs.CV 2026-07 conditional novelty 7.0 of 10

    An anchor-guided MLLM generates complex Chinese vector glyphs from one or a few style exemplars, decoupling coarse layout from Bézier curve completion to improve structure and editability.

  2. Will It Still Be True Tomorrow? Multilingual Evergreen Question Classification to Improve Trustworthy QA

    cs.CL 2025-05 conditional novelty 7.0 of 10

    EverGreenQA and EG-E5 provide a multilingual, human-labeled evergreen question classifier that improves self-knowledge estimation and QA dataset curation.

  3. Reasoning about Uncertainty: Do Reasoning Models Know When They Don't Know?

    cs.AI 2025-06 conditional novelty 6.0 of 10

    Reasoning language models are systematically overconfident, deeper reasoning makes them more overconfident, and a two-stage introspective prompting method improves calibration for some models.

  4. Entropy Sentinel: Probing Entropy Traces for LLM Monitoring

    cs.CL 2026-01 conditional novelty 5.0 of 10

    Top-k decoding-entropy profiles can estimate and rank held-out domain accuracy for most tested LLMs, with difficulty-diverse training data the main success factor.

  5. ChemAU: Harness the Reasoning of LLMs in Chemical Research with Adaptive Uncertainty Estimation

    cs.AI 2025-06 reject novelty 4.0 of 10

    ChemAU adds a position penalty to token-level uncertainty estimates so that flagged reasoning steps are corrected by a fine-tuned chemistry model, reporting improved accuracy on GPQA, MMLU-Pro, and SuperGPQA chemistry...

Pith tools