Pith. sign in

REVIEW 6 cited by

Quantifying Uncertainty in Answers from any Language Model and Enhancing their Trustworthiness

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2308.16175 v2 pith:RQIUB36D submitted 2023-08-30 cs.CL cs.AI

classification cs.CLcs.AI
keywords confidenceresponsesuncertaintyanswersbsdetectorevaluationextralanguage
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We introduce BSDetector, a method for detecting bad and speculative answers from a pretrained Large Language Model by estimating a numeric confidence score for any output it generated. Our uncertainty quantification technique works for any LLM accessible only via a black-box API, whose training data remains unknown. By expending a bit of extra computation, users of any LLM API can now get the same response as they would ordinarily, as well as a confidence estimate that cautions when not to trust this response. Experiments on both closed and open-form Question-Answer benchmarks reveal that BSDetector more accurately identifies incorrect LLM responses than alternative uncertainty estimation procedures (for both GPT-3 and ChatGPT). By sampling multiple responses from the LLM and considering the one with the highest confidence score, we can additionally obtain more accurate responses from the same LLM, without any extra training steps. In applications involving automated evaluation with LLMs, accounting for our confidence scores leads to more reliable evaluation in both human-in-the-loop and fully-automated settings (across both GPT 3.5 and 4).

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 7 citations worldwide. Full citation record

  1. HalluField: Detecting LLM Hallucinations via Field-Theoretic Modeling

    cs.LG 2025-09 conditional novelty 6.0 of 10

    HalluField flags LLM hallucinations using a hand-weighted temperature-perturbation of token-level 'free energy' (negative log-likelihood) and Shannon entropy.

  2. Conformal Language Model Reasoning with Coherent Factuality

    cs.CL 2025-05 conditional novelty 6.0 of 10

    A conformal prediction filter over deducibility graphs keeps language model reasoning coherently factual at user-chosen coverage levels on MATH and FELM problems.

  3. Entropy Sentinel: Probing Entropy Traces for LLM Monitoring

    cs.CL 2026-01 conditional novelty 5.0 of 10

    Top-k decoding-entropy profiles can estimate and rank held-out domain accuracy for most tested LLMs, with difficulty-diverse training data the main success factor.

  4. The Consistency Hypothesis in Uncertainty Quantification for Large Language Models

    cs.CL 2025-06 conditional novelty 5.0 of 10

    LLM generations that resemble other generations for the same question are more likely to be correct, and this rule can be used to build competitive black-box confidence scores.

  5. From Calibration to Collaboration: LLM Uncertainty Quantification Should Be More Human-Centered

    cs.CL 2025-06 conditional novelty 5.0 of 10

    LLM uncertainty quantification should be judged by whether it improves real human decisions, not by calibration scores on trivia benchmarks.

  6. ChemAU: Harness the Reasoning of LLMs in Chemical Research with Adaptive Uncertainty Estimation

    cs.AI 2025-06 reject novelty 4.0 of 10

    ChemAU adds a position penalty to token-level uncertainty estimates so that flagged reasoning steps are corrected by a fine-tuned chemistry model, reporting improved accuracy on GPQA, MMLU-Pro, and SuperGPQA chemistry...

Pith tools