Pith. sign in

REVIEW 2 cited by

MAQA: Evaluating Uncertainty Quantification in LLMs Regarding Data Uncertainty

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2408.06816 v2 pith:2YZVGOPX submitted 2024-08-13 cs.AI cs.CL

classification cs.AIcs.CL
keywords uncertaintyquantificationdatamethodsllmssettingsmaqapresence
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Despite the massive advancements in large language models (LLMs), they still suffer from producing plausible but incorrect responses. To improve the reliability of LLMs, recent research has focused on uncertainty quantification to predict whether a response is correct or not. However, most uncertainty quantification methods have been evaluated on single-labeled questions, which removes data uncertainty: the irreducible randomness often present in user queries, which can arise from factors like multiple possible answers. This limitation may cause uncertainty quantification results to be unreliable in practical settings. In this paper, we investigate previous uncertainty quantification methods under the presence of data uncertainty. Our contributions are two-fold: 1) proposing a new Multi-Answer Question Answering dataset, MAQA, consisting of world knowledge, mathematical reasoning, and commonsense reasoning tasks to evaluate uncertainty quantification regarding data uncertainty, and 2) assessing 5 uncertainty quantification methods of diverse white- and black-box LLMs. Our findings show that previous methods relatively struggle compared to single-answer settings, though this varies depending on the task. Moreover, we observe that entropy- and consistency-based methods effectively estimate model uncertainty, even in the presence of data uncertainty. We believe these observations will guide future work on uncertainty quantification in more realistic settings.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MedBayes-Lite: A Clinical Uncertainty Governance Layer for Risk-Aware Medical Decision Support

    cs.AI 2025-11 reject novelty 3.0 of 10

    A retraining-free uncertainty layer is claimed to reduce overconfident clinical QA errors, but the key derivation is invalid and the abstract and full text report different results.

  2. Assessing GPT Model Uncertainty in Mathematical OCR Tasks via Entropy Analysis

    cs.IT 2024-12 reject novelty 3.0 of 10

    The paper reports that GPT-4o's token-level uncertainty, computed as the negative log-likelihood of its output, rises monotonically as image resolution falls from 300 to 72 dpi on a single test page.

Pith tools