Pith. sign in

REVIEW 4 cited by

Finetuning Language Models to Emit Linguistic Expressions of Uncertainty

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2409.12180 v1 pith:XF2LF7Q6 submitted 2024-09-18 cs.CL cs.LG

classification cs.CLcs.LG
keywords modelsexpressionsllmsuncertaintyfinetuninglanguagelinguisticpredictions
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large language models (LLMs) are increasingly employed in information-seeking and decision-making tasks. Despite their broad utility, LLMs tend to generate information that conflicts with real-world facts, and their persuasive style can make these inaccuracies appear confident and convincing. As a result, end-users struggle to consistently align the confidence expressed by LLMs with the accuracy of their predictions, often leading to either blind trust in all outputs or a complete disregard for their reliability. In this work, we explore supervised finetuning on uncertainty-augmented predictions as a method to develop models that produce linguistic expressions of uncertainty. Specifically, we measure the calibration of pre-trained models and then fine-tune language models to generate calibrated linguistic expressions of uncertainty. Through experiments on various question-answering datasets, we demonstrate that LLMs are well-calibrated in assessing their predictions, and supervised finetuning based on the model's own confidence leads to well-calibrated expressions of uncertainty, particularly for single-claim answers.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Confidence Elicitation: A New Attack Vector for Large Language Models

    cs.LG 2025-02 conditional novelty 6.0 of 10

    Confidence elicitation, asking a model to verbalize its uncertainty, can be used as a soft-label signal to craft stronger black-box word-substitution attacks on LLMs.

  2. Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety

    cs.SE 2025-06 accept novelty 5.0 of 10

    A new survey organizes LLM interpretation methods by workflow stage and connects them to safety enhancement strategies and tools, covering around 70 works.

  3. Knowledge Boundary of Large Language Models: A Survey

    cs.CL 2024-12 conditional novelty 5.0 of 10

    A survey that formalizes the knowledge boundary of LLMs into a four-type taxonomy and reviews detection and mitigation methods.

  4. Towards Harmonized Uncertainty Estimation for Large Language Models

    cs.CL 2025-05 conditional novelty 4.0 of 10

    CUE combines a supervised correctness classifier with existing LLM uncertainty scores to improve indication, balance, and calibration, reporting AUROC and ECE gains across models and datasets.

Pith tools