Pith. sign in

REVIEW 3 cited by

Calibrating Verbalized Probabilities for Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.06707 v1 pith:RW7E2CWB submitted 2024-10-09 cs.CL cs.AI

classification cs.CLcs.AI
keywords verbalizedllmsprobabilitiesdistributionsscalingcalibratingcalibrationcapability
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Calibrating verbalized probabilities presents a novel approach for reliably assessing and leveraging outputs from black-box Large Language Models (LLMs). Recent methods have demonstrated improved calibration by applying techniques like Platt scaling or temperature scaling to the confidence scores generated by LLMs. In this paper, we explore the calibration of verbalized probability distributions for discriminative tasks. First, we investigate the capability of LLMs to generate probability distributions over categorical labels. We theoretically and empirically identify the issue of re-softmax arising from the scaling of verbalized probabilities, and propose using the invert softmax trick to approximate the "logit" by inverting verbalized probabilities. Through extensive evaluation on three public datasets, we demonstrate: (1) the robust capability of LLMs in generating class distributions, and (2) the effectiveness of the invert softmax trick in estimating logits, which, in turn, facilitates post-calibration adjustments.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification

    cs.CL 2026-08 conditional novelty 7.0 of 10

    Verbalized LLM confidence scores are sparse enough that the interpolation method used to compute AUARC reverses method rankings, and a simple logprobs-weighted digit expectation (verbalization logprobs) outperforms va...

  2. Aligning Language Models with Selective Prediction

    cs.LG 2026-07 accept novelty 7.0 of 10

    RLSR aligns LLMs via a lifted AURC reward and batch ranking inside GRPO, producing better risk-coverage curves than accuracy- or calibration-based RL on in- and out-of-domain tasks.

  3. Empirical Characterization of Inference-Time Elicited Probability Transformations in Large Language Models

    cs.CL 2026-02 conditional novelty 6.0 of 10

    Across 4,975 reasoning problems and multiple LLM families, post-evidence answer probabilities follow an approximate log-ratio relation log q̃ ≈ α(log q + log b) + c with mean R² ≈ 0.76, where α varies by prompting con...

Pith tools