REVIEW 3 cited by
Calibrating Verbalized Probabilities for Large Language Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Calibrating verbalized probabilities presents a novel approach for reliably assessing and leveraging outputs from black-box Large Language Models (LLMs). Recent methods have demonstrated improved calibration by applying techniques like Platt scaling or temperature scaling to the confidence scores generated by LLMs. In this paper, we explore the calibration of verbalized probability distributions for discriminative tasks. First, we investigate the capability of LLMs to generate probability distributions over categorical labels. We theoretically and empirically identify the issue of re-softmax arising from the scaling of verbalized probabilities, and propose using the invert softmax trick to approximate the "logit" by inverting verbalized probabilities. Through extensive evaluation on three public datasets, we demonstrate: (1) the robust capability of LLMs in generating class distributions, and (2) the effectiveness of the invert softmax trick in estimating logits, which, in turn, facilitates post-calibration adjustments.
Forward citations
Cited by 3 Pith papers
-
Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification
Verbalized LLM confidence scores are sparse enough that the interpolation method used to compute AUARC reverses method rankings, and a simple logprobs-weighted digit expectation (verbalization logprobs) outperforms va...
-
Aligning Language Models with Selective Prediction
RLSR aligns LLMs via a lifted AURC reward and batch ranking inside GRPO, producing better risk-coverage curves than accuracy- or calibration-based RL on in- and out-of-domain tasks.
-
Empirical Characterization of Inference-Time Elicited Probability Transformations in Large Language Models
Across 4,975 reasoning problems and multiple LLM families, post-evidence answer probabilities follow an approximate log-ratio relation log q̃ ≈ α(log q + log b) + c with mean R² ≈ 0.76, where α varies by prompting con...
Discussion (0). Continue with ORCID to comment.