Pith. sign in

REVIEW 1 cited by

Improving Predictor Reliability with Selective Recalibration

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.05407 v1 pith:6O36NSLV submitted 2024-10-07 cs.LG cs.AI

classification cs.LGcs.AI
keywords recalibrationmodelselectivespacecalibrationconfidencedatadifficult
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

A reliable deep learning system should be able to accurately express its confidence with respect to its predictions, a quality known as calibration. One of the most effective ways to produce reliable confidence estimates with a pre-trained model is by applying a post-hoc recalibration method. Popular recalibration methods like temperature scaling are typically fit on a small amount of data and work in the model's output space, as opposed to the more expressive feature embedding space, and thus usually have only one or a handful of parameters. However, the target distribution to which they are applied is often complex and difficult to fit well with such a function. To this end we propose \textit{selective recalibration}, where a selection model learns to reject some user-chosen proportion of the data in order to allow the recalibrator to focus on regions of the input space that can be well-captured by such a model. We provide theoretical analysis to motivate our algorithm, and test our method through comprehensive experiments on difficult medical imaging and zero-shot classification tasks. Our results show that selective recalibration consistently leads to significantly lower calibration error than a wide range of selection and recalibration baselines.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Identifying and Calibrating Overconfidence in Noisy Speech Recognition

    eess.AS 2025-09 conditional novelty 6.0 of 10

    In noisy speech, Whisper often assigns high confidence to wrong tokens; a selective token-level temperature-scaling calibrator reduces ECE by 58% on the R-SPIN low-SNR range.

Pith tools