Pith. sign in

REVIEW 2 cited by

Calibrated Selective Classification

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2208.12084 v2 pith:7AFL6DTY submitted 2022-08-25 cs.LG

classification cs.LG
keywords predictionsselectiveclassificationcalibratedcalibrationestimatesmodelsuncertainty
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Selective classification allows models to abstain from making predictions (e.g., say "I don't know") when in doubt in order to obtain better effective accuracy. While typical selective models can be effective at producing more accurate predictions on average, they may still allow for wrong predictions that have high confidence, or skip correct predictions that have low confidence. Providing calibrated uncertainty estimates alongside predictions -- probabilities that correspond to true frequencies -- can be as important as having predictions that are simply accurate on average. However, uncertainty estimates can be unreliable for certain inputs. In this paper, we develop a new approach to selective classification in which we propose a method for rejecting examples with "uncertain" uncertainties. By doing so, we aim to make predictions with {well-calibrated} uncertainty estimates over the distribution of accepted examples, a property we call selective calibration. We present a framework for learning selectively calibrated models, where a separate selector network is trained to improve the selective calibration error of a given base model. In particular, our work focuses on achieving robust calibration, where the model is intentionally designed to be tested on out-of-domain data. We achieve this through a training strategy inspired by distributionally robust optimization, in which we apply simulated input perturbations to the known, in-domain training data. We demonstrate the empirical effectiveness of our approach on multiple image classification and lung cancer risk assessment tasks.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 5 citations worldwide. Full citation record

  1. Identifying and Calibrating Overconfidence in Noisy Speech Recognition

    eess.AS 2025-09 conditional novelty 6.0 of 10

    In noisy speech, Whisper often assigns high confidence to wrong tokens; a selective token-level temperature-scaling calibrator reduces ECE by 58% on the R-SPIN low-SNR range.

  2. Interpretable Failure Detection with Human-Level Concepts

    cs.CV 2025-02 conditional novelty 6.0 of 10

    ORCA ranks concept activations from CLIP and uses the rank-weighted agreement of the top-K concepts with the predicted category as its confidence score, improving failure-detection FPR on several benchmarks.

Pith tools