Pith. sign in

REVIEW 2 cited by

The Calibration Generalization Gap

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2210.01964 v2 pith:MFBPUAXA submitted 2022-10-05 cs.LG cs.AIcs.CVstat.ML

classification cs.LGcs.AIcs.CVstat.ML
keywords calibrationgeneralizationerrormodeltrainaugmentationcalibrateddata
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Calibration is a fundamental property of a good predictive model: it requires that the model predicts correctly in proportion to its confidence. Modern neural networks, however, provide no strong guarantees on their calibration -- and can be either poorly calibrated or well-calibrated depending on the setting. It is currently unclear which factors contribute to good calibration (architecture, data augmentation, overparameterization, etc), though various claims exist in the literature. We propose a systematic way to study the calibration error: by decomposing it into (1) calibration error on the train set, and (2) the calibration generalization gap. This mirrors the fundamental decomposition of generalization. We then investigate each of these terms, and give empirical evidence that (1) DNNs are typically always calibrated on their train set, and (2) the calibration generalization gap is upper-bounded by the standard generalization gap. Taken together, this implies that models with small generalization gap (|Test Error - Train Error|) are well-calibrated. This perspective unifies many results in the literature, and suggests that interventions which reduce the generalization gap (such as adding data, using heavy augmentation, or smaller model size) also improve calibration. We thus hope our initial study lays the groundwork for a more systematic and comprehensive understanding of the relation between calibration, generalization, and optimization.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Distillation Scaling Laws

    cs.LG 2025-02 conditional novelty 7.0 of 10

    A distillation scaling law predicts student cross-entropy from teacher loss, student size, and data, and gives compute-optimal teacher-student allocations.

  2. Rethinking Early Stopping: Refine, Then Calibrate

    cs.LG 2025-01 conditional novelty 6.0 of 10

    Stopping training on the loss after temperature scaling (a refinement estimate) instead of the raw validation loss, then applying temperature scaling afterwards, lowers test logloss.

Pith tools