Pith. sign in

REVIEW 3 cited by

Rethinking Early Stopping: Refine, Then Calibrate

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2501.19195 v2 pith:RTYK4HCG submitted 2025-01-31 cs.LG cs.AI

classification cs.LGcs.AI
keywords calibrationerrorrefinementcalibratedifferentduringminimizingpredictions
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Machine learning classifiers often produce probabilistic predictions that are critical for accurate and interpretable decision-making in various domains. The quality of these predictions is generally evaluated with proper losses, such as cross-entropy, which decompose into two components: calibration error assesses general under/overconfidence, while refinement error measures the ability to distinguish different classes. In this paper, we present a novel variational formulation of the calibration-refinement decomposition that sheds new light on post-hoc calibration, and enables rapid estimation of the different terms. Equipped with this new perspective, we provide theoretical and empirical evidence that calibration and refinement errors are not minimized simultaneously during training. Selecting the best epoch based on validation loss thus leads to a compromise point that is suboptimal for both terms. To address this, we propose minimizing refinement error only during training (Refine,...), before minimizing calibration error post hoc, using standard techniques (...then Calibrate). Our method integrates seamlessly with any classifier and consistently improves performance across diverse classification tasks.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Eigenvalue Calibration for Semantic Embeddings of Large Language Models

    cs.LG 2026-07 conditional novelty 6.5 of 10

    Temperature scaling of density-matrix eigenvalues from LLM semantic embeddings optimizes proper-score calibration and corrects systematic overconfidence so entropy equals risk.

  2. When Shift Happens - Confounding Is to Blame

    cs.LG 2025-05 conditional novelty 6.0 of 10

    Under hidden confounding shifts, predictive information reduces to conditional informativeness minus a residual, a result the authors use to explain ERM's surprising OOD competitiveness and the value of all-covariate models.

  3. The Well-Tempered Classifier: Some Elementary Properties of Temperature Scaling

    stat.ML 2026-02 conditional novelty 5.0 of 10

    Temperature scaling is the unique accuracy-preserving linear recalibrator, and per-step tempering of a toy LLM can make sequence entropy non-monotonic.

Pith tools