Pith. sign in

REVIEW 1 cited by

Towards Understanding Why Label Smoothing Degrades Selective Classification and How to Fix It

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.14715 v3 pith:MMS5MDND submitted 2024-03-19 cs.LG cs.AIcs.CV

classification cs.LGcs.AIcs.CV
keywords degradesclassificationcorrectdemonstrateeffectiveexplanationlabellikely
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Label smoothing (LS) is a popular regularisation method for training neural networks as it is effective in improving test accuracy and is simple to implement. ``Hard'' one-hot labels are ``smoothed'' by uniformly distributing probability mass to other classes, reducing overfitting. Prior work has suggested that in some cases LS can degrade selective classification (SC) -- where the aim is to reject misclassifications using a model's uncertainty. In this work, we first demonstrate empirically across an extended range of large-scale tasks and architectures that LS consistently degrades SC. We then address a gap in existing knowledge, providing an explanation for this behaviour by analysing logit-level gradients: LS degrades the uncertainty rank ordering of correct vs incorrect predictions by suppressing the max logit more when a prediction is likely to be correct, and less when it is likely to be wrong. This elucidates previously reported experimental results where strong classifiers underperform in SC. We then demonstrate the empirical effectiveness of post-hoc logit normalisation for recovering lost SC performance caused by LS. Furthermore, linking back to our gradient analysis, we again provide an explanation for why such normalisation is effective.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. BayesAdapter: enhanced uncertainty estimation in CLIP few-shot adaptation

    cs.CV 2024-12 conditional novelty 6.0 of 10

    Applying variational Bayesian inference to the CLAP linear probe adapter improves calibration and high-confidence coverage in CLIP few-shot classification, with a modest accuracy trade-off.

Pith tools