Pith. sign in

REVIEW 8 cited by

Calibration in Deep Learning: A Survey of the State-of-the-Art

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2308.01222 v4 pith:IGZYJHKQ submitted 2023-08-02 cs.LG cs.AI

classification cs.LGcs.AI
keywords calibrationmodelsdeepmodelmethodscalibratingrecentcalibrated
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Calibrating deep neural models plays an important role in building reliable, robust AI systems in safety-critical applications. Recent work has shown that modern neural networks that possess high predictive capability are poorly calibrated and produce unreliable model predictions. Though deep learning models achieve remarkable performance on various benchmarks, the study of model calibration and reliability is relatively under-explored. Ideal deep models should have not only high predictive performance but also be well calibrated. There have been some recent advances in calibrating deep models. In this survey, we review the state-of-the-art calibration methods and their principles for performing model calibration. First, we start with the definition of model calibration and explain the root causes of model miscalibration. Then we introduce the key metrics that can measure this aspect. It is followed by a summary of calibration methods that we roughly classify into four categories: post-hoc calibration, regularization methods, uncertainty estimation, and composition methods. We also cover recent advancements in calibrating large models, particularly large language models (LLMs). Finally, we discuss some open issues, challenges, and potential directions.

Discussion (0). Sign in to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 14 citations worldwide. Full citation record

  1. Bayesian-LoRA: Probabilistic Low-Rank Adaptation of Large Language Models

    cs.AI 2026-01 conditional novelty 7.0 of 10

    Treating LoRA adapters as sparse-GP random variables with a normalizing flow yields up to 84% ECE and 76% NLL reduction at ~1.2x training cost, reverting to standard LoRA in a point-mass limit.

  2. Cut Costs, Not Accuracy: LLM-Powered Data Processing with Guarantees

    cs.DB 2025-09 conditional novelty 7.0 of 10

    BARGAIN uses betting-based anytime-valid tests and adaptive, target-aware sampling to set model-cascade thresholds, delivering non-asymptotic quality guarantees and up to 86% greater cost savings than SUPG.

  3. PUF: Plug-and-Play Uncertainty-Aware Fusion for Online 3D Scene Graph Generation

    cs.CV 2026-07 conditional novelty 6.0 of 10

    PUF replaces deterministic 2D-to-3D scene graph fusion with probabilistic node association and Dirichlet evidence accumulation, yielding substantial accuracy gains on 3DSSG and ReplicaSSG at 15ms/frame.

  4. Towards Understanding The Calibration Benefits of Sharpness-Aware Minimization

    cs.LG 2025-05 reject novelty 6.0 of 10

    SAM's calibration benefit is attributed to implicit entropy maximization, but the proof rests on an unstated gradient-norm assumption; CSAM shows further ECE reductions.

  5. CRS-Triage: Confidence- and Reliability-Aware Selective Triage under Incomplete Clinical Evidence

    cs.LG 2026-08 conditional novelty 5.0 of 10

    CRS-Triage, which jointly models modality reliability, cross-modal consistency, and a learned confidence score, improves triage accuracy and reduces under-triage on MIMIC-IV-ED compared with evidential and fusion baselines.

  6. Calibrated and Robust Foundation Models for Vision-Language and Medical Image Tasks Under Distribution Shift

    cs.CV 2025-07 reject novelty 4.0 of 10

    StaRFM reuses the authors' earlier CalShift penalties, extends them to 3D medical segmentation with patch-wise and voxel-wise variants, and claims large gains that are not consistently supported by the paper's own tables.

  7. Why Uncertainty Calibration Matters for Reliable Perturbation-based Explanations

    cs.LG 2025-06 conditional novelty 4.0 of 10

    Calibration of a model under the exact perturbations used by an explanation method improves explanation fidelity, and ReCalX achieves this with per-perturbation-strength temperature scaling.

  8. Reality Check: A New Evaluation Ecosystem Is Necessary to Understand AI's Real World Effects

    cs.CY 2025-05 conditional novelty 4.0 of 10

    A position paper argues that understanding AI's second-order effects requires moving from static benchmarks to an ecosystem of field testing, red teaming, and contextual evaluation.

Pith tools