Pith. sign in

REVIEW 5 cited by

Calibration of Pre-trained Transformers

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2003.07892 v3 pith:QCAXLNMO submitted 2020-03-17 cs.CL cs.LG

classification cs.CLcs.LG
keywords calibrationin-domainmodelsout-of-domainpre-trainedcalibratedempiricalerror
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Pre-trained Transformers are now ubiquitous in natural language processing, but despite their high end-task performance, little is known empirically about whether they are calibrated. Specifically, do these models' posterior probabilities provide an accurate empirical measure of how likely the model is to be correct on a given example? We focus on BERT and RoBERTa in this work, and analyze their calibration across three tasks: natural language inference, paraphrase detection, and commonsense reasoning. For each task, we consider in-domain as well as challenging out-of-domain settings, where models face more examples they should be uncertain about. We show that: (1) when used out-of-the-box, pre-trained models are calibrated in-domain, and compared to baselines, their calibration error out-of-domain can be as much as 3.5x lower; (2) temperature scaling is effective at further reducing calibration error in-domain, and using label smoothing to deliberately increase empirical uncertainty helps calibrate posteriors out-of-domain.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. How Good are LLM-based Rerankers? An Empirical Analysis of State-of-the-Art Reranking Models

    cs.CL 2025-08 conditional novelty 7.0 of 10

    On a new benchmark of post-April 2025 queries, LLM rerankers show a 5-15% performance drop compared with familiar benchmarks, and lightweight models match them on efficiency and sometimes accuracy.

  2. Rationales Are Not Silver Bullets: Measuring the Impact of Rationales on Model Performance and Reliability

    cs.CL 2025-05 conditional novelty 7.0 of 10

    Rationale-augmented finetuning can hurt accuracy while improving calibration, with the sizes of both effects tied linearly to task difficulty.

  3. Enhancing Trustworthy GUI Grounding via Self-Critiqued Reinforcement Learning

    cs.CV 2025-10 conditional novelty 5.0 of 10

    HyperClick trains GUI grounding models with GRPO to output clicks plus confidence scores, jointly rewarding correct clicks and Brier-calibrated confidence, and reports SOTA accuracy on six of seven benchmarks with bet...

  4. Temporalizing Confidence: Evaluation of Chain-of-Thought Reasoning with Signal Temporal Logic

    cs.LG 2025-06 conditional novelty 4.0 of 10

    Stepwise CoT confidence is reshaped and scored with signal temporal logic robustness to produce better calibrated confidence estimates on Gaokao math questions.

  5. How Knowledge Popularity Influences and Enhances LLM Knowledge Boundary Perception

    cs.CL 2025-05 conditional novelty 4.0 of 10

    Entity popularity and entity co-occurrence in Wikipedia correlate with LLM QA accuracy, confidence, and calibration, and combining them with confidence improves answer-correctness prediction by 5.24% on average.

Pith tools