Pith. sign in

REVIEW 10 cited by

Calibration of Pre-trained Transformers

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2003.07892 v3 pith:QCAXLNMO submitted 2020-03-17 cs.CL cs.LG

classification cs.CLcs.LG
keywords calibrationin-domainmodelsout-of-domainpre-trainedcalibratedempiricalerror
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Pre-trained Transformers are now ubiquitous in natural language processing, but despite their high end-task performance, little is known empirically about whether they are calibrated. Specifically, do these models' posterior probabilities provide an accurate empirical measure of how likely the model is to be correct on a given example? We focus on BERT and RoBERTa in this work, and analyze their calibration across three tasks: natural language inference, paraphrase detection, and commonsense reasoning. For each task, we consider in-domain as well as challenging out-of-domain settings, where models face more examples they should be uncertain about. We show that: (1) when used out-of-the-box, pre-trained models are calibrated in-domain, and compared to baselines, their calibration error out-of-domain can be as much as 3.5x lower; (2) temperature scaling is effective at further reducing calibration error in-domain, and using label smoothing to deliberately increase empirical uncertainty helps calibrate posteriors out-of-domain.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. How Good are LLM-based Rerankers? An Empirical Analysis of State-of-the-Art Reranking Models

    cs.CL 2025-08 conditional novelty 7.0 of 10

    On a new benchmark of post-April 2025 queries, LLM rerankers show a 5-15% performance drop compared with familiar benchmarks, and lightweight models match them on efficiency and sometimes accuracy.

  2. Rationales Are Not Silver Bullets: Measuring the Impact of Rationales on Model Performance and Reliability

    cs.CL 2025-05 conditional novelty 7.0 of 10

    Rationale-augmented finetuning can hurt accuracy while improving calibration, with the sizes of both effects tied linearly to task difficulty.

  3. Fact-Level Confidence Calibration and Self-Correction

    cs.CL 2024-11 conditional novelty 6.0 of 10

    The paper measures LLM confidence per atomic fact, weights correctness by relevance, and uses high-confidence facts from the same response to correct low-confidence facts.

  4. Enhancing Trustworthy GUI Grounding via Self-Critiqued Reinforcement Learning

    cs.CV 2025-10 conditional novelty 5.0 of 10

    HyperClick trains GUI grounding models with GRPO to output clicks plus confidence scores, jointly rewarding correct clicks and Brier-calibrated confidence, and reports SOTA accuracy on six of seven benchmarks with bet...

  5. The Capabilities and Limitations of Weak-to-Strong Generalization: Generalization and Calibration

    cs.LG 2025-02 reject novelty 5.0 of 10

    The paper derives generalization and calibration bounds for weak-to-strong generalization and extends a known regression result from squared loss to KL divergence.

  6. Temporalizing Confidence: Evaluation of Chain-of-Thought Reasoning with Signal Temporal Logic

    cs.LG 2025-06 conditional novelty 4.0 of 10

    Stepwise CoT confidence is reshaped and scored with signal temporal logic robustness to produce better calibrated confidence estimates on Gaokao math questions.

  7. How Knowledge Popularity Influences and Enhances LLM Knowledge Boundary Perception

    cs.CL 2025-05 conditional novelty 4.0 of 10

    Entity popularity and entity co-occurrence in Wikipedia correlate with LLM QA accuracy, confidence, and calibration, and combining them with confidence improves answer-correctness prediction by 5.24% on average.

  8. A Survey on Uncertainty Quantification of Large Language Models: Taxonomy, Open Research Challenges, and Future Directions

    cs.CL 2024-12 conditional novelty 4.0 of 10

    A review that organizes LLM uncertainty quantification into token-level, self-verbalized, semantic-similarity, and mechanistic interpretability categories.

  9. Distributed Collaborative Inference System in Next-Generation Networks and Communication

    cs.NI 2024-11 reject novelty 3.0 of 10

    A multi-level cloud-edge-end inference system combining early exit, attention-based pruning, and confidence-based offloading reduces BERT sentiment-analysis latency by up to 17%, but only at measurable accuracy cost.

  10. A Comprehensive Guide to Explainable AI: From Classical Models to LLMs

    cs.LG 2024-12 unverdicted novelty 1.0 of 10

    A survey-style XAI book with code examples, covering standard interpretability methods and models, but no new scientific contributions.

Pith tools