Pith. sign in

REVIEW 7 cited by

Thermometer: Towards Universal Calibration for Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.08819 v2 pith:64BT7HGN submitted 2024-02-20 cs.LG cs.CLstat.ML

classification cs.LGcs.CLstat.ML
keywords calibrationllmstasksthermometercalibratingchallengeslanguagelarge
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We consider the issue of calibration in large language models (LLM). Recent studies have found that common interventions such as instruction tuning often result in poorly calibrated LLMs. Although calibration is well-explored in traditional applications, calibrating LLMs is uniquely challenging. These challenges stem as much from the severe computational requirements of LLMs as from their versatility, which allows them to be applied to diverse tasks. Addressing these challenges, we propose THERMOMETER, a calibration approach tailored to LLMs. THERMOMETER learns an auxiliary model, given data from multiple tasks, for calibrating a LLM. It is computationally efficient, preserves the accuracy of the LLM, and produces better-calibrated responses for new tasks. Extensive empirical evaluations across various benchmarks demonstrate the effectiveness of the proposed method.

Discussion (0). Sign in to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Beyond Medical Diagnostics: How Medical Multimodal Large Language Models Think in Space

    cs.CV 2026-03 conditional novelty 6.0 of 10

    SpatialMed provides the first CT-based benchmark of 3D spatial reasoning for medical MLLMs, on which 14 models perform near chance, particularly for distance and volume estimation.

  2. Towards Objective Fine-tuning: How LLMs' Prior Knowledge Causes Potential Poor Calibration?

    cs.CL 2025-05 conditional novelty 6.0 of 10

    Fine-tuning on data aligned with an LLM's prior knowledge induces overconfidence, and CogCalib mitigates this by gating a calibration loss to known data.

  3. The Well-Tempered Classifier: Some Elementary Properties of Temperature Scaling

    stat.ML 2026-02 conditional novelty 5.0 of 10

    Temperature scaling is the unique accuracy-preserving linear recalibrator, and per-step tempering of a toy LLM can make sequence entropy non-monotonic.

  4. CDE: Curiosity-Driven Exploration for Efficient Reinforcement Learning in Large Language Models

    cs.CL 2025-09 conditional novelty 5.0 of 10

    Adding actor perplexity and multi-head critic variance as intrinsic exploration bonuses improves RLVR math reasoning accuracy by roughly +2 to +3 points on AIME benchmarks.

  5. Overconfidence in LLM-as-a-Judge: Diagnosis and Confidence-Driven Solution

    cs.AI 2025-08 unverdicted novelty 5.0 of 10

    LLM-as-a-Judge systems report confidence that overstates their accuracy, and the paper's TH-Score plus LLM-as-a-Fuser improves calibration.

  6. Insights into User Interface Innovations from a Design Thinking Workshop at deRSE25

    cs.HC 2025-08 conditional novelty 4.0 of 10

    A workshop at deRSE25 produced seven user-interface sketches for LLMs that emphasize branching, context management, and user weighting, which the authors map onto their whiteboard-based interface concept.

  7. Towards Harmonized Uncertainty Estimation for Large Language Models

    cs.CL 2025-05 conditional novelty 4.0 of 10

    CUE combines a supervised correctness classifier with existing LLM uncertainty scores to improve indication, balance, and calibration, reporting AUROC and ECE gains across models and datasets.

Pith tools