REVIEW 7 cited by
Thermometer: Towards Universal Calibration for Large Language Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We consider the issue of calibration in large language models (LLM). Recent studies have found that common interventions such as instruction tuning often result in poorly calibrated LLMs. Although calibration is well-explored in traditional applications, calibrating LLMs is uniquely challenging. These challenges stem as much from the severe computational requirements of LLMs as from their versatility, which allows them to be applied to diverse tasks. Addressing these challenges, we propose THERMOMETER, a calibration approach tailored to LLMs. THERMOMETER learns an auxiliary model, given data from multiple tasks, for calibrating a LLM. It is computationally efficient, preserves the accuracy of the LLM, and produces better-calibrated responses for new tasks. Extensive empirical evaluations across various benchmarks demonstrate the effectiveness of the proposed method.
Forward citations
Cited by 7 Pith papers
-
Beyond Medical Diagnostics: How Medical Multimodal Large Language Models Think in Space
SpatialMed provides the first CT-based benchmark of 3D spatial reasoning for medical MLLMs, on which 14 models perform near chance, particularly for distance and volume estimation.
-
Towards Objective Fine-tuning: How LLMs' Prior Knowledge Causes Potential Poor Calibration?
Fine-tuning on data aligned with an LLM's prior knowledge induces overconfidence, and CogCalib mitigates this by gating a calibration loss to known data.
-
The Well-Tempered Classifier: Some Elementary Properties of Temperature Scaling
Temperature scaling is the unique accuracy-preserving linear recalibrator, and per-step tempering of a toy LLM can make sequence entropy non-monotonic.
-
CDE: Curiosity-Driven Exploration for Efficient Reinforcement Learning in Large Language Models
Adding actor perplexity and multi-head critic variance as intrinsic exploration bonuses improves RLVR math reasoning accuracy by roughly +2 to +3 points on AIME benchmarks.
-
Overconfidence in LLM-as-a-Judge: Diagnosis and Confidence-Driven Solution
LLM-as-a-Judge systems report confidence that overstates their accuracy, and the paper's TH-Score plus LLM-as-a-Fuser improves calibration.
-
Insights into User Interface Innovations from a Design Thinking Workshop at deRSE25
A workshop at deRSE25 produced seven user-interface sketches for LLMs that emphasize branching, context management, and user weighting, which the authors map onto their whiteboard-based interface concept.
-
Towards Harmonized Uncertainty Estimation for Large Language Models
CUE combines a supervised correctness classifier with existing LLM uncertainty scores to improve indication, balance, and calibration, reporting AUROC and ECE gains across models and datasets.
Discussion (0). Sign in to comment.