Pith. sign in

REVIEW 5 cited by

Batch Calibration: Rethinking Calibration for In-Context Learning and Prompt Engineering

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2309.17249 v3 pith:OAGLHKDU submitted 2023-09-29 cs.CL cs.AIcs.LG

classification cs.CLcs.AIcs.LG
keywords calibrationbiaslearningperformancepromptbatchcontextualin-context
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Prompting and in-context learning (ICL) have become efficient learning paradigms for large language models (LLMs). However, LLMs suffer from prompt brittleness and various bias factors in the prompt, including but not limited to the formatting, the choice verbalizers, and the ICL examples. To address this problem that results in unexpected performance degradation, calibration methods have been developed to mitigate the effects of these biases while recovering LLM performance. In this work, we first conduct a systematic analysis of the existing calibration methods, where we both provide a unified view and reveal the failure cases. Inspired by these analyses, we propose Batch Calibration (BC), a simple yet intuitive method that controls the contextual bias from the batched input, unifies various prior approaches, and effectively addresses the aforementioned issues. BC is zero-shot, inference-only, and incurs negligible additional costs. In the few-shot setup, we further extend BC to allow it to learn the contextual bias from labeled data. We validate the effectiveness of BC with PaLM 2-(S, M, L) and CLIP models and demonstrate state-of-the-art performance over previous calibration baselines across more than 10 natural language understanding and image classification tasks.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. From Words to Widgets for Controllable LLM Generation

    cs.HC 2026-04 unverdicted novelty 6.5 of 10

    Reifying ad-hoc preference phrases as GUI widgets, steered by log-probability modulation at decode time, improves precision and perceived control of LLM writing over natural-language prompting alone.

  2. Surprise Calibration for Better In-Context Learning

    cs.CL 2025-06 conditional novelty 6.0 of 10

    Surprise Calibration uses the model's own surprise at each demonstration's label to dynamically correct class priors in in-context learning, improving accuracy on eight NLP benchmarks.

  3. CausalAbstain: Enhancing Multilingual LLMs with Causal Reasoning for Trustworthy Abstention

    cs.CL 2025-05 conditional novelty 6.0 of 10

    CausalAbstain filters multilingual self-feedback by comparing how much it changes the model's abstention decision, improving abstention accuracy over baselines on two benchmarks.

  4. Lost in Context: Addressing Context Anxiety in Large Language Models

    cs.AI 2026-05 reject novelty 5.0 of 10

    Context anxiety — abandoning solvable tasks over perceived token limits — is measurable and reducible by fine-tuning on anxiety-free reasoning traces, but the paper's causal mechanism is not actually tested.

  5. Probabilistic Concept-Aware Steering for Trustworthy LLM Inference

    cs.AI 2026-05 reject novelty 4.0 of 10

    PCS improves steering direction accuracy by adaptively sampling the intervention coefficient from a cosine-similarity-conditioned Gaussian, but its evaluation is partly circular because the optimal coefficient is chos...

Pith tools