Pith. sign in

REVIEW 4 cited by

Which Attention Heads Matter for In-Context Learning?

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2502.14010 v1 pith:E7ELTB7A submitted 2025-02-19 cs.LG cs.AIcs.CL

classification cs.LGcs.AIcs.CL
keywords headsinductionlearningmodelsdrivesin-contextlanguagemechanism
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large language models (LLMs) exhibit impressive in-context learning (ICL) capability, enabling them to perform new tasks using only a few demonstrations in the prompt. Two different mechanisms have been proposed to explain ICL: induction heads that find and copy relevant tokens, and function vector (FV) heads whose activations compute a latent encoding of the ICL task. To better understand which of the two distinct mechanisms drives ICL, we study and compare induction heads and FV heads in 12 language models. Through detailed ablations, we discover that few-shot ICL performance depends primarily on FV heads, especially in larger models. In addition, we uncover that FV and induction heads are connected: many FV heads start as induction heads during training before transitioning to the FV mechanism. This leads us to speculate that induction facilitates learning the more complex FV mechanism that ultimately drives ICL.

Discussion (0). Sign in to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Compressed Sensing for Capability Localization in Large Language Models

    cs.CL 2026-02 conditional novelty 6.0 of 10

    LLM capabilities are concentrated in small sets of attention heads, and a compressed-sensing method can find those heads efficiently.

  2. Predicting the Emergence of Induction Heads in Language Model Pretraining

    cs.CL 2025-11 conditional novelty 6.0 of 10

    A fitted law UPT = T/√(BC) predicts when induction heads emerge, and a Pareto frontier in bigram repetition frequency/reliability separates data that produce induction heads from data that do not.

  3. CompressKV: Semantic Retrieval Heads Know What Tokens are Not Important Before Generation

    cs.CL 2025-08 unverdicted novelty 6.0 of 10

    A method that uses a handful of 'semantic retrieval heads' instead of all attention heads to decide which key-value cache entries can be dropped, plus layer-wise cache budgeting, reportedly beats prior KV compression ...

  4. Distinct Computations Emerge From Compositional Curricula in In-Context Learning

    cs.LG 2025-06 conditional novelty 6.0 of 10

    When transformer models see easy component examples before a harder combined math problem in one prompt, they solve unseen versions of the combined problem and store intermediate steps internally, unlike models traine...

Pith tools