Pith. sign in

REVIEW 1 cited by

Comparing Bottom-Up and Top-Down Steering Approaches on In-Context Learning Tasks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2411.07213 v1 pith:JLGVJULL submitted 2024-11-11 cs.LG

classification cs.LG
keywords steeringbottom-uptaskstop-downin-contextmethodsapproachesarxiv
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

A key objective of interpretability research on large language models (LLMs) is to develop methods for robustly steering models toward desired behaviors. To this end, two distinct approaches to interpretability -- ``bottom-up" and ``top-down" -- have been presented, but there has been little quantitative comparison between them. We present a case study comparing the effectiveness of representative vector steering methods from each branch: function vectors (FV; arXiv:2310.15213), as a bottom-up method, and in-context vectors (ICV; arXiv:2311.06668) as a top-down method. While both aim to capture compact representations of broad in-context learning tasks, we find they are effective only on specific types of tasks: ICVs outperform FVs in behavioral shifting, whereas FVs excel in tasks requiring more precision. We discuss the implications for future evaluations of steering methods and for further research into top-down and bottom-up steering given these findings.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Understanding (Un)Reliability of Steering Vectors in Language Models

    cs.LG 2025-05 conditional novelty 5.0 of 10

    Steering vectors are unreliable when the target behavior does not correspond to a coherent, well-separated linear direction in activation space, and this coherence can be measured from training data.

Pith tools