Pith. sign in

REVIEW 7 cited by

Steering Large Language Model Activations in Sparse Spaces

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2503.00177 v1 pith:SIXU6TAN submitted 2025-02-28 cs.LG cs.AI

classification cs.LGcs.AI
keywords sparseactivationfeaturesspacessteeringactivationsbehaviorbehaviors
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

A key challenge in AI alignment is guiding large language models (LLMs) to follow desired behaviors at test time. Activation steering, which modifies internal model activations during inference, offers a potential solution. However, prior work in dense activation spaces struggles with superposition, wherein multiple features become entangled, limiting interpretability and precise control. In contrast, sparse representations provide an untapped opportunity for more interpretable behavior modulation. In this work, we introduce sparse activation steering (SAS), a method that leverages sparse autoencoders (SAEs) to steer LLM behavior in sparse spaces. By isolating behavior-specific features through a contrastive prompt-pairing approach, we define a set of features that can selectively reinforce or suppress behaviors. Experiments on Gemma 2 LLMs show that SAS vectors enable nuanced behavioral modulation and finer-grained control. Furthermore, scaling SAEs improves monosemanticity of SAS vectors, suggesting more reliable and interpretable interventions.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SparSEEty: Extracting Tokens from Sparsity-Exploiting LLM Serving Systems via Deterministic Side Channels

    cs.CR 2026-08 conditional novelty 7.0 of 10

    Using page-fault side channels, an attacker can observe which FFN neurons a sparsity-exploiting LLM activates and invert those binary traces to recover prompt and response tokens with BLEU above 0.95.

  2. Sparse Autoencoders are Capable LLM Jailbreak Mitigators

    cs.CR 2026-02 conditional novelty 6.0 of 10

    CC-Delta defends LLMs against jailbreaks by statistically selecting and steering sparse-SAEs features that change when harmful prompts are embedded in jailbreak contexts, outperforming dense activation steering across...

  3. Can LLMs Lie? Investigation beyond Hallucination

    cs.LG 2025-09 conditional novelty 6.0 of 10

    The paper localizes LLM lying to sparse attention heads and chat-template 'dummy tokens', and shows steering vectors can modulate deception, but the evidence is weakened by selection and small samples.

  4. Resa: Transparent Reasoning Models via SAEs

    cs.CL 2025-06 conditional novelty 6.0 of 10

    SAE-Tuning, a sparse-autoencoder-guided SFT procedure, elicits RL-comparable reasoning in 1.5B models from CoT-free QA data at about $1 and 20 minutes of training.

  5. Beyond Interpretability: When, Why, and How Sparse Autoencoders Enable Label-Free Visual Steering

    cs.CV 2025-06 unverdicted novelty 6.0 of 10

    VS2 constructs steering vectors from sparse SAE features on unlabeled in-domain activations to improve zero-shot accuracy of CLIP models by 0.93-4.12% on CIFAR-100, CUB-200, and Tiny-ImageNet while remaining forward-p...

  6. STEER-BENCH: A Benchmark for Evaluating the Steerability of Large Language Models

    cs.CL 2025-05 conditional novelty 6.0 of 10

    STEER-BENCH is a Reddit-derived benchmark of 5,552 multiple-choice questions on which the best of 13 large language models scores near 65 percent, versus human experts near 81 percent.

  7. Quantifying Conversation Drift in MCP via Latent Polytope

    cs.CL 2025-08 reject novelty 4.0 of 10

    SecMCP flags MCP conversation drift by thresholding per-layer activation distances from benign anchors, reporting AUROC above 0.915 on Llama3, Vicuna, and Mistral.

Pith tools