Pith. sign in

REVIEW 5 cited by

Extending Activation Steering to Broad Skills and Multiple Behaviours

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.05767 v1 pith:QF6QSDU4 submitted 2024-03-09 cs.LG cs.AIcs.CLcs.CY

classification cs.LGcs.AIcs.CLcs.CY
keywords steeringbehavioursskillsactivationmultipleabilitybecomebroad
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Current large language models have dangerous capabilities, which are likely to become more problematic in the future. Activation steering techniques can be used to reduce risks from these capabilities. In this paper, we investigate the efficacy of activation steering for broad skills and multiple behaviours. First, by comparing the effects of reducing performance on general coding ability and Python-specific ability, we find that steering broader skills is competitive to steering narrower skills. Second, we steer models to become more or less myopic and wealth-seeking, among other behaviours. In our experiments, combining steering vectors for multiple different behaviours into one steering vector is largely unsuccessful. On the other hand, injecting individual steering vectors at different places in a model simultaneously is promising.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Failure by Interference: Language Models Make Balanced Parentheses Errors When Faulty Mechanisms Overshadow Sound Ones

    cs.CL 2025-06 conditional novelty 6.0 of 10

    Language models fail at balanced parentheses because unreliable internal components that promote wrong tokens can outvote reliable ones, and amplifying reliable components fixes the errors.

  2. Improved Representation Steering for Language Models

    cs.CL 2025-05 conditional novelty 6.0 of 10

    RePS, a reference-free bidirectional preference optimization objective, improves representation steering and suppression for Gemma models, outperforming language-modeling objectives and approaching prompting performance.

  3. OPIUM: Mitigating Steering Externalities and Over-Refusal via Dual Objective Latent Optimization

    cs.LG 2026-07 conditional novelty 5.0 of 10

    OPIUM optimizes steering vectors in activation space so LLMs keep their intended behavior while shedding safety externalities and over-refusal, improving safety–utility tradeoff on Qwen-2.5 and LLaMA-3.1.

  4. Beyond Prompt Engineering: Robust Behavior Control in LLMs via Steering Target Atoms

    cs.CL 2025-05 conditional novelty 5.0 of 10

    STA selects sparse autoencoder features by activation amplitude and frequency to build steering vectors that improve LLM safety control over prompt engineering and standard steering.

  5. Probabilistic Concept-Aware Steering for Trustworthy LLM Inference

    cs.AI 2026-05 reject novelty 4.0 of 10

    PCS improves steering direction accuracy by adaptively sampling the intervention coefficient from a cosine-similarity-conditioned Gaussian, but its evaluation is partly circular because the optimal coefficient is chos...

Pith tools