Pith. sign in

REVIEW 4 cited by

Steering Large Language Models using Conceptors: Improving Addition-Based Activation Engineering

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.16314 v4 pith:ZR7RAZK6 submitted 2024-10-09 cs.NE cs.LG

classification cs.NEcs.LG
keywords conceptorssteeringactivationengineeringlanguagelargellmsmethods
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large language models have transformed AI, yet reliably controlling their outputs remains a challenge. This paper explores activation engineering, where outputs of pre-trained LLMs are controlled by manipulating their activations at inference time. Unlike traditional methods using a single steering vector, we introduce conceptors - mathematical constructs that represent sets of activation vectors as ellipsoidal regions. Conceptors act as soft projection matrices and offer more precise control over complex activation patterns. Our experiments demonstrate that conceptors outperform traditional methods across multiple steering tasks. We further use Boolean operations on conceptors for combined steering goals that empirically outperform additively combining steering vectors on a set of tasks. These results highlight conceptors as a promising tool for more effective steering of LLMs. Our code is available on github.com/jorispos/conceptorsteering.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Temporal Preference Concepts and their Functions in a Large Language Model

    cs.LG 2026-05 unverdicted novelty 6.5 of 10

    Temporal preference in Qwen3-4B-Instruct-2507 localizes to layers 17–35 (especially L24 attention), has curved residual-stream geometry, is behaviorally unstable, and can be bidirectionally steered.

  2. Latent-IM: Latent Interaction Management for Speech LLMs

    cs.CL 2026-07 conditional novelty 6.0 of 10

    A streaming residual-stream controller plus move-specific activation steering recovers selection and realization of five conversational moves in frozen speech LLMs, matching fine-tuning on human-move accuracy.

  3. Where Steering Signals Come From: Activation Source Selection in Activation Steering

    cs.CL 2026-07 conditional novelty 6.0 of 10

    Activation steering works best when the signal comes from the state where the model is about to produce the target behavior, not from text that already shows it.

  4. Probabilistic Concept-Aware Steering for Trustworthy LLM Inference

    cs.AI 2026-05 reject novelty 4.0 of 10

    PCS improves steering direction accuracy by adaptively sampling the intervention coefficient from a cosine-similarity-conditioned Gaussian, but its evaluation is partly circular because the optimal coefficient is chos...

Pith tools