Pith. sign in

REVIEW 2 major objections 5 minor 14 references

Chat-calibrated additive steering reaches agent residual streams at near-full strength, yet its behavioral grip is rescaled model-by-model with no universal factor.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-13 05:01 UTC pith:GOO63JXW

load-bearing objection First careful chat-to-agent transfer study of additive steering: direction survives, behavioral coupling rescales per model (can amplify or attenuate), with solid controls and honest limits. the 2 major comments →

arxiv 2607.09156 v1 pith:GOO63JXW submitted 2026-07-10 cs.LG

Present but Rescaled: Chat-to-Agent Transfer of Additive Activation Steering

classification cs.LG
keywords activation steeringadditive injectionReAct agentschat-to-agent transferrefusal bypassresidual streamrepresentation engineeringbehavioral coupling
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper shows that additive activation steering—injecting a chat-extracted residual-stream direction—survives the move from single-turn chat into tool-using ReAct agents as a near-full representational signal. What does not survive is the behavioral coupling: the same surviving direction can amplify refusal bypass by up to 2× on some models and attenuate it on others. The rescaling is already locked in by the ReAct format scaffold itself, before any tool observation arrives, and it is specific to continuous additive injection rather than permanent ablation of the same axis. Because models are increasingly deployed as agents, a chat safety or control number cannot be assumed to bound agent behavior under the same intervention. The practical upshot is that monitoring still works while control must be recalibrated per model and per deployment format.

Core claim

Transfer of additive activation steering from chat to ReAct agents is real but dissociated: the injected direction reaches late layers at near-full strength (install-site agent-over-chat ratios 0.83–1.16), while behavioral coupling is reset per model, spanning amplification (Gemma-2-9B T=2.00) to attenuation (Yi-1.5-9B T=0.43) with no universal constant or sign. The rescaling localizes to the ReAct format scaffold before tool observations and is additive-specific: directional ablation of the same axis does not amplify while additive injection does.

What carries the argument

The matched-information ladder (plain chat C0 versus ReAct-with-tool C3) paired with the transfer ratio T = Δagent/Δchat, under matched-norm random-direction controls and full transcript re-encoding every turn. It holds the instruction byte-identical while isolating the deployment wrapper, letting representation survival and behavioral coupling be measured separately.

Load-bearing premise

The parser that scores refusal or compliance is equally sensitive on free-form chat replies and on ReAct Thought tokens, so the transfer ratio is not an artifact of the scorer working better in one format than the other.

What would settle it

Re-run the uniform-protocol experiment on held-out model families under the same ladder and random controls: if behavioral T is statistically indistinguishable from 1 with tight intervals while install-site projection ratios remain near 1, or if removing only the ReAct format scaffold leaves T unchanged while tool-observation insertion moves it, the claimed dissociation and format-priming localization fail.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Agentic deployment can amplify steering-based refusal bypass by up to 2× on some models, so chat-only safety evaluations understate the hazard.
  • Activation probes trained in chat continue to fire at near-chat strength in agents; control coefficients must be recalibrated in the deployment context.
  • Coupling has no universal constant or sign across families; each model requires its own agent-side characterization.
  • Rescaling is set by format priming, not by tool observations or residual-norm dilution, so observation-content interventions miss the main effect.
  • Additive injection and directional ablation transfer differently: the former must be re-asserted every step, the latter does not.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • Other structured agent formats (code-agent loops, plan-act-observe variants) may each impose their own distinct rescale factors that should be measured before deployment.
  • The existence of both clean amplifiers and a clean attenuator points to alignment-training geometry as the likely determinant of sign; a broader census could map which recipes produce attenuation.
  • Any long-horizon setting whose residual norm grows may systematically under-power one-shot or prefill-only steering relative to continuous per-token injection.
  • Red-teaming that ranks models only under chat-calibrated additive jailbreaks will mis-order them once the same models run as agents.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 5 minor

Summary. The paper presents the first systematic study of chat-to-agent transfer for additive residual-stream steering. Using a matched-information ladder (C0 plain chat vs C3 ReAct with a deterministic tool), matched-norm random controls, full-transcript re-encoding each turn, and a dual readout (late-layer projection survival + parser-based behavioral transfer ratio T = Δagent/Δchat), it reports a dissociation: the injected direction reaches late layers at near-full strength (install-site agent-over-chat ratios 0.83–1.16 across Qwen2.5-7B, Llama-3.1-8B, Gemma-2-9B) while behavioral coupling is reset per model, spanning amplification (Gemma T=2.00, Qwen T=1.41–1.45) to attenuation (Yi-1.5-9B T=0.43). Additive injection amplifies while directional ablation of the same axis does not (Φ=20.1 pt gap); two pre-registered instruments localize the rescaling to the ReAct format scaffold before any tool observation. Safety implication: agentic deployment can amplify or attenuate refusal-bypass steering in a model-specific way that chat calibration does not predict.

Significance. If the dissociation and two-sided coupling distribution hold, the result is immediately useful for both mechanistic interpretability and deployment safety. Additive steering is already proposed as a control/monitoring primitive; the paper shows that representation survival does not imply behavioral grip transfer, that the effect is additive-specific (vs ablation), and that the rescaling is set by format priming rather than observation dilution. Strengths that raise the bar include pre-registration of live/die criteria and dose grids, item-paired bootstrap CIs (B=10k), matched-norm random bands at every cell, a capability-matched verbosity control, convergent localization instruments, and explicit disclosure of the two-pass finer-ladder rescue that recovered the attenuator. These make the central claims falsifiable and reproducible rather than post-hoc.

major comments (2)
  1. [The Setting-Invariant Metric; Table 1] The transfer ratio T rests on a setting-invariant binary parser (refuses/complies) applied identically to free-form chat replies and ReAct Thought tokens. The manuscript reports 83.8% agreement with blind human labels on a stratified 100-item slice, but does not break that agreement down by rung (C0 vs C3) or report false-positive/false-negative rates under the ReAct scaffold. Because T is a ratio of deltas, differential parser sensitivity or ceiling effects across formats could inflate or deflate the reported amplification/attenuation. A short per-rung confusion matrix or human re-label of a C3-only slice would close this load-bearing measurement risk.
  2. [Per-Model Distribution; Figure 2; Limitations] The claim of 'no universal sign' is anchored on a single clean attenuator (Yi-1.5-9B T=0.43), recovered only after a pre-registered finer-ladder second pass. While the two-pass structure and pre-registration of Yi's sign are disclosed, five multi-dose families plus one single-dose boundary family leave the two-sided distribution asymmetrically supported. The safety conclusion that 'a deployment cannot assume a given model is safe' is still warranted by the amplifiers alone, but the stronger phrasing of no universal sign should be tempered or supported by at least one additional attenuating family if available.
minor comments (5)
  1. [Table 1] Table 1 footnote on Qwen2-7B single-dose T and the gate-fail families is dense; a short column or legend clarifying which T values are AUC-over-doses versus single-dose point ratios would help readers.
  2. [Abstract; Behavioral leg: sycophancy] Sycophancy T=0.78 CI spans 1; the text correctly treats the cross-behavior interaction as the formal result, but the abstract and safety paragraphs still lean on 'attenuate' language that should stay qualified.
  3. [Representational leg; Table 3] Install-site ratios are reported as 0.83–1.16 across three families; the main text and Table 3 give more granular layer-wise numbers for Qwen. A compact cross-family install-site table would make the survival half easier to cite.
  4. [The Matched-Information Ladder] C4 is mentioned as held out for tau2-bench but never used; either drop the rung from the ladder description or note that external-benchmark transfer is future work more prominently in the design section.
  5. [Abstract; Introduction] Minor typography: several compound terms appear without spaces or hyphens in the abstract and early sections (e.g., 'Additiveactivationsteering', 'chat-to-agenttransfer'); these are likely PDF extraction artifacts but should be cleaned for the camera-ready version.

Circularity Check

0 steps flagged

No significant circularity: transfer ratios T, install-site survival, additive-vs-ablation gaps, and frame-priming localization are measured quantities under pre-registered matched designs, not forced by definition or self-citation.

full rationale

This is a self-contained empirical measurement paper. The central quantities (T = Δagent/Δchat from parser-scored refusal/sycophancy rates on matched items, install-site residual projections onto the unit steering direction, AUC-over-doses under a uniform protocol, and the 20.1-point additive-vs-ablation gain difference) are computed from observed behavioral and representational read-outs after injection; they are not algebraically identical to any fitted input or definitional identity. Operating layers and sub-saturation dose grids are fixed from chat pilots before any agent cell runs (explicitly pre-registered, including the two-pass rescue structure and Yi’s attenuating sign from a held-out prior run). Matched-norm random-direction bands, KV recompute, capability-matched verbosity controls, and item-paired bootstrap CIs further gate direction-specificity rather than bake in the result. The Belief-Dynamics head-to-head treats agent cells as out-of-sample predictions from a chat-fitted slope and reports large gaps (e.g., 67 points on induce). No uniqueness theorem, ansatz, or load-bearing self-citation by the sole author is invoked to force the dissociation or the two-sided coupling distribution. The paper therefore satisfies the hard rule for an honest non-finding of circularity.

Axiom & Free-Parameter Ledger

4 free parameters · 5 axioms · 2 invented entities

The paper is empirical; its load-bearing content is measurement design rather than free parameters or new ontological entities. Free parameters are the pilot-chosen layers, dose grids, and the AMP_MARGIN threshold that labels amplify/boundary/attenuate. Domain assumptions are the standard residual-stream difference-of-means extraction and the claim that a parser-based binary is setting-invariant. No new physical or mathematical entities are postulated; T and the ladder are operational definitions.

free parameters (4)
  • per-family operating layer and sub-saturation dose grid
    Fixed from chat pilots before agent cells; different for each model (e.g., Qwen L16 c12/c20, Gemma L20 c80/c88, Yi L30 c8–c11). Choice affects which families pass the gate and the numerical value of Tfamily.
  • AMP_MARGIN = 1.10
    Pre-committed threshold that converts a continuous T interval into the discrete labels AMPLIFY / BOUNDARY / ATTENUATE used in the distribution claim.
  • injection coefficient α (c8, c12, c16, …)
    Chosen to produce 25–60 pp chat swing while remaining sub-saturation in the agent frame; the registered primary coefficient for Qwen bypass is c16.
  • read-layer and install-site prefix length
    Late-layer projection site and fixed-length system-prompt prefix used for the representation-survival ratio; behavior-independent by construction but still a design choice.
axioms (5)
  • domain assumption Difference-of-means residual-stream directions extracted on last-token activations of harmful vs harmless (or trait+/trait−) completions are the correct objects for additive steering.
    Standard recipe of Zou et al. / Arditi et al. / Panickssery et al.; adopted without re-derivation.
  • domain assumption A validated parser-based binary (refuses/complies, agrees/disagrees) applied identically to natural-language output is setting-invariant across plain chat and ReAct Thought tokens.
    83.8% human agreement on a 100-item slice is the only external validation; the transfer ratio T inherits any residual non-invariance.
  • domain assumption Re-encoding the full transcript from scratch each turn fully excludes KV-cache contamination as a mechanism.
    Stated as construction; rules out Kang et al. (2026) but assumes no other cross-turn state leakage.
  • domain assumption Matched-norm random unit vectors (n_rand ≥ 5) at the same coefficient constitute a sufficient direction-specificity gate.
    Used throughout; real effect must exceed max |random effect|.
  • standard math Bootstrap item-paired CIs with B=10 000 correctly capture uncertainty for the transfer ratios and gain differences.
    Standard non-parametric bootstrap; no parametric model of the generative process is assumed.
invented entities (2)
  • matched-information five-rung ladder (C0–C4) independent evidence
    purpose: Holds the harmful instruction byte-identical while varying only the deployment wrapper so that chat-to-agent differences can be attributed to format/context rather than content.
    Operational experimental design, not a new theoretical object; independent evidence is the internal consistency of the localization instruments.
  • transfer ratio T = Δagent / Δchat independent evidence
    purpose: Scalar summary of behavioral coupling under matched items.
    Defined from observed rate deltas; not an unobserved mediator.

pith-pipeline@v1.1.0-grok45 · 22625 in / 3617 out tokens · 34880 ms · 2026-07-13T05:01:00.908189+00:00 · methodology

0 comments
read the original abstract

Additive activation steering (injecting a scaled residual-stream direction during generation) is calibrated almost entirely in single-turn chat, yet the models it targets are increasingly deployed as tool-using ReAct agents. We present the first systematic chat-to-agent transfer study of additive steering, coupling behavioral measurement with a representation read-out in a matched-information design: the same items rendered as plain chat or as a ReAct tool-use episode, with matched-norm random-direction controls and the transcript re-encoded every turn to exclude KV-cache contamination. Transfer is real but rescaled, and the right description is a dissociation: the injected direction reaches the late layers at near-full strength in every setting and model tested (install-site agent-over-chat ratios 0.83-1.16 across three families), while the behavioral coupling is reset per model and context. On Qwen2.5-7B a refusal bypass vector amplifies in the agent (T = 1.45, CI [1.20, 1.78], N = 300); across a powered uniform-protocol distribution the coupling spans amplification (Gemma-2-9B T = 2.00) to attenuation (Yi-1.5-9B T = 0.43, CI [0.29, 0.60]), with no universal constant and a single clean attenuator against a universal sign. Directional ablation of the same axis does not amplify (T = 0.93, CI including 1) while additive injection amplifies (T = 1.50), a 20.1-point gain difference (CI [13.4, 26.8]) that identifies an additive-specific mechanism. Two pre-registered instruments converge to localize the rescaling to the ReAct format scaffold, before any tool observation, rather than to the observation boundary where a dilution account would predict it. The safety implication is immediate and unpredictable: agentic deployment amplifies steering-based refusal bypass by up to 2.00x on some models while others attenuate, so a deployment cannot assume a given model is safe under additive steering.

Figures

Figures reproduced from arXiv: 2607.09156 by Lucas Pinto.

Figure 1
Figure 1. Figure 1: Setup schematic. A chat-extracted direction [PITH_FULL_IMAGE:figures/full_fig_p004_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: The two-sided per-model coupling distribution [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Frame-priming localization results. Left: Nested input-frame ablation (Instrument 1). Each point is the marginal coupling step from adding one frame ingredient; the ReAct format scaffold carries 104% of the endpoint gain; all other steps are within the random band. Right: Forward-pass phase restriction (Instrument 2). Pre-observation block injection reproduces 0.88 of the full agent gain; post-observation … view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

14 extracted references · 13 linked inside Pith

  1. [1]

    RefusalinLanguageMod- els Is Mediated by a Single Direction.arXiv, 2406.11717

    Arditi, A.; Obeso, O.; Syed, A.; Paleka, D.; Panickssery, N.; Gurnee,W.;andNanda,N.2024. RefusalinLanguageMod- els Is Mediated by a Single Direction.arXiv, 2406.11717. Bigelow, E.; Wurgaft, D.; Wang, Y.; Goodman, N.; Ullman, T.;Tanaka,H.;andLubana,E.S.2025. BeliefDynamicsRe- veal the Dual Nature of In-Context Learning and Activation Steering.arXiv, 2511.0...

  2. [2]

    Chen,Y.;Siu,V.;Liu,Y.;Song,D.;andWang,C.2026

    PersonaVectors:MonitoringandControllingCharac- ter Traits in Language Models.arXiv, 2507.21509. Chen,Y.;Siu,V.;Liu,Y.;Song,D.;andWang,C.2026. Con- trollingToolUsewithHeading-SpecificActivationSteering. arXiv, 2607.05790. Cristofano, T

  3. [3]

    Deng,Y.2026.GEMS:GeometricConstraintsEnableMulti- Semantic Superposition in LLMs.arXiv, 2606.19946

    Universal Refusal Circuits Across LLMs: Cross-Model Transfer via Trajectory Replay and Concept-Basis Reconstruction.arXiv, 2601.16034. Deng,Y.2026.GEMS:GeometricConstraintsEnableMulti- Semantic Superposition in LLMs.arXiv, 2606.19946. Fomin, M.; David, E.; and LeVi, A

  4. [4]

    Galeone, C.; Ettorre, A.; Park, M.; Ettorre, G.; and Ligorio, D

    Internal-State Probes Read the Situation, Not the Action: Three Nega- tiveResultsforPre-ActionMisalignmentMonitoring.arXiv, 2606.30449. Galeone, C.; Ettorre, A.; Park, M.; Ettorre, G.; and Ligorio, D

  5. [5]

    Steering in Language Models.arXiv, 2606.24952

    Perfect Detection, Failed Control: The Geome- try of Knowing vs. Steering in Language Models.arXiv, 2606.24952. Google DeepMind

  6. [6]

    Kang, D.; Liu, Z.; Ma, N.; Huang, Y.; Tan, Z.; and Jiang, M

    Gemma 2: Improving Open Lan- guage Models at a Practical Size.arXiv, 2408.00118. Kang, D.; Liu, Z.; Ma, N.; Huang, Y.; Tan, Z.; and Jiang, M

  7. [7]

    Kumar,A.;andMaple,C.2026

    Prompt-Activation Duality: Improving Activa- tion Steering via Attention-Level Interventions.arXiv, 2605.10664. Kumar,A.;andMaple,C.2026. RefusedinChat,Writtenin Code:Workflow-LevelJailbreakConstructioninIDECoding Agents.arXiv, 2607.03968. Lermen,S.;Dziemian,M.;andPimpale,G.2024. Applying Refusal-Vector Ablation to Llama 3.1 70B Agents.arXiv, 2410.10871. ...

  8. [8]

    Moskvoretskii, V.; Glandorf, D.; Medina Moreira, J.; Käser, T.;andWest,R.2026.TracingPersonaVectorsthroughLLM Pretraining.arXiv, 2605.13329

    The Llama 3 Herd of Models.arXiv, 2407.21783. Moskvoretskii, V.; Glandorf, D.; Medina Moreira, J.; Käser, T.;andWest,R.2026.TracingPersonaVectorsthroughLLM Pretraining.arXiv, 2605.13329. Nguyen,T.;Nguyen,T.A.;Alemohammad,S.;andBaraniuk, R. G

  9. [9]

    Panickssery,N.;Gabrieli,N.;Schulz,J.;Tong,M.;Hubinger, E.;andTurner,A.M.2024

    Minimizing Collateral Damage in Activation Steering.arXiv, 2605.01167. Panickssery,N.;Gabrieli,N.;Schulz,J.;Tong,M.;Hubinger, E.;andTurner,A.M.2024. SteeringLlama2viaContrastive Activation Addition.arXiv, 2312.06681. Tan, D.; Chanin, D.; Lynch, A.; Kanoulas, D.; Paige, B.; Garriga-Alonso, A.; and Kirk, R

  10. [10]

    Analyzing the Generalization and Reliability of Steering Vectors.arXiv, 2407.12404. Team, Q

  11. [11]

    Turner, A

    Qwen2.5 Technical Report.arXiv, 2412.15115. Turner, A. M.; Thiergart, L.; Leech, G.; Udell, D.; Vazquez, J.J.;Mini,U.;andMacDiarmid,M.2023.SteeringLanguage Models With Activation Engineering.arXiv, 2308.10248. Walsh, C.; and Barkett, E

  12. [12]

    arXiv, 2605.25151

    Representation Without Control:TestingtheRealizationEffectinLanguageModels. arXiv, 2605.25151. Yap,J.Q.2026.BehavioralSteeringina35BMoELanguage Model via SAE-Decoded Probe Vectors: One Agency Axis, Not Five Traits.arXiv, 2603.16335. Zhong, V.; and Li, Q

  13. [13]

    Zou, A.; Phan, L.; Chen, S.; Campbell, J.; Guo, P.; Ren, R.; Turner, A

    Refusal Lives Downstream of Persona in Chat Models.ICML 2026 Mechanistic Inter- pretability Workshop / arXiv, 2606.26161. Zou, A.; Phan, L.; Chen, S.; Campbell, J.; Guo, P.; Ren, R.; Turner, A. M.; Robey, B.; Kolter, Z.; Fredrikson, M.; and Hendrycks, D

  14. [14]

    Gate-fail

    Representation Engineering: A Top- Down Approach to AI Transparency.arXiv, 2310.01405. Preregistrations and Reproducibility Every behavioral experiment was pre-registered before pilot launch, with the live/die crite- ria and the verdict-mapping engine committed to the public repository before any data were collected. The registration documents indocs/incl...