{"id":"d0c9b74c-02f5-4c1f-b153-8d0faa97e020","arxiv_id":"2608.12892","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Low-dose causal responses at strength 0.1 predict later selective steering outcomes at strengths 0.25 and 0.5 better than static localization features, enabling abstention-aware coefficient selection.","lead":"This paper introduces a way to predict whether an activation-steering direction in a language model will work selectively before applying the full intervention. It shows that a small, cheap low-dose steering measurement is far more informative than static localization features for forecasting later outcomes, and that a risk-aware policy can choose a coefficient or abstain.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The forecast claim holds on its own margin-level terms, but the paper's practical 'selective intervention' framing relies on an untested generation-level transfer that its own S7 indicates is absent.","rationale":"I read the paper as a carefully scoped empirical study whose central claim is about margin-level outcome forecasting. The internal support is strong: the strength-disjoint weak response, grouped record/dataset/domain splits, record-paired bootstrap intervals, cross-model confirmation, and the layer-selection exclusion all support the conclusion that a low-dose causal response is substantially more informative than static localization geometry for forecasting later margin-level path outcomes. Static localization and AGOP geometry add little once the weak response is observed, and the paper does not overclaim on that point. The genuine soft spot is the generation-level transfer: the motivating language of 'selective intervention paths' and 'risk-aware intervention decision' extends beyond the measured teacher-forced margins. The paper is transparent about this limitation, explicitly stating in Section 5.5 and Appendix S7 that free-generation stress tests do not establish reliable wrong-to-right control. That transparency does not remove the load-bearing nature of the surrogate for the practical claim; it simply means the concern is disclosed rather than hidden. The reader's conditional verdict is appropriate: the margin-level forecast claim stands, but the selective-intervention framing should be evidence-bound or softened. My proposed test is a direct replay of the actual PML policy at the generation level, which would settle whether the surrogate concern is fatal to the practical framing or merely a scope clarification.","tokens_in":20255,"tokens_out":13916,"duration_ms":149560,"concrete_test":"Re-run Appendix S7 exactly with the PML M+A+R selector: on held-out records, let the policy choose a coefficient or abstain for suppression and enhancement, then decode free-form responses and measure wrong-to-right correction, semantic-neighbor preservation, and capability scores at the selected coefficient versus no intervention and versus the train-tuned fixed policy. Also recompute Table S18 restricted to the coefficients PML actually selects, with abstention counted as no intervention. If selected-coefficient generations show no reliable target correction or collateral benefit, the paper's practical framing should be reduced to margin-level forecasting; if they do show generation-level improvement, the surrogate concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central result is that a low-dose causal response forecasts later selective behavior. The paper defines 'selective behavior' at the level of teacher-forced answer-margin movements ΔT/ΔN/ΔC and explicitly disclaims reliable free-generation control (Section 5.5). This scope is internally consistent, so the forecast claim itself is not false. The load-bearing step is external: the abstract, introduction, and contribution statements promise a 'risk-aware intervention decision' and 'selective intervention paths.' Those practical conclusions require margin movements to transfer to decoded output. Appendix S7 (Table S18) tests exactly this transfer: on Qwen3-1.7B, learned directions yield 0–1 correctness gains and 0–1 damages across 24–28 paths, with 25–37% text changes; on Qwen3.5, gains (2–3) are balanced or exceeded by damages (2–4). Thus the model's margin moves without reliable wrong-to-right correction. If the intended contribution is margin-level forecasting, the paper is sound; if it is selective intervention, the surrogate is load-bearing and unsupported. The reader's weakest assumption correctly identifies this, and the manuscript's own S7 supplies the relevant evidence.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces Predictive Memory Localization (PML), a framework that treats the measured coefficient-grid response of an activation-steering direction as a sequence of random-calibrated events (target, semantic-neighbor damage, capability damage, and clean). PML combines baseline margins, metadata, static localization features, supervised geometry, and a strength-disjoint low-dose response at |α|=0.1 to forecast later outcomes at |α|∈{0.25,0.5}. The study covers 3,000 records from nine datasets on Qwen3-1.7B plus 500-record residual-norm-matched confirmations on three models. The main findings are that learned directions modestly improve target and clean-path incidence over random controls; that the low-dose causal response is a substantially stronger predictor of later path outcomes than static localization; and that a held-out selector improves utility relative to a train-tuned fixed-strength policy while reducing neighbor damage and avoiding most dense-scan evaluations. The paper explicitly scopes its primary evidence to teacher-forced answer-margin outcomes and reports free-generation stress tests in the supplement.","tokens_in":20450,"tokens_out":14783,"duration_ms":145709,"significance":"If the results hold, the paper makes a useful empirical contribution by quantifying the relative predictive value of static localization versus an inexpensive causal probe for intervention-path outcomes. The strength-disjoint design, frozen protocol, grouped train/test splits, record-paired bootstrap intervals, audit sensitivity analysis, and residual-norm-matched cross-model confirmations are careful and are genuine strengths. The finding that a low-dose causal response dominates static geometry as a forecast of later margin-level outcomes is clearly supported by the reported AUROC/AP comparisons. The paper is honest about its margin-level scope and about the limitations of the free-generation endpoint, which should be credited. The main significance is as a benchmark and a cautionary result for the localization-to-control hypothesis, rather than as a new mechanism or theory of localization.","major_comments":[{"comment":"The claim that 'the residual-norm-matched confirmations preserve the same qualitative pattern at model-dependent magnitudes' is contradicted by the Qwen3-1.7B 500-record subset at layer 11. In that panel, the mean-difference Clean delta is -0.6 percentage points and the RFM/AGOP Clean delta is -0.4 percentage points, with Target deltas of only +0.4 and +0.2 points, whereas the primary 3,000-record study at the same layer reports Clean deltas of +2.3 and +1.5 points for these families. As written, the cross-model confirmation claim overstates the support for 'learned directions improve target leverage and clean-path incidence over random controls.' The authors should report bootstrap intervals for the subset deltas, restrict the claims to the shallower block, or explicitly discuss the layer-11 discrepancy and its sampling variability.","section":"§5.2, Table 1 (panel B, layer 11)"},{"comment":"The abstract and introduction present PML as providing 'selective intervention paths' and a 'risk-aware intervention decision,' yet the operationalized endpoint is teacher-forced answer-margin utility (Eq. 4), and the paper's own free-generation stress tests (Table S18) show that learned directions do not produce reliable wrong-to-right correction (Qwen3-1.7B: 0–1 gains and 0–1 damages across 24–28 paths; Qwen3.5: gains 2–3 are balanced or exceeded by damages 2–4). Since the manuscript explicitly disclaims reliable free-generation control, the contribution framing should be consistently and clearly limited to margin-level outcome forecasting and margin-utility-based coefficient selection, rather than suggesting generation-level selective control. This is a scope-and-interpretation issue rather than a technical error in the forecasting experiments, but it affects how the contribution will be read and cited.","section":"§5.5, Table S18, Abstract"}],"minor_comments":[{"comment":"The decision score combines predicted utility with a 0.1 clean-bonus weight and a 0.01 magnitude penalty; these weights are declared rather than derived. The weight-sensitivity analysis in Table S17 is reassuring, but a sentence in the main text explaining the default values would improve readability.","section":"Eq. (5)"},{"comment":"The sentence 'The logged matched_norm_random entry is numerically identical and retained only for auditability' is unclear; please specify whether this entry is a duplicate row in the released logs and how it differs from the ordinary random control.","section":"§4.1"},{"comment":"The 'Scalar' control is described as training-free, but it would help to state explicitly that this is the signed |α|=0.1 response without any scaling by the strength ratio, so that the comparison with the multivariate R-only model is unambiguous.","section":"Table 3"},{"comment":"The main text reports that the selector improves utility 'relative to a train-tuned fixed coefficient' but does not state the fixed coefficient's value or sign for suppression/enhancement; Table S14 gives this information, so a cross-reference would help the reader interpret the magnitude of the gains.","section":"§5.4, Table 4"}],"recommendation":"major_revision","confidential_remarks":"The paper is a carefully executed empirical study, and the central margin-level forecast claim is defensible. The two major comments concern overreach in the cross-model confirmation claims and in the selective-intervention framing. Both are addressable within the manuscript's scope; no fundamental flaw in the forecasting experiments is apparent. I suggest requesting a revision that tightens these claims and adds the needed qualifications, after which the paper could be suitable for publication."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Things you should know: this paper actually delivers on its margin-level forecast claim, but its own appendix undercuts the broader 'selective intervention' framing. Worth reading for the evaluation design, not for a practical control tool.\n\nThe new thing here is the measured-grid path as the predictive object and the strength-disjoint low-dose response at |α|=0.1 as the dominant predictor of later outcomes at |α|=0.25/0.5. Static localization and AGOP geometry add almost nothing once the weak response is observed. That's a genuinely useful negative result for the localization community, and it's backed by a careful protocol: random-calibrated thresholds, record/dataset/domain grouped splits, bootstrap intervals, and replication on three norm-matched models. The AUROC range (0.80–0.85) is solid for this kind of sparse multi-label prediction. The paper is also unusually honest—Section 5.5 and the Limitations section state plainly that free-generation wrong-to-right correction is not reliable.\n\nSoft spots, in order of weight. First, the practical framing overreaches. The abstract and introduction promise 'risk-aware intervention decisions' and 'selective intervention paths,' but the endpoint is teacher-forced answer-margin movement. Appendix S7 tests transfer to decoded text and finds it largely absent: on Qwen3-1.7B, learned directions produce 0–1 correctness gains and 0–1 damages across 24–28 paths; on Qwen3.5, gains are balanced or exceeded by damages. The authors disclose this, but it should be in the main text as the boundary, not relegated to the supplement. The forecast claim itself is fine—it's margin-level and falsifiable—but the intervention language oversells it.\n\nSecond, there is no empirical comparison to the closest related method, SteerBoost (Fan et al. 2026). The paper distinguishes itself in text, but given SteerBoost also predicts steerability, a head-to-head on the same paths would have made the contribution crisper.\n\nThird, no code or artifacts are actually linked. 'Released protocol' and 'released artifacts' are mentioned, but no URL. For a paper whose contribution is largely methodological, that's a real gap.\n\nMinor: the target probes are semantically close to the direction-fitting statements (surface-disjoint but near paraphrases), so the 'selective' separation between target and neighbor/capability deserves a few sentences of discussion. The utility weights are hand-declared; the sensitivity table helps, but that's still a free choice.\n\nNet: the central measurement claim holds up, and the evaluation design is careful. A serious referee should engage; the paper will need revision to align the framing with the margin-level evidence, add the SteerBoost comparison, and ship code. I'd accept it conditional on those changes.","headline":"Solid within-subfield evaluation study: the margin-level forecast claim holds and is well tested, but the selective-intervention framing needs to be scaled back to what S7 shows.","tokens_in":21019,"tokens_out":2486,"would_cite":true,"duration_ms":22498,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"One small probe forecasts later steering outcomes best","keywords":["Predictive Memory Localization","activation steering","low-dose causal response","selective intervention","RFM/AGOP direction","random calibration","measured-grid intervention path","risk-aware strength selection"],"falsifier":"Run the same path-prediction protocol with free-generation correctness as the outcome: decode completions at $|\\alpha|\\in\\{0.25,0.5\\}$ on the 3,000 records and label wrong-to-right target flips and neighbor damage from the decoded text. If the low-dose margin response predicts these generation-level flips with AUROC near 0.5, or if no policy achieves more corrections than random directions, the surrogate assumption behind PML's practical claims is falsified.","tokens_in":20012,"feed_emoji":"🎯","tokens_out":11085,"duration_ms":94822,"temperature":0.7,"pith_summary":"This paper is trying to establish that the most useful thing to look at when predicting whether an activation-steering intervention will work is a cheap, small causal response, not a static map of where the memory is localized. It introduces Predictive Memory Localization (PML), which treats the whole measured intervention path—target effect, semantic-neighbor damage, capability damage, and clean strengths—as the object to forecast, and uses a response measured at a tiny strength ($|\\alpha|=0.1$) to predict outcomes at disjoint stronger strengths ($|\\alpha|\\in\\{0.25,0.5\\}$). Across 3,000 frozen records from nine datasets and fourteen domains, static localization and supervised geometry add almost no predictive power, while the low-dose response produces the dominant gain and replicates across three residual-norm-matched base models with record-held-out macro AUROC of 0.801–0.828. If right, this shifts evaluation practice: steering researchers should measure a small cheap perturbation before betting on a coefficient, rather than trusting localization geometry alone.","feed_headline":"One small probe forecasts later steering outcomes best","feed_subtitle":"A cheap low-dose causal measurement, not static geometry, tells which steering coefficients are safe and useful.","key_machinery":"The load-bearing object is the measured-grid intervention path: a record-specific unit direction injected at a selected layer at signed strengths $\\alpha\\in\\{-0.5,-0.25,-0.1,0,0.1,0.25,0.5\\}$, with target, semantic-neighbor, and capability answer-margin changes calibrated against random-direction 95th-percentile thresholds. Path events at held-out strengths $|\\alpha|\\in\\{0.25,0.5\\}$ are the prediction labels, while the low-dose response features are measured only at $|\\alpha|=0.1$, so probe and label strengths are disjoint. Direction families include the mean-difference activation-addition vector, linear and logistic probe weights, and the RFM/AGOP leading eigenvector of the average gradient outer product; static localization and geometry features are the competing predictive evidence. The policy layer scores each candidate coefficient with predicted utility $\\hat u_i(\\alpha)+0.1\\,\\hat p_i(\\text{clean}|\\alpha)-0.01|\\alpha|$ and abstains when the best score is nonpositive.","core_discovery":"The central result is stated directly: a small causal response is the most useful forecast of later selective behavior. Concretely, the paper defines a random-calibrated measured-grid path for each record, direction, and layer, recording whether an intervention at each coefficient crosses thresholds for target leverage, semantic-neighbor damage, capability damage, and clean operation. Learned directions such as the mean-difference vector, logistic probe, and the RFM/AGOP direction raise target-any and clean-any path incidence relative to random (at layer 7, RFM/AGOP reaches 13.1% target-any and 12.3% clean-any versus 9.5% and 8.9% for random), while collateral-damage differences remain statistically unresolved. When these paths serve as prediction labels, dropping static localization changes AUROC by only about +0.002 on average, whereas adding the strength-disjoint low-dose response changes it by about +0.177; the same hierarchy holds under record-, dataset-, and domain-grouped splits and across three residual-norm-matched base models. The paper therefore claims that localization becomes actionable only when paired with a cheap causal measurement of the specific path, and that a held-out policy using those forecasts can select a coefficient or abstain, improving utility and reducing neighbor damage relative to a fixed-strength policy.","pith_inferences":["This extends the paper's margin-level result: because the low-dose response is cheap and strength-disjoint, one could build an online adaptive controller that re-measures the weak response per prompt and chooses or abstains from an intervention in real time.","The margin-to-generation gap suggests a testable refinement: use the low-dose predictor as a filter that sends only high-confidence paths to expensive free-generation evaluation, potentially converting margin-level forecasts into generation-level control.","The same hierarchy—static localization is a weak prior, a small causal probe is a strong one—may apply to knowledge editing and unlearning; the paper's own transfer test to parameter editing implies activation controllability does not equal editability, so a cheap causal diagnostic could screen editability too.","A cross-model caveat follows from the paper's own numbers: learned-vs-random damage differences are unresolved, so deployment should claim reduced damage at the selected coefficient, as the policy does, rather than claiming directions are intrinsically safer."],"forward_implications":["Activation steering evaluation should adopt a cheap low-dose probe as a diagnostic: a measurement at $|\\alpha|=0.1$ forecasts later target and damage outcomes better than any static localization feature tested.","A fixed intervention strength is the wrong default: a predictor-driven selector that can abstain improves utility and reduces semantic-neighbor damage relative to the best train-tuned fixed coefficient, while evaluating about 90% fewer coefficients than a dense scan.","Learned directions create more usable intervention paths without generally reducing collateral movement: target and clean incidence rise, but damage differences from random are not statistically resolved, so selectivity must be checked per path rather than assumed from how the direction was built.","The diagnostic hierarchy transfers across base models when intervention budgets are aligned by residual norm, indicating the low-dose dominance is not an artifact of one architecture's scale.","PML's forecasts are margin-level: the paper explicitly keeps free-generation correction out of its primary claims, so the forecasting result and the generation-control result should be read separately."],"supporting_citations":[{"why":"Supplies the RFM/AGOP direction and spectral features used as the supervised geometry signal.","marker":"Radhakrishnan et al. 2022"},{"why":"Defines activation steering by injecting directions; the mean-difference direction family builds on this activation-addition baseline.","marker":"Turner et al. 2023"},{"why":"Contrastive activation addition direction family used as one of the learned interventions.","marker":"Rimsky et al. 2024"},{"why":"Representation engineering contrast directions motivate the supervised direction constructions.","marker":"Zou et al. 2023"},{"why":"Prior evidence that localization does not inform editing; PML's boundary claim about controllability versus editability rests on this contrast.","marker":"Hase et al. 2023"},{"why":"Closest prior work predicting steerability from first-token dynamics; PML distinguishes its path-level, strength-disjoint design from it.","marker":"Fan et al. 2026"},{"why":"Locating factual associations in transformer feed-forward layers motivates the memory-localization framing.","marker":"Meng et al. 2022"},{"why":"Contributes MMLU-Pro records to the frozen 3,000-record benchmark used for the empirical claims.","marker":"Wang et al. 2024b"},{"why":"Contributes MMLU-Redux 2.0 records and the answer-margin evaluation format used in the benchmark.","marker":"Gema et al. 2024"}],"fun_headline_variants":["Low-dose causal probe tops static geometry for steering prediction","Tiny causal signal forecasts selective steering outcomes best","Small causal response beats geometry in forecasting steering selectivity","Cheap causal probe predicts safe and useful steering coefficients","Strength-disjoint causal response best forecasts intervention success"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The paper's forecasts, utility scores, and policy decisions are all measured on teacher-forced answer-margin movements; if those margin movements do not translate into reliable wrong-to-right changes in freely generated text, the framework's practical value as a selective intervention tool is reduced, even though the margin-level forecasts themselves may remain accurate.","fun_headline_variants_meta":{"raw":{"variants":["Low-dose causal probe tops static geometry for steering prediction","Tiny causal signal forecasts selective steering outcomes best","Small causal response beats geometry in forecasting steering selectivity","Cheap causal probe predicts safe and useful steering coefficients","Strength-disjoint causal response best forecasts intervention success"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000276,"raw_usage":{"total_tokens":1725,"prompt_tokens":1104,"completion_tokens":621,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":720,"completion_tokens_details":{"reasoning_tokens":548}},"tokens_in":720,"tokens_out":621,"duration_ms":6342,"temperature":1.0,"reasoning_tokens":548,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T21:06:58.745890+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same path-prediction protocol with free-generation correctness as the outcome: decode completions at $|\\alpha|\\in\\{0.25,0.5\\}$ on the 3,000 records and label wrong-to-right target flips and neighbor damage from the decoded text. If the low-dose margin response predicts these generation-level flips with AUROC near 0.5, or if no policy achieves more corrections than random directions, the surrogate assumption behind PML's practical claims is falsified.","supporting_citations":[{"cited_title":"Advances in Neural Information Processing Systems , volume =","cited_arxiv_id":null,"evidence_quote":"Prior evidence that localization does not inform editing; PML's boundary claim about controllability versus editability rests on this contrast."},{"cited_title":"Advances in Neural Information Processing Systems , volume =","cited_arxiv_id":null,"evidence_quote":"Locating factual associations in transformer feed-forward layers motivates the memory-localization framing."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Contributes MMLU-Redux 2.0 records and the answer-margin evaluation format used in the benchmark."}],"review_version":1}