{"id":"a9fdb9ed-7e7c-42b0-aae7-80c09eea381f","arxiv_id":"2608.11735","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"A cue-induced activation contrast at selected layers tracks implicit personalization in LLM recommendations and can be ablated to suppress it, with effects that are model- and attribute-specific.","lead":"Using matched cued and neutral conversations in five instruction-tuned LLMs, this paper shows that the magnitude of the activation difference between cued and neutral contexts tracks how much a model's recommendations shift (up to r=0.87), and that removing this internal direction can suppress the shift, often better than prompting. It also finds that when two demographic cues co-occur, the internal signals combine linearly while the output changes are sub-additive.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Topic confound: cue-minus-neutral ΔA conflates demographic association with topical/lexical content, so the tracked and ablated 'demographic direction' may be a topic direction; the paper concedes this but provides no topic-matched control.","rationale":"The reader identified the cue-vs-neutral construct as the weakest assumption; my read concurs, and I judge it the single most load-bearing point because it gates the interpretation of both halves of the central claim. The magnitude-tracking result and the ablation result are internally consistent and, to the paper's credit, extensively controlled: sham directions (including a magnitude-matched sham, Appendix G) rule out generic perturbation; the White-cue universal null provides a falsifiable anchor showing that raw topical difference alone does not move behavior; the vector-source robustness (dim/intersect/explicit, Table 11) shows the ablation is not an artifact of one contrast definition; and the CAR metric is human-validated (80% pooled agreement) and scores stereotype-aligned content by construction, so the behavioral DV is well-defined. None of these, however, isolates the demographic component of the internal signal. The explicit-source robustness is the closest control — a demographically explicit context carries less cue-typical topical content — but the paper reports it as the most variable and weakest source in some cells (e.g., Mistral race×age explicit +0.022 vs. dim +0.154), so it mitigates without resolving. I also considered the reader's in-sample selection concern (L⋆ from patching and per-cell α⋆ chosen on the same data used to report correlations and reductions) and treat it as secondary: L⋆ is selected on patching NIE rather than on the magnitude–behavior correlation; the ablation effects are present, if smaller, at exact projection α=1 without grid search (Table 10: 11/18 SED targets significant); and the full grids are disclosed, making any inflation auditable. The topic confound, by contrast, is structural and untested — the paper's own Limitations text acknowledges it. A topic-matched filler condition is the decisive experiment: if correlations and CAR reductions survive topic-matching, the demographic interpretation stands; if they collapse, the paper's contribution re-scopes from 'implicit personalization' to topic-conditioned output shift. The verdict stays CONDITIONAL; the concern sharpens the condition rather than changing it.","tokens_in":27362,"tokens_out":16508,"duration_ms":172470,"concrete_test":"One experiment: rebuild the 9×50 matched triplets with topic-matched, register-matched neutral fillers — for each demographic cue, a filler matched for topic, register, and conversational function but stripped of demographic association (e.g., 'I'm headed to a Kwanzaa celebration this week' vs. 'I'm headed to a winter holiday celebration this weekend'; 'no cap' vs. generational-neutral slang of the same register). Then recompute the Table 1 per-sample correlations and the §4.3 ablation α-grid with target-CAR reductions, and measure the cosine similarity at L⋆ between the sample-averaged cue-condition ΔA direction and the topic-matched-filler ΔA direction.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim attaches a demographic meaning to an internal quantity: ΔA = A(cue) − A(neutral) is called the cue direction and used as an inference-time target for suppressing implicit personalization. But each cue phrase varies three things at once — demographic association, conversation topic, and lexical/register specificity — while the neutral fillers are topic-unmatched (e.g., 'I'm headed to a Kwanzaa celebration this week' vs. 'Been getting into knitting lately'). The paper concedes exactly this: the Limitations state that 'a demographic cue intrinsically carries topical content and lexical specificity' and, citing Neplenbroek et al. (2026), that 'conversation topic can dominate output shifts'; Section 5 hedges that tracking is attributed to the cue 'though not to demographic content specifically rather than the topical content correlated with it.' Both legs of the central claim are therefore at risk: (1) the per-sample correlation (r up to 0.87) may track topic-driven output movement rather than demographic representation, and (2) the ablated direction v̂ may remove a topic axis whose suppression merely removes the topic that activates stereotype-tagged recommendations, leaving the 'demographic direction' mis-specified. Genuine mitigating evidence exists — the White-cue universal null shows raw topical difference alone does not shift behavior, and the vk=explicit source robustness (Table 11) shows a demographically explicit contrast also suppresses targets — but neither closes the gap. No topic-matched control was run, and the explicit source is itself 'the most variable across cells and the weakest on some' (Appendix H.2, e.g., Mistral race×age explicit +0.022 vs. dim +0.154). The load-bearing assumption is therefore untested, and the authors' own hedge means the demographic interpretation is unsupported as stated.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies implicit personalization in LLMs: user utterances that never name a demographic nevertheless shift recommendation outputs toward stereotype-aligned content. Using matched cued/neutral multi-turn dialogues across five 7B–14B instruction-tuned models, the authors define a raw residual-stream activation contrast ΔA between cued and neutral contexts and report three findings: (1) the normalized magnitude of ΔA at patching-selected layers correlates per-sample with behavioral shift, up to r=0.87; (2) in mixed-cue contexts, the internal contrast is substantially reconstructible as a linear combination of single-cue contrasts, while the output shift is sub-additive; and (3) projecting out a sample-averaged cue direction during generation suppresses target-dimension stereotype content, often more than an explicit \"ignore demographics\" prompt, with only small MMLU degradation. The paper positions these results as establishing an internal, causally manipulable signal underlying implicit personalization.","tokens_in":27541,"tokens_out":3731,"duration_ms":45986,"significance":"If the central claims hold, the paper would make a useful contribution: it connects a well-documented behavioral phenomenon to a measurable and intervenable internal quantity, and it provides a multi-cue, dose-matched experimental design that is more careful than much of the existing activation-steering literature. Strengths worth emphasizing are the matched-dialogue construction, dose-matched factorial controls, random-direction and magnitude-matched shamming, pairwise negative controls for the composition analysis, BH correction, and the vector-source robustness check in Appendix H.2. The paper is also unusually candid about its limitations, including the topic confound and the model-specificity of selective ablation. However, two methodological reuse issues—the in-sample selection of L⋆ for the correlation analysis and the in-sample estimation of the ablated direction—bear directly on the strength and generality of the headline claims, and the topic-confound issue affects the demographic interpretation of the internal direction. For these reasons the paper needs substantial revision before its central claims can be taken at face value.","major_comments":[{"comment":"The per-sample correlations in Table 1 are computed on the same n=50 triplets used to select the layer set L⋆: Appendix C.4 ranks layers by NIE averaged over the retained samples, and §4.1 then averages sΔA over those selected layers and correlates with behavior on the same samples. Because the NIE ranking is a causal measure of how much each layer mediates the cued-vs-neutral output divergence, selecting layers on the evaluation samples can inflate the reported correlations even though the ranking uses next-token JSD rather than SED-desc or CAR. The magnitude of this inflation is not quantified. The authors should report a nested or split-sample procedure (e.g., selecting L⋆ on one half of the triplets and correlating on the other, or a leave-one-out scheme) and show that the headline r values survive. This is load-bearing because the first contribution is precisely the claim that the internal magnitude tracks the behavioral shift.","section":"§4.1 and Appendix C"},{"comment":"The ablated direction v̂dim is computed by averaging the single-cue activation contrasts over the same n=50 responses on which the ablation is then evaluated (Tables 2, 3, 8, 9, and Figure 4). This is an in-sample estimate of the direction, so the demonstration does not establish the stated \"inference-time target\" property: a user arriving with a new conversation would not have the sample-averaged direction available unless it was estimated on a separate development set. The vector-source robustness in Appendix H.2 is helpful, but the explicit and intersect sources are also evaluated on the same samples. The authors should either split the data for direction estimation vs. evaluation or provide a clear statement that the reported effect sizes are in-sample and may overstate out-of-sample suppression.","section":"§4.3, Eq. (2)"},{"comment":"The cue-minus-neutral contrast ΔA conflates demographic association with topical content and lexical specificity. The neutral fillers are not topic-matched to the cue phrases, and the paper itself concedes in the Limitations that \"a demographic cue intrinsically carries topical content and lexical specificity\" and, citing Neplenbroek et al. (2026), that \"conversation topic can dominate output shifts in real conversations.\" Section 5 also hedges that tracking is attributed to the cue \"though not to demographic content specifically rather than the topical content correlated with it.\" This is not merely a presentational caveat: the paper's central framing names the ablated direction a \"demographic direction,\" and the title and abstract promise control of \"implicit personalization.\" Without a topic-matched control (e.g., neutral fillers matched for topic and register while varying demographic association, or a stratified analysis across cue topics), the tracked and ablated direction could be a topic/lexical axis whose suppression removes stereotype-aligned content by removing the topic that activates it. The White-cue null and the vk=explicit robustness partially mitigate this concern, but they do not establish that the internal quantity is specifically demographic rather than topical. This issue should be addressed with additional controls or with a substantially narrowed claim.","section":"§3.4, §5, Limitations"}],"minor_comments":[{"comment":"The human-annotation agreement of 63% on age, with all disagreements at the adult/neutral boundary, is lower than the pooled 80% and is worth stating in the main text near the CAR definition, since the age condition is used in several headline results.","section":"Appendix B"},{"comment":"The caption says the dashed line represents the \"original random-direction sham,\" but the text also discusses a matched-norm sham; please clarify which sham is shown and how the two shams differ.","section":"Figure 4"},{"comment":"The significance test behind the per-condition counts uses uncorrected p<.05; since the counts are used as a descriptive landscape rather than a corrected claim, this is acceptable, but the lack of correction should be stated next to the counts in §4.1.","section":"Appendix D"},{"comment":"The two behavioral metrics are described as complementary, but Table 1 shows several cells where SED and CAR correlations disagree sharply (e.g., Mistral male cue: SED r=0.37, CAR r=-0.16). A sentence in §4.1 explaining how such divergences are interpreted would help the reader avoid over-reliance on any single cell.","section":"§3.3"},{"comment":"The use of three different layer-set sizes (top-10 for correlation, pairwise union for decomposition, top-2 for ablation) is reasonable, but the main text should remind the reader that the same NIE ranking underlies all three choices, so the ablation and the correlation are not independent layer selections.","section":"Appendix C.5"}],"recommendation":"major_revision","confidential_remarks":"The paper's empirical scaffolding is strong and unusually honest, but the in-sample layer selection and in-sample direction estimation are exactly the kind of details that a careful reader will seize on, and the topic-confound caveat sits in direct tension with the demographic framing of the title and abstract. I would advise the editor that the revisions requested in the major comments are feasible within the manuscript's scope: a split-sample correlation analysis, a held-out direction evaluation, and either a topic-matched control or a recalibrated claim would materially change the strength of the contribution."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThis paper is worth your time. It does something genuinely new: it links the well-documented behavioral bias of implicit personalization to a localized internal activation signal and shows that signal can be used for causal control. The empirical design is thoughtful—matched cued/neutral dialogues, dose-matched controls, random-direction and magnitude-matched shams, BH correction, and pairwise negative controls. The per-sample correlations (up to r=0.87) and the sub-additive composition finding are real contributions, and the paper is honest about model-specificity and the limits of selectivity.\n\nThe main soft spot is not the topic confound, though that is real; it is that the evaluation is in-sample. L* is selected by activation patching on the same matched triplets used for the correlations, and the ablation strength alpha is chosen per cell on the same data that reports the effect. Reported effect sizes are therefore likely to overstate out-of-sample performance. This is fixable with cross-validation or held-out cue sets, and it does not break the core finding, but the numbers should be treated as upper bounds.\n\nThe topic confound is acknowledged in the Limitations: a demographic cue carries topical content and lexical specificity, and topic can dominate output shifts. The paper's own hedge—that tracking is attributed to the cue 'though not to demographic content specifically'—is the right reading. The White-cue null and the explicit-source robustness (Table 11) mitigate, but no topic-matched control was run, so the demographic interpretation is not fully supported. The ablation target is better described as a cue direction than a demographic direction.\n\nTwo smaller points. The abstract's 'often more effectively than prompting' is stronger than the CAR evidence, where many cells are null or reversed; a softer aggregate claim would be more accurate. And withholding the cue templates blocks exact reproduction; understandable ethically, but it raises the bar for independent verification.\n\nOverall: this is a serious, carefully controlled study whose central claim holds up in conditional form. It deserves a real referee. I'd send it to review and ask for out-of-sample validation, a topic-matched control or a re-framed claim, and a more measured abstract. It is the kind of paper I would bring to reading group.","headline":"A careful, honest empirical study that plausibly links implicit personalization to a localizable internal signal, but the demographic interpretation is not fully supported and the in-sample evaluation likely overstates effect sizes.","tokens_in":28260,"tokens_out":2124,"would_cite":true,"duration_ms":22539,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that implicit personalization in LLMs is tracked and controllable through a localized activation signal, with per-sample correlations up to $r=0.87$.","keywords":["implicit personalization","activation contrast","direction ablation","stereotype bias","LLM recommendation","intersectional cues","representation steering","causal mediation"],"falsifier":"Build a matched control in which each demographic cue phrase is replaced by a topic-equivalent but demographically neutral phrase (e.g., a Kwanzaa cue vs. a similarly specific secular holiday-planning cue), rerun the per-sample magnitude–shift correlation and the direction ablation, and check whether the correlation and suppression effect survive; if they collapse, the signal tracks topic, not demography.","tokens_in":27039,"feed_emoji":"🧠","tokens_out":8210,"duration_ms":78828,"temperature":0.7,"pith_summary":"Large language models quietly tailor recommendations to demographic signals users never state—a Kwanzaa mention, 'no cap', a question about playground games. This paper claims that implicit personalization is not just an output-level behavior: it has a localized internal signature. The size of the cue-minus-neutral activation contrast at the layers that mediate the shift tracks how much the recommendations move, with per-sample correlations up to $r=0.87$, and it falls toward zero where no shift occurs. The paper further claims that projecting the averaged cue direction out of the residual stream during generation suppresses the stereotype-tagged content, often beating an explicit 'ignore demographics' prompt, and that when two cues co-occur their internal signals combine nearly linearly while the behavioral shift stays sub-additive. If these claims hold, silent, invisible personalization becomes an inspectable and controllable mechanism rather than a black-box output effect.","feed_headline":"An internal LLM signal tracks hidden-cue bias up to r=0.87","feed_subtitle":"Erasing that signal reduces stereotype-tagged recommendations, often beating 'ignore demographics' prompts.","key_machinery":"The central object is the cue-induced activation contrast $\\Delta A^{(i)}_L = A_L(C_{\\text{cue}}^{(i)}, x) - A_L(C_{\\text{neutral}}^{(i)}, x)$: the difference in post-attention residual-stream activations at the final query-token position between a matched cued and neutral multi-turn dialogue. Its normalized magnitude $s^{(i)}_{\\Delta A} = \\|\\Delta A^{(i)}_L\\|_2 / \\|A_L(x)\\|_2$, averaged over the causally implicated layers $L^\\star$, is the detection signal that tracks the behavioral shift. Its sample-averaged unit direction $\\hat{v}_{\\text{dim}}$ is the intervention target, removed at inference time with a forward pre-hook projection $h' = (I - \\alpha \\hat{v}_{\\text{dim}}\\hat{v}_{\\text{dim}}^{\\top})h$. The layers $L^\\star$ are located by activation patching with a normalized indirect effect on the next-token Jensen–Shannon divergence between cued and neutral runs. The two behavioral metrics—semantic embedding distance on movie plot descriptions and the content alignment ratio of stereotype-tagged items—supply the per-sample targets that the internal signal is shown to track.","core_discovery":"The paper's central discovery is that the cue-induced activation contrast $\\Delta A = A_L(\\text{cued}) - A_L(\\text{neutral})$, read at the last query-token position of the residual stream at layers $L^\\star$ identified by activation patching, carries the information that drives implicit personalization, and does so in two usable forms. Its normalized magnitude $s_{\\Delta A}$ correlates per sample with the behavioral shift—up to $r=0.87$ on SED-desc and up to $0.85$ on stereotype-tagged content (CAR)—across five 7B–14B instruction-tuned LLMs, with the correlation scaling with effect size and collapsing in the universal null condition (white-cue). Its sample-averaged direction, when projected out of the residual stream during generation via $h' = (I - \\alpha \\hat{v}\\hat{v}^{\\top})h$, reduces stereotype-tagged recommendations on several model-by-dimension combinations, matching or exceeding an explicit 'ignore demographics' prompt (up to $9.2\\times$ on Qwen3-8B SED-desc) while keeping MMLU within 0.6 points at exact projection and at most 2.8 points at $\\alpha=5$. The paper also establishes a dissociation between internal geometry and external behavior: mixed-cue contrasts reconstruct well from the sample's own single-cue components (mean fit 0.576–0.704, vs 0.095 with a different sample's basis), but the behavioral shift is strictly sub-additive, 28–38% below the sum of single-cue shifts. Selectivity—removing one dimension while sparing a co-present one—holds on some models (Mistral gender→race, Llama race→gender) but not others (Qwen3-8B), so the paper frames suppression as more reliable than selectivity.","pith_inferences":["If the magnitude–shift correlation holds in live traffic, the same $\\Delta A$ measurement could be computed on real user conversations without matched neutrals by using a reference distribution of neutral continuations, turning the per-sample correlation into a monitoring statistic for silent personalization.","The sub-additive behavior with near-linear internals suggests a fixed-output-budget mechanism: with five recommendations, cues compete for slots, which predicts that lengthening the output list should reduce the 28–38% shortfall; this is testable by varying list length.","Where rank-one ablation fails to be selective (Qwen3-8B), nullspace-projection or closed-form linear erasure methods that remove a subspace rather than a single direction are the natural next test, and the paper's decomposition fits provide the subspace geometry needed to formulate it.","If the activation-contrast direction is truly demographic rather than topical, the same method could map which layers encode which demographic axis across models, and whether the universal null (white-cue) condition corresponds to a genuinely absent direction or to a direction that the model does not route through the mediating layers."],"forward_implications":["Implicit personalization can be detected before the model answers: the per-sample normalized contrast magnitude predicts how far recommendations will move, so the same signal could serve as a runtime audit of when a conversation is about to elicit stereotype drift.","Mixed cues do not behave in a single-axis way: because internal contrasts combine two directions while outputs compress sub-additively, evaluations that vary one demographic cue at a time will both overestimate (additive prediction) and misattribute the effect of co-occurring cues.","Prompt-free control is feasible in some settings: projection ablation suppresses targeted stereotype content where 'ignore demographics' prompts are flat or backfire, and at exact projection leaves general benchmark accuracy essentially unchanged.","The dissociation between linear internal composition and sub-additive behavior means output-level audits cannot reveal the representational state: two contexts that produce identical recommendation shifts can have different internal mixtures.","Direction-source robustness suggests the effect is not an artifact of one contrast definition: re-deriving the ablated direction from mixed-cue or explicitly-stated demographic contrasts reproduces the suppression, with cell-specific differences in strength."],"supporting_citations":[{"why":"Supplies the documented implicit personalization behavior and the cue taxonomy the paper's matched dialogues are built on.","marker":"Neplenbroek et al., 2025"},{"why":"Provides the concrete harm (dialect-conditioned decisions) that motivates connecting the behavior to an internal, controllable signal.","marker":"Hofmann et al., 2024"},{"why":"Defines the phenomenon of implicit personalization that the paper sets out to locate internally.","marker":"Jin et al., 2024"},{"why":"Provides the contrastive activation construction that the cue-minus-neutral contrast and steering direction adapt.","marker":"Zou et al., 2023"},{"why":"Supplies the single-direction mediation and projection-ablation approach used for the inference-time intervention.","marker":"Arditi et al., 2024"},{"why":"Establishes the causal mediation analysis framework used for activation patching to select the mediating layers $L^\\star$.","marker":"Vig et al., 2020"},{"why":"Provides the contrast magnitude formulation that the paper adapts as its per-sample detection signal.","marker":"Dherin et al., 2025"},{"why":"Shows that neutralizing attribute directions can outperform anti-bias prompts, the baseline the ablation results are compared against.","marker":"Karvonen and Marks, 2025"},{"why":"Cited in the Limitations to support the concern that conversation topic can dominate output shifts, which bounds the demographic interpretation of the cue contrast.","marker":"Neplenbroek et al., 2026"}],"fun_headline_variants":["LLM bias traced to a localized activation signal","Erasing one signal reduces LLM bias more than prompting","Implicit personalization maps to an internal signal, r=0.87","Surgical activation removal curbs stereotype-tagged output","Hidden cues' effect lives in a single internal signal"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the cue-minus-neutral activation contrast isolates demographic association rather than the topic, wording, or lexical register of the cue phrase—the paper's neutral fillers are not matched to the cues on those dimensions, so if topic continuation drives the shift, the 'demographic direction' is mis-specified.","fun_headline_variants_meta":{"raw":{"variants":["LLM bias traced to a localized activation signal","Erasing one signal reduces LLM bias more than prompting","Implicit personalization maps to an internal signal, r=0.87","Surgical activation removal curbs stereotype-tagged output","Hidden cues' effect lives in a single internal signal"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000761,"raw_usage":{"total_tokens":3455,"prompt_tokens":1095,"completion_tokens":2360,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":711,"completion_tokens_details":{"reasoning_tokens":2279}},"tokens_in":711,"tokens_out":2360,"duration_ms":17162,"temperature":1.0,"reasoning_tokens":2279,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T00:28:52.722589+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Build a matched control in which each demographic cue phrase is replaced by a topic-equivalent but demographically neutral phrase (e.g., a Kwanzaa cue vs. a similarly specific secular holiday-planning cue), rerun the per-sample magnitude–shift correlation and the direction ablation, and check whether the correlation and suppression effect survive; if they collapse, the signal tracks topic, not demography.","supporting_citations":[],"review_version":1}