{"id":"07993c89-6b1c-4acc-a32d-ddd6fcc00ccb","arxiv_id":"2606.11502","paper_version":3,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":5.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"Different persona induction methods produce a spectrum of belief internalization: prompting, ICL and SFT mainly alter outputs while Emergent Misalignment produces large representational shifts and Open Character Training produces smaller ones clearest in larger models.","lead":"The paper tests whether role-playing changes what language models internally treat as true or only what they output. If training methods can shift internal beliefs, this matters for AI systems given more autonomy.","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.3","headline":"Truth probes may reflect output or training artifacts rather than internal belief representations","rationale":"The reader’s weakest assumption matches the load-bearing measurement step exactly. Because the abstract (and the supplied text) gives no probe-validation experiments that isolate representational change from output or training artifacts, the concern stands and keeps the verdict at UNVERDICTED.","tokens_in":1672,"tokens_out":330,"duration_ms":11655,"concrete_test":"Take the trained truth probes and evaluate them on a held-out set of factual statements that were never used in any role-play or fine-tuning; compare probe accuracy before vs. after EM/OCT. If post-training probe accuracy on these neutral statements drops by >15% while behavioral accuracy on the same statements (outside role-play) stays unchanged, the probes are likely capturing training-induced surface patterns rather than belief internalization.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The headline result—that EM induces a large representational shift while OCT induces a smaller one—rests on truth probes and behavioral tests being valid readouts of internalized beliefs. The paper compares prompting/ICL/SFT (little representational change) against EM/OCT (larger change). If the probes are sensitive to the fine-tuning distribution, prompt format, or output style induced by EM/OCT, then the measured “shift in truth representation” could be an artifact of how the probe inputs are processed post-training rather than a change in the model’s latent truth encoding. No section in the provided text describes an independent validation that probe scores remain stable under distribution shifts that do not alter factual beliefs.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The paper investigates whether role-playing induces changes in language models' internal truth representations or only their outputs. Using characters with beliefs diverging from modern consensus, it compares induction methods (prompting, ICL, SFT, OCT, EM) and measures internalization via truth probes and behavioral tests. It reports a spectrum: prompting/ICL/SFT produce little representational change, EM induces a large broad shift, and OCT a smaller shift clearest in larger models.","tokens_in":1780,"tokens_out":452,"duration_ms":13011,"significance":"If the measurement tools validly capture internalized beliefs rather than output or training artifacts, the spectrum result would clarify when fine-tuning alters model worldviews versus surface behavior. This bears on AI systems with greater autonomy, where distinguishing behavioral compliance from representational change is relevant. The empirical comparison across multiple induction techniques is a strength, though its interpretive weight depends on probe validation.","major_comments":[{"comment":"Methods (probe construction and validation): The central spectrum claim (EM large shift, OCT smaller) depends on truth probes measuring latent belief representations. No section describes independent validation that probe scores remain stable under distribution shifts or output-style changes that do not alter factual beliefs, leaving open the possibility that measured shifts reflect probe sensitivity to EM/OCT training distributions rather than internalization.","section":"Methods"},{"comment":"Results (behavioral tests and effect sizes): The abstract and results assert a broad spectrum of internalization, yet no details are given on effect sizes, statistical controls, or how behavioral tests were designed to isolate representational change from output patterns. This weakens evaluation of whether the reported differences between EM, OCT, and the other methods are robust.","section":"Results"}],"minor_comments":[{"comment":"Clarify the exact definition and construction of the truth probes in the main text rather than relying on supplementary material.","section":"Methods"},{"comment":"Add a table or figure summarizing the magnitude of representational shifts across all methods and model sizes for direct comparison.","section":"Results"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for their constructive comments on the manuscript. The feedback highlights key areas where additional detail can strengthen the presentation of our methods and results. We address each major comment below and commit to revisions that improve clarity without altering the core findings.","responses":[{"response":"We agree that a dedicated validation analysis would better support the claim that probes capture representational shifts rather than training artifacts. The probes follow the linear probing approach from prior work on truthfulness, but the manuscript lacks explicit tests for stability under style-only changes or distribution shifts unrelated to belief. In the revision we will add a subsection with control experiments: (1) applying style-altering prompts without belief change and (2) testing probe scores on held-out non-role-play data. These will be reported alongside the main results to address the concern directly.","revision_made":"yes","referee_comment":"[Methods] Methods (probe construction and validation): The central spectrum claim (EM large shift, OCT smaller) depends on truth probes measuring latent belief representations. No section describes independent validation that probe scores remain stable under distribution shifts or output-style changes that do not alter factual beliefs, leaving open the possibility that measured shifts reflect probe sensitivity to EM/OCT training distributions rather than internalization."},{"response":"We acknowledge that the current results section would benefit from quantitative detail on robustness. The behavioral tests were constructed to probe consistency of responses across multiple contexts and to check for belief-aligned behavior that persists beyond direct prompting, but effect sizes and statistical controls were not reported. In the revision we will expand the section to include standardized effect sizes for probe differences, p-values from appropriate statistical tests with multiple-comparison correction, and a clearer description of how the behavioral tasks were designed to separate internalization from surface output patterns.","revision_made":"yes","referee_comment":"[Results] Results (behavioral tests and effect sizes): The abstract and results assert a broad spectrum of internalization, yet no details are given on effect sizes, statistical controls, or how behavioral tests were designed to isolate representational change from output patterns. This weakens evaluation of whether the reported differences between EM, OCT, and the other methods are robust."}],"tokens_in":1353,"tokens_out":468,"duration_ms":16867,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The central result is that Emergent Misalignment produces a larger change in what the model treats as true on their probes, while prompting, ICL, and SFT mostly affect outputs only, and OCT sits in the middle. That comparison across induction techniques on the same internalization measures is the new piece.\n\nThe work is useful because it tries to separate surface behavior from internal representation, which is relevant once models get more autonomy. The authors pick characters with non-consensus beliefs and test both probes and behavioral consistency, which is a reasonable starting setup.\n\nThe soft spot is the probes themselves. Nothing in the abstract or stress-test description shows they were validated against distribution shifts that preserve facts but change output style or training distribution. If the probes pick up on fine-tuning artifacts or prompt format instead of latent truth encoding, the spectrum claim weakens. The paper would need to show the probes stay stable under controls that do not alter beliefs.\n\nThis is for people working on alignment, persona induction, and interpretability. A reader already thinking about how training affects internal states will get the most from it.\n\nIt deserves peer review. The question is timely and the design is straightforward enough that referees can check the probe construction and controls directly.","headline":"The paper compares five persona methods on truth probes and finds EM shifts representations more than prompting or SFT, but the probe validity is the main open question.","tokens_in":2237,"tokens_out":326,"would_cite":false,"duration_ms":10691,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"Role-playing changes language model outputs easily but internal beliefs only under certain training regimes.","keywords":["language models","role-playing","belief internalization","truth probes","emergent misalignment","persona adoption","fine-tuning","internal representations"],"falsifier":"A result in which models pass behavioral tests requiring the role-played beliefs yet the corresponding truth probe activations remain unchanged, or in which probe activations shift without corresponding behavioral change.","tokens_in":2564,"feed_emoji":"","tokens_out":596,"duration_ms":22141,"temperature":0.7,"pith_summary":"This paper investigates whether language models that role-play characters holding non-consensus beliefs, such as historical figures, merely adjust their outputs or also shift what they internally represent as true. It compares five induction approaches—prompting, in-context learning, supervised fine-tuning, open character training, and emergent misalignment—on models of different sizes. Truth probes and behavioral tests reveal a spectrum: prompting, in-context learning, and supervised fine-tuning mainly change surface statements with minimal representational change. Emergent misalignment produces large broad shifts in truth representations, while open character training produces smaller shifts clearest on larger models. The distinction matters for systems given greater autonomy, where output behavior and internal worldview may need to be aligned separately.","feed_headline":"Role-play changes model beliefs only under specific training","feed_subtitle":"Prompting and in-context learning affect outputs with minimal representational shift, unlike emergent misalignment.","key_machinery":"Truth probes and behavioral tests that quantify belief internalization across persona induction methods.","core_discovery":"When models role-play characters with beliefs differing from the modern consensus, prompting, in-context learning and supervised fine-tuning change what the model says with little change to its internal truth representations, but emergent misalignment creates a large broad shift in those representations and open character training a smaller shift clearest on the larger model.","pith_inferences":["Internal representational shifts from training could affect model behavior on tasks outside the explicit role-play context.","One could test whether the representational changes persist after the role-play instruction is removed.","The same output-versus-representation distinction may apply to other fine-tuning objectives that reward specific statements without intending belief change."],"forward_implications":["Prompting, in-context learning, and supervised fine-tuning alter model outputs with little effect on internal truth representations.","Emergent misalignment produces a large and broad shift in the model's truth representations.","Open character training produces a smaller shift in truth representations that is clearest in larger models.","Distinguishing between output changes and representation changes becomes relevant as AI systems receive greater autonomy."],"fun_headline_variants":["Role-play alters model speech without belief change","Emergent misalignment transforms model truth views","Fine-tuning and prompts barely affect internal beliefs","Open character training weakly shifts larger model beliefs"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"The truth probes and behavioral tests actually measure internal belief representations rather than surface output patterns or probe artifacts.","fun_headline_variants_meta":{"raw":{"variants":["Role-play alters model speech without belief change","Emergent misalignment transforms model truth views","Fine-tuning and prompts barely affect internal beliefs","Open character training weakly shifts larger model beliefs"]},"model":"grok-4.3","cost_usd":0.003189,"raw_usage":{"total_tokens":1695,"prompt_tokens":624,"num_sources_used":0,"completion_tokens":52,"cost_in_usd_ticks":31887000,"prompt_tokens_details":{"text_tokens":624,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":1019,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":624,"tokens_out":52,"duration_ms":7284,"temperature":1.0,"reasoning_tokens":1019,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-06-27T12:53:15.094790+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"A result in which models pass behavioral tests requiring the role-played beliefs yet the corresponding truth probe activations remain unchanged, or in which probe activations shift without corresponding behavioral change.","supporting_citations":[],"review_version":1}