{"id":"e415196e-e102-4000-ac4a-c9c80c0b8b0e","arxiv_id":"2607.20708","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"In a reward-free active inference agent, Φr localizes to the slow perspective latent g, its aggregate magnitude comes from the recurrent architecture rather than learning, and learning only becomes visible as a decoupling sign flip plus regime-invariant stability.","lead":"This paper measures Integrated Information Decomposition (Φr) inside a reward-free active inference agent, finding that the slow 'perspective' latent carries the temporal signal, but the scalar value mostly reflects the recurrent architecture rather than learning. It argues that Φr must be read at the atom-composition level: learning flips whole-to-whole 'decoupling' from negative to positive and makes it stable across regime changes.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Fiedler bipartition instability across cohorts could make decoupling sign flip a partition artifact; need fixed partition or stability audit.","rationale":"The reader's conditional verdict hinges on the same estimator-validity concern I identify. The load-bearing piece of the paper's central claim—that learning qualitatively reorganizes Φr composition (decoupling flips sign and becomes regime-invariant) rather than merely scaling it—depends on the decoupling atom being comparable across trained/untrained and pre/post conditions. Because the Fiedler bisection is re-estimated per episode from the thresholded MI graph, the two-node reduction's 'whole' is not guaranteed to be the same object across conditions. If the partition shifts, the sign flip may be an artifact of relabeling rather than a genuine change in whole-to-whole causal contribution. The paper explicitly defers a partition-stability audit, so this is a known gap. A concrete test with a fixed partition would settle whether the sign flip survives. Since the reader already flags the estimator as the weakest assumption and asks for conditional acceptance, my concern reinforces that conditionality without escalating it. Thus the verdict remains CONDITIONAL (UNCHANGED).","tokens_in":9457,"tokens_out":3340,"duration_ms":25359,"concrete_test":"Fix the bipartition across all conditions—e.g., use the trained-control Fiedler partition for every trajectory—and recompute the decoupling, downward causation, and part-driven groups. If the decoupling sign flip (untrained −0.51 → trained +0.12) or its regime-invariance (trained Δ=+0.04, p=0.23) disappears, the compositional claim is a partition artifact. Additionally, report the fraction of episodes whose Fiedler partition matches the trained-control partition, and the dimension membership of A/B across cohorts.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central compositional claim is that learning shifts the decoupling atom from negative to positive and makes it regime-invariant (Sections 3.3–3.4). Decoupling is defined on a two-node reduction Y=(Y_A,Y_B) where the bipartition is chosen per episode by the Fiedler vector of the thresholded MI matrix (Section 2.4, Eqs. 2–3). The 'whole' is therefore not fixed: the set of g-dimensions in A vs B can differ between untrained and trained agents and between pre- and post-switch windows. A sign flip in the decoupling atom could then reflect a change in which dimensions are grouped into the whole, rather than a change in whole-to-whole causal contribution. The paper's own limitations list a 'partition-stability audit' as future work, but the current cross-condition atom comparisons presuppose partition comparability. Until the partition is shown stable (or the analysis is repeated with a fixed partition), the sign flip cannot be unambiguously attributed to learning-induced reorganization of causal structure.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies an active-inference agent whose architecture separates a fast perceptual latent z from a slow GRU-based global latent g, with g structurally decoupled from policy gradients. In a reward-free regime-switching gridworld, the author computes a local pointwise Integrated Information Decomposition (Phi_r) and reports three main findings: (1) Phi_r concentrates in g rather than z; (2) aggregate Phi_r(g) is largely architectural and decreases with training; (3) learning reorganizes the atom composition of Phi_r(g), shifting the decoupling atom from negative to positive and making it approximately regime-invariant, while downward causation carries the main regime-dependent change. The paper argues that scalar Phi_r should not be read as a direct index of learned integration and that the atom-level decomposition is needed to see what learning changes.","tokens_in":9671,"tokens_out":4789,"duration_ms":40651,"significance":"If the results hold, the paper offers a useful, sobering counterpoint to the recent use of scalar Phi_r as a measure of learned integration in RL agents: the architectural substrate alone can supply high aggregate Phi_r, and the meaningful learning effect only appears in the composition of PhiID atoms. The control design is strong in conception: within-episode shuffling isolates temporal structure, the untrained baseline isolates learning, and the pre/post regime comparison tests adaptivity. The effects are large and internally consistent, and the manuscript is candid about its limitations, including the need for a partition-stability audit and targeted ablations. Code and data availability further support reproducibility. However, the central compositional claim depends on two assumptions that are not yet validated: the stability of the per-episode Fiedler bipartition across conditions, and the meaningfulness of the sign of the local pointwise PhiID atom groups. The statistical tests also suffer from pseudo-replication. These issues are fixable but currently make the evidence suggestive rather than conclusive.","major_comments":[{"comment":"The inferential statistics treat 300 episodes from 30 seeds as independent samples. Welch's and paired t-tests over episodes (e.g., Sec. 3.2: t=-25.1; Sec. 3.3: t=8.6; Sec. 3.4: t=6.6) ignore nesting of episodes within seeds, making p-values anti-conservative. The large effect sizes likely survive a seed-level analysis, but the quantitative support and especially the regime-invariance claim for decoupling (Sec. 3.4, p=0.23) need re-analysis at the seed level or with a mixed-effects model before being accepted.","section":"Sec. 2.3 and 3.2-3.4"},{"comment":"The decoupling atom is defined on a two-node reduction whose bipartition is chosen per episode by the Fiedler vector of the thresholded MI matrix. Cross-cohort comparisons (untrained vs trained, pre- vs post-switch) therefore compare whole-to-whole contributions computed over potentially different partitions. A sign flip in decoupling could reflect a change in which g dimensions are grouped into A vs B rather than a change in whole-to-whole causal contribution. The manuscript itself lists a partition-stability audit as future work, but the cross-condition comparisons in Secs. 3.3-3.4 presuppose partition comparability. Please report partition stability across cohorts/conditions, use a fixed a priori partition, or show the sign flip survives on a common partition.","section":"Sec. 2.4 Eq. (3); Sec. 3.3-3.4"},{"comment":"The atom-level interpretation, especially the sign and relative magnitude of the decoupling atom, rests entirely on the local pointwise Gaussian PhiID estimator with within-partition averaging and median summarization. The paper correctly notes that this differs from the non-negative Rosas-style construction, but it does not validate whether the sign of these local atoms is a faithful measure of whole-to-whole causal contribution in this setting. A synthetic benchmark with known ground-truth atoms, or a comparison against a non-negative PhiID formulation, is needed before the sign flip can be read as a mechanism rather than an estimator artifact.","section":"Sec. 2.4, Eqs. (2-4)"}],"minor_comments":[{"comment":"The same paired t-value (43.6) is reported for the localization comparison and the shuffle collapse. These are different pairs of distributions; please verify that both statistics are correct or clarify why they coincide.","section":"Sec. 3.1"},{"comment":"A forced-split diagnostic on the joint [z,g] trajectory is described but no results are reported for it. Either present the diagnostic or remove it from the methods text.","section":"Sec. 2.4"},{"comment":"Means and SDs are reported at the episode level, which conflates within-seed and between-seed variance. Reporting seed-level means and confidence intervals would make the magnitude of the effects easier to assess.","section":"Sec. 3.2-3.4"},{"comment":"The claim that trained decoupling is 'almost unchanged' across the regime switch is a null result with p=0.23. Given the pseudo-replication concern, please report a seed-level effect size and a confidence interval, or an equivalence test, to support the claim of invariance.","section":"Sec. 3.4"}],"recommendation":"major_revision","confidential_remarks":"The paper is largely a single-author self-contained analysis, but it leans heavily on the author's own prior work for the architecture and the 'perspective latent' construct. The central claim is defensible, but the Fiedler partition issue is especially important because the main novel finding is a sign flip in a quantity defined on a data-dependent partition. If the author can provide a fixed-partition or stability-audited analysis together with seed-level statistics, the paper would be suitable for publication."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague — this paper is worth reading, but the headline compositional claim needs a fixed partition before we can trust it. It applies the Φr/ΦID atom decomposition to an active inference agent with a slow recurrent 'perspective' latent g, and reports three things: Φr localizes to g, aggregate Φr is higher in untrained agents than in trained ones, and atom-level decoupling flips from negative to positive with training while staying roughly constant across a regime switch, with downward causation carrying the adjustment. The localization and shuffle controls are clean: the 129x contrast between g and z and the 99.4% collapse under temporal shuffle are large and internally consistent. The author is also honest in the limitations section, and code/data are provided.\n\nThe soft spots are real, and two are not minor. First, the statistics treat 300 episodes from only 30 seeds as independent in paired t-tests; that overstates significance, though the effect sizes are large enough that a seed-level analysis would likely still find the aggregate effects. Second, the claim that decoupling is regime-invariant rests on failing to reject the null (p=0.23); without an equivalence test that is weak support. Third—and the stress-test note gets this right—the Fiedler bipartition is recomputed per episode, so the two-node 'whole' is not fixed across trained/untrained or pre/post-switch conditions. A sign flip in decoupling could be a partition artifact rather than a genuine change in whole-to-whole contribution. The author lists a partition-stability audit as future work, but the cross-condition comparisons in the current paper presuppose comparability.\n\nI would still send it to referees: the caution against scalar-Φr readings is worth airing in public, and the paper knows its own limits. But I would ask for a fixed-partition analysis or stability audit, a seed-level statistical treatment, and an equivalence test for invariance before accepting the compositional story.","headline":"A useful, honest caution against reading scalar Φr as learned integration, but the sign-flip claim needs a fixed bipartition before it can be believed.","tokens_in":10180,"tokens_out":1879,"would_cite":true,"duration_ms":22521,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The causal-emergence signal in this active inference agent lives in the slow 'perspective' latent, not the fast perceptual code, and learning reshapes its composition rather than its magnitude.","keywords":["causal emergence","integrated information decomposition","active inference","perspective latent","slow latent variable","recurrent neural network","regime switching","information dynamics"],"falsifier":"If the negative-to-positive decoupling flip disappears under a non-Gaussian PhiID estimator, or when within-partition averaging is replaced by a different reduction, the central dissociation would be undercut. The most direct experiment is ablating the stop-gradient separation: a trained agent without architectural decoupling from policy gradients should not show the stable positive decoupling if the paper's claim is right.","tokens_in":9265,"feed_emoji":"🧠","tokens_out":4634,"duration_ms":38003,"temperature":0.7,"pith_summary":"This paper asks where the integrated-information-decomposition signal Phi_r appears inside an active inference agent and whether learning creates it. Measuring Phi_r on the fast perception latent z and the slow global latent g, it finds that Phi_r concentrates in g, and that an untrained copy of the same recurrent architecture shows higher aggregate Phi_r than the trained agent. Learning's real effect is compositional: the decoupling atom flips from negative to positive and becomes stable across an environmental regime switch, while downward causation shrinks and carries the regime-dependent adjustment. The paper concludes that scalar Phi_r is not a reliable index of learned integration; only atom-level decomposition reveals what learning changes.","feed_headline":"Untrained agents look more integrated than trained ones","feed_subtitle":"Atom-level decomposition shows training flips whole-to-whole coupling positive even as total integration drops.","key_machinery":"The central object is the perspective latent g, a GRU-gated slow variable in the agent's world-modeling path, updated from the fast perceptual latent and previous action, trained by one-step prediction error and a smoothness regularizer, and shielded from policy gradients by stop-gradient operators. The analysis applies a pointwise PhiID estimate to g: the trajectory is standardized, a thresholded lag-1 Gaussian mutual information matrix is built, a Fiedler bipartition splits the latent dimensions into two nodes, within-partition averages form a two-node trajectory, and local PhiID atoms are computed. The atoms are grouped into decoupling (single whole-to-whole atom), downward causation (thr","core_discovery":"The central finding is that Phi_r-relevant temporal organization in this active inference agent is localized in the slow perspective latent g and is largely supplied by the recurrent architecture itself, not by learning. Training reduces aggregate Phi_r(g), but it reorganizes the atom composition: decoupling flips from negative to positive and stays stable across a regime switch, while downward causation decreases and carries the response to the switch. The paper argues that scalar Phi_r conflates these components and should not be read as a direct measure of learned integration.","pith_inferences":["A direct test the paper does not run: remove the stop-gradient separation between g and policy gradients, retrain, and check whether the decoupling sign flip disappears; if it does, the flip is tied to the architectural shielding of the perspective latent.","The same atom-composition analysis could be applied to reward-driven reinforcement learning agents that show Phi_r growing with training; the untrained-baseline contrast predicts that part of that growth is compositional, not an increase in integration amount.","If the decoupling flip is produced by predictive smoothing alone, then any recurrent latent trained on temporal prediction would show it; comparing against a recurrent latent lacking the perspective-specific coupling would settle whether perspective-specific organization is needed.","The post-training regime-invariance of decoupling is a candidate operational signature of a stable 'world model' state that persists while lower-level predictive engagement adjusts."],"forward_implications":["Scalar Phi_r should not be used alone as an integration or causal-emergence index, because all three reported effects are invisible to magnitude alone.","Untrained-architecture baselines should accompany Phi_r measurements, since the recurrent substrate alone can produce high values that look like integration.","In reward-free active inference, training can lower raw Phi_r while qualitatively changing its composition, so learned organization shows up only at the atom level.","The perspective latent g is the architectural locus of Phi_r-relevant slow temporal organization, whereas the fast perceptual latent contributes almost none.","After learning, the whole-to-whole decoupling component is regime-invariant while whole-to-part downward causation adapts to regime change, suggesting a division between stable integration and adaptive engagement."],"fun_headline_variants":["Training flips coupling sign as total Phi_r drops","Architecture, not learning, drives integration measure","Slow latent carries Phi_r; learning only reshapes it","Phi_r drops with training, but coupling turns positive","Untrained agents look more integrated; training flips coupling"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The whole interpretation rests on the pointwise Gaussian PhiID estimator—with its Fiedler bipartition, two-node averaging, and median summarization—being faithful enough that the decoupling atom's sign flip reflects a real change in whole-to-whole causal contribution rather than an artifact of the estimator or partition choice.","fun_headline_variants_meta":{"raw":{"variants":["Training flips coupling sign as total Phi_r drops","Architecture, not learning, drives integration measure","Slow latent carries Phi_r; learning only reshapes it","Phi_r drops with training, but coupling turns positive","Untrained agents look more integrated; training flips coupling"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000101,"raw_usage":{"total_tokens":821,"prompt_tokens":673,"completion_tokens":148,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":417,"completion_tokens_details":{"reasoning_tokens":81}},"tokens_in":417,"tokens_out":148,"duration_ms":3217,"temperature":1.0,"reasoning_tokens":81,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T09:34:51.136095+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"If the negative-to-positive decoupling flip disappears under a non-Gaussian PhiID estimator, or when within-partition averaging is replaced by a different reduction, the central dissociation would be undercut. The most direct experiment is ablating the stop-gradient separation: a trained agent without architectural decoupling from policy gradients should not show the stable positive decoupling if the paper's claim is right.","supporting_citations":[],"review_version":1}