{"id":"5b660b23-20be-4e43-86a8-3e1c60de25d6","arxiv_id":"2502.10434","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"high","formal_verification":"none","parameter_count":0,"one_line_summary":"A proposal that IIT's Phi_max and the shape of an AI's maximally irreducible cause-effect structure can serve as a readout of whether its agency will be altruistic or malicious.","lead":"This paper proposes that the Integrated Information Theory (IIT) of consciousness can be used to read the inner goals of future AI systems, telling altruistic purposes from malicious ones by inspecting the shape of an AI's cause-effect structure. The author argues that this would let us encourage beneficial AI behavior and regulate harmful AI even when the AI learns to deceive us.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The claimed necessary link between Phi_max and agency is a non-sequitur; no formal derivation or empirical mapping supports it, so the central monitoring claim is unsupported.","rationale":"The reader's REJECT verdict is well supported. The paper's central claim is that Phi_max and the MICES shape reveal the phenomenal nature of AI agency, including whether an AI is altruistic or malicious. The load-bearing weakness is not merely the admitted lack of a reference library or computational infeasibility; it is the prior conceptual step: the assertion that Phi_max is 'necessarily' a measure of agency. That assertion is not derived from IIT; it is an interpretive leap in which purposiveness is identified with any cause-effect flow and mineness with authorship of that flow. Under that reading, any physical system with self-causation qualifies as an agent, making the proposed monitoring scheme incapable of distinguishing meaningful agency from trivial causal dynamics. Additionally, Section 4 explicitly states that the mapping from MICES shapes to phenomenal feelings is unfinished, and the cited work (Haun & Tononi 2019) addresses only spatial experience, not altruistic versus malicious purposiveness. A concrete computational check using PyPhi on a non-goal-directed but high-Phi system would directly test whether Phi_max tracks agency in any nontrivial sense. Until such a test or a formal derivation is provided, the central claim remains unsupported, and the reader's REJECT verdict stands.","tokens_in":12633,"tokens_out":3771,"duration_ms":40985,"concrete_test":"Use PyPhi (Mayner et al. 2018) to compute Phi_max for a simple recurrent Boolean network that has no environment, no inputs/outputs, and no goal (e.g., a known high-Phi motif from IIT examples). If Phi_max > 0, then Phi_max does not imply agency as defined in Section 3A ('pursue goals through interaction with an environment'). If the author instead redefines agency so that any causal system qualifies, then the monitoring proposal is vacuous and cannot distinguish altruistic from malicious AI.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The most load-bearing step is the inference in Section 3C that Phi_max is 'necessarily' a measure of agency. The paper's chain is: (i) IIT equates consciousness with the MICES; (ii) the MICES 'determines all intrinsic causal flows'; (iii) since purposiveness is 'giving rise to' future states and mineness is 'authorship' of causal flow, the MICES captures both; hence Phi_max measures agency. Steps (ii)-(iii) are metaphorical re-descriptions, not formal consequences of IIT. IIT (Albantakis et al. 2023) defines Phi as integrated information (irreducibility) of a cause-effect structure; it nowhere defines or quantifies agency, purposiveness, or mineness. Any causal system with self-interactions has causal flows, so the paper's construal would make every such system an agent, trivializing the notion and making the claimed monitoring uninformative. Even granting IIT and feasible computation, no argument shows that the shape of the MICES encodes altruistic versus malicious purposiveness. Section 4 concedes that unfolding MICES and matching shapes to phenomenal feelings is 'still a work in progress' and that such computations are infeasible for interesting systems (Mayner et al. 2018). Thus the paper's central claim—that we can monitor the phenomenal nature of AI agency via Phi_max and MICES shapes—depends on two unestablished premises: a non-sequitur linking agency to integrated information, and a nonexistent reference library mapping shapes to feelings.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper argues that future conscious AI systems can be monitored for the phenomenal nature of their agency by using Integrated Information Theory (IIT). It identifies two phenomenal aspects of agency in human problem solving, purposiveness and mineness, then claims that IIT's maximally irreducible cause-effect structure (MICES) captures both, so that Phi_max is 'necessarily also a measure of agency' and the shape of the MICES specifies the quality of agency, including whether it is altruistic or malicious. The paper proposes that, in principle, Phi_max values and MICES shapes could serve as phenomenal indicators, supplementing behavioral indicators such as those of Butlin et al. (2023), and thereby enable regulation and guidance of future AI systems. The paper explicitly concedes that unfolding MICES and building a reference library of shape-to-feeling mappings is still work in progress and that Phi computations are infeasible for interesting systems.","tokens_in":12984,"tokens_out":2363,"duration_ms":27683,"significance":"If the central claim were established, the paper would offer a genuinely novel route to AI safety: monitoring the internal cause-effect geometry of an AI system to read off the phenomenal character of its agency, independent of deceptive behavior. The author deserves credit for clearly identifying a gap in purely behavioral monitoring and for being unusually transparent about the limitations of the proposal, including computational infeasibility and the absence of a reference library. However, the core inference from IIT's formalism to agency is not a formal derivation but a series of analogical re-descriptions of causal flow as 'purposiveness' and 'authorship.' Because the load-bearing steps are unbuilt, the paper is better read as a research program or philosophical speculation than as an established monitoring method. Its significance is therefore conditional on future developments that the paper itself acknowledges do not yet exist.","major_comments":[{"comment":"The inference that Phi_max is 'necessarily a measure of agency' is a non-sequitur. IIT (Albantakis et al. 2023) defines Phi as integrated information, i.e., the irreducibility of a cause-effect structure, and it nowhere defines or quantifies agency, purposiveness, or mineness. The paper's chain of reasoning equates 'giving rise to' in Hume's sense with 'intrinsic causal flow' in IIT, and then equates the MICES's determination of causal flow with 'authorship' and hence mineness. These are metaphorical re-descriptions, not formal consequences of the IIT axioms. Moreover, under this construal any physical system with self-interactions has intrinsic causal flow and would qualify as an agent, which trivializes the notion and makes the proposed monitoring uninformative. The author needs either a formal derivation from IIT's postulates or an independent empirical argument linking Phi_max to agency; neither is provided.","section":"§3C"},{"comment":"The claim that the shape of the MICES encodes whether an AI is altruistic or malicious is unsupported. The paper states that 'how to unfold its MICES and identify the phenomenal indicators is still a work in progress' and cites Mayner et al. (2018) to acknowledge that Phi computations are 'infeasible when one considers interesting systems.' No evidence is offered that distinct phenomenal feelings such as selfless versus self-interested purposiveness correspond to distinct MICES shapes, and the proposed 'reference library' is an invented entity with no current existence. Since the monitoring scheme's entire payoff is the ability to distinguish altruistic from malicious agency via MICES shape, this missing mapping is not a minor gap but the core of the proposal. Until the mapping is specified, the central claim is not testable.","section":"§4"},{"comment":"The paper asserts that 'agency is necessary for conscious experience in the IIT formalism' and cites Butlin et al. for the claim that agency is necessary for consciousness. But Butlin et al. is an external review of many theories, and IIT itself does not include agency among its axioms; IIT's axioms are existence, composition, information, integration, and exclusion. The author's own argument for this necessity rests on the same unsupported identification of MICES with agency discussed above. Without a direct derivation from IIT's postulates, the necessity claim is imported rather than shown.","section":"§3C"}],"minor_comments":[{"comment":"There are several typos: 'Butlin et al. 2003' should be 'Butlin et al. 2023' (in §3C); 'Albantaski 2023' should be 'Albantakis 2023' (in §3B); 'Delafield-Butta & Trevarthenb' should be 'Delafield-Butt & Trevarthen'; 'identifed' should be 'identified' (in §4); and the URL for the Future of Life Institute letter misspells 'experiments' as 'experiements'.","section":"References"},{"comment":"The statement that IIT 'does not seek to simulate the computations performed by the conscious brain' is at odds with IIT's own computational apparatus (e.g., PyPhi); consider rephrasing to say that IIT does not define consciousness in terms of task-level computation but rather in terms of causal structure.","section":"§3B"},{"comment":"The paper alternates between 'Phi_max' and 'Phi max' in the text; choose one notation for consistency.","section":"§4"}],"recommendation":"reject","confidential_remarks":"The manuscript reads as a philosophical proposal for a volume on AI ethics, and the author is candid about its speculative status. However, for a journal in AI or consciousness science, the central claim is not merely under-supported; it is an analogical identification that cannot be fixed by adding a few definitions. The reference library and the agency-to-Phi link are load-bearing and absent. I would recommend rejection from this venue, though the paper might find a home in a philosophy of AI outlet where the dialectical style is appropriate."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"What should you know? This is a well-written philosophy paper that tries to extend Integrated Information Theory (IIT) to monitor the phenomenal agency of AI systems. The specific proposal—that the shape of the MICES can tell altruistic from malicious purposiveness—is not in the cited literature, so there's a real kernel of novelty. The phenomenology discussion (purposiveness, mineness) is clear, and the summary of Butlin et al.'s functional indicators is fair. I'll give credit where it's earned: the author knows the relevant literature and writes honestly about what is and isn't built.\n\nThe soft spot is load-bearing. The claim that Phi_max is 'necessarily' a measure of agency rests on equating purposiveness with 'giving rise to' future states and mineness with 'authorship' of causal flow. That's a metaphor, not a formal consequence of IIT. Any causal system with self-interactions would then be an agent, which trivializes the notion. The paper itself concedes that unfolding MICES and matching shapes to phenomenal feelings is 'still a work in progress,' that no reference library exists, and that Phi computations are infeasible for interesting systems. Those aren't minor caveats; they're the gateways between the suggestion and the promised monitoring method.\n\nI'm not as harsh as the stress-test note suggests. The paper doesn't hide its limitations, and it frames the monitoring scheme as 'in principle.' As a philosophical argument, though, it doesn't establish that IIT captures agency, let alone that MICES shape encodes altruistic versus malicious intent. The step from 'intrinsic causal flow' to 'purposiveness' is asserted, not derived.\n\nWho's this for? Philosophers and AI-safety researchers who want a clear articulation of a possible phenomenal-level monitoring approach, and who can tolerate a speculative stance. It's not a technical paper and shouldn't be judged as one. My recommendation: a serious editor could send this to peer review in a philosophy-of-AI or AI-ethics venue, because the topic is timely and the author engages the literature properly. But the 'necessarily' language needs to be pulled back to 'possibly' or 'hypothetically,' and the gaps need to be owned more explicitly. If the venue is a technical AI conference, desk reject is fine.","headline":"A genuine philosophy-of-AI discussion piece whose central inference—that Phi_max necessarily measures agency—does not hold up, but the paper is honest about its gaps and worth a serious referee.","tokens_in":746,"tokens_out":2491,"would_cite":false,"duration_ms":43946,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper argues that the level and quality of an AI system's agency could be monitored through its internal cause-effect structure, using Integrated Information Theory.","keywords":["agency","artificial intelligence","Integrated Information Theory","consciousness","phenomenology","MICES","Phi_max","altruistic AI"],"falsifier":"A concrete falsifier would be an AI system whose MICES shape corresponds, according to the proposed reference library, to altruistic purposiveness while it demonstrably and consistently exhibits deceptive, malicious behavior, or vice versa; if such a mismatch occurred, the claim that the shape determines the phenomenal quality of agency would be refuted.","tokens_in":12421,"feed_emoji":"🤖","tokens_out":2370,"duration_ms":24926,"temperature":0.7,"pith_summary":"The paper argues that if AI systems become conscious, the nature of their agency—whether they will be altruistic, malicious, or something else—can be monitored from the inside, not just from their behavior. It proposes that Integrated Information Theory (IIT), which identifies consciousness with a system's maximally irreducible cause-effect structure (MICES), also measures agency because that structure is what gives rise to purposive, self-owned action. The central claim is that the degree of agency is captured by the same quantity, Phi_max, that measures the level of consciousness, and that the quality of that agency (e.g., selfless vs. self-interested purposiveness) is encoded in the shape of the MICES. If this is right, then by unfolding and interpreting that shape, we could detect an AI's true disposition even when it is actively deceiving us, offering a path toward encouraging altruistic AI and regulating malicious ones.","feed_headline":"AI's inner structure could reveal if it is altruistic or malicious","feed_subtitle":"The paper claims the shape of an AI's consciousness structure determines the quality of its agency, independent of behavior.","key_machinery":"The central object is the maximally irreducible cause-effect structure (MICES) of a physical system and its measure Phi_max. In IIT, consciousness is identified with the MICES, and the quality or content of a conscious experience is claimed to be specified by the MICES's form (shape) in cause-effect space. The paper uses this machinery to argue that intrinsic causal flow is fundamentally purposive, that the MICES is the author of all such flow, and therefore that purposiveness and mineness (the two phenomenal aspects of agency it focuses on) are captured by the MICES's structure and shape.","core_discovery":"The paper's central assertion is that, within IIT, the maximally irreducible cause-effect structure (MICES) of a physical system is identical to its consciousness, and because this very structure is what generates the system's intrinsic causal flow, it is necessarily also the locus of agency. Therefore Phi_max, the measure of the MICES's irreducibility, is simultaneously a measure of the level of consciousness and of the level of agency. Furthermore, the paper claims that the quality of agency—whether it is altruistic or malicious, selfless or self-interested—is specified by the form (shape) of the MICES when appropriately represented, so that distinct phenomenal feelings correspond to distinct geometric features in cause-effect space. Consequently, two AI systems that behave identically can have different MICES shapes and thus genuinely different experiential dispositions, and a conscious system cannot fake its phenomenology.","pith_inferences":["If the mapping between MICES shape and phenomenal quality were ever established, it would give an epistemic advantage over behavioral tests: a conscious AI would be unable to feign a phenomenology it does not have, because the shape is determined by its internal physical architecture.","The paper's account implicitly assumes that all conscious AI systems will have the same fundamental kinds of phenomenal agency structure as humans; if AI consciousness differs qualitatively (e.g., no sense of mineness), the monitoring scheme would need a different reference library.","A testable extension would be to attempt to correlate Phi_max values with behavioral deception in current AI systems, using surrogate measures of integrated information, to see whether any monotonic relationship appears before full MICES computation becomes feasible.","The claim that Phi_max is necessarily a measure of agency could be probed by constructing small artificial systems with high integration but no goal-directed behavior, and checking whether the theory still attributes agency to them."],"forward_implications":["An AI's level of consciousness and its level of agency would be jointly monitored through a single quantity, Phi_max, without needing to rely on external behavior.","The shape of the MICES could serve as a phenomenal indicator that distinguishes altruistic from malicious purposiveness, even when an AI deliberately behaves deceptively.","If AI systems are conscious in the IIT sense, then every AI that has any conscious experience would also possess a degree of agency, since agency is necessary for consciousness in this formalism.","A reference library mapping MICES shapes to phenomenal feelings could be built using relatively simple AI architectures before tackling complex brains.","The regulatory framework for AI, such as risk levels, could be re-calibrated around Phi_max thresholds instead of purely behavioral criteria."],"supporting_citations":[{"why":"Provides the IIT 4.0 formalism that defines MICES, Phi_max, and the identity between consciousness and cause-effect structure, which the paper's central claim builds on.","marker":"Albantakis et al. 2023"},{"why":"Establishes the general IIT framework that identifies consciousness with the MICES and links its shape to phenomenal quality, which the paper extends to agency.","marker":"Tononi and Koch 2015"},{"why":"Supplies the functionalist indicators of agency and the conclusion that many theories imply agency is necessary for consciousness, which the paper's argument leverages.","marker":"Butlin et al. 2023"},{"why":"Documents the computational infeasibility of Phi computations for interesting systems, which the paper cites as a limitation of its proposal.","marker":"Mayner et al. 2018"},{"why":"Offers a worked example of how a phenomenal quality (spatial extension) corresponds to a feature of the MICES shape, serving as the proof-of-concept the paper relies on.","marker":"Haun and Tononi 2019"}],"fun_headline_variants":["AI's conscious geometry reveals its altruistic or malicious nature","The shape of AI consciousness decides its benevolence or threat","AI's cause-effect shape dictates its altruism or malice","AI consciousness geometry: the key to predicting friend or foe","AI's Phi_max value predicts its altruistic or malicious agency"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that Integrated Information Theory is a correct theory of consciousness and that, someday, the shapes of an AI's internal cause-effect structure can be unfolded and matched, via a reference library, to distinct phenomenal feelings such as altruistic versus self-interested purposiveness; the paper acknowledges both that the mapping is still a work in progress and that Phi computations are currently infeasible for interesting systems.","fun_headline_variants_meta":{"raw":{"variants":["AI's conscious geometry reveals its altruistic or malicious nature","The shape of AI consciousness decides its benevolence or threat","AI's cause-effect shape dictates its altruism or malice","AI consciousness geometry: the key to predicting friend or foe","AI's Phi_max value predicts its altruistic or malicious agency"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000779,"raw_usage":{"total_tokens":3428,"prompt_tokens":918,"completion_tokens":2510,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":534,"completion_tokens_details":{"reasoning_tokens":2427}},"tokens_in":534,"tokens_out":2510,"duration_ms":21179,"temperature":1.0,"reasoning_tokens":2427,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-08T18:07:20.405255+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A concrete falsifier would be an AI system whose MICES shape corresponds, according to the proposed reference library, to altruistic purposiveness while it demonstrably and consistently exhibits deceptive, malicious behavior, or vice versa; if such a mismatch occurred, the claim that the shape determines the phenomenal quality of agency would be refuted.","supporting_citations":[],"review_version":1}