{"id":"45e7e852-992f-4486-9a83-acbeaef77eab","arxiv_id":"2412.02579","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Structural independence in a factored space is equivalent to conditional independence in all product distributions, generalizing d-separation to deterministic functions.","lead":"This paper introduces factored space models, a graph-free way to represent randomness as independent factors, and proves that a notion called structural independence exactly matches statistical independence in every product distribution. It matters because this extends the classical d-separation theorem to variables that are deterministic functions of other variables, which ordinary causal graphs cannot handle.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The main theorem is proved only for finite discrete variables, and Definition 6.1's zero-probability convention makes the claimed continuous generalization false: with continuous Z, every conditioning event has probability zero, so conditional independence becomes vacuous.","rationale":"The reader's weakest-assumption analysis already identified finiteness. I agree and want to sharpen why it is load-bearing rather than cosmetic. Within the stated finite/discrete domain, the proof appears coherent: soundness and completeness for events are supplied in Appendix C, and the lifting to variables via Lemma 4.9 uses finiteness of value spaces. The issue is the manuscript's own scope claims. The abstract and introduction motivate continuous quantities, but Section 4 restricts to discrete variables and Definition 6.1 conditions on atomic events {Z=z} with a zero-probability convention. In any non-atomic distribution this makes every conditional independence statement vacuously true, so the theorem cannot be a generalization of d-separation to continuous variables. The concrete counterexample with Z=U1+U2 shows the equivalence fails as soon as one factor is continuous, independently of any proof-technical gap. This does not invalidate Theorem 6.2; it requires the authors to either extend the formalism or explicitly scope the central claim and the motivating examples to finite discrete systems. Therefore I leave the reader's CONDITIONAL verdict unchanged.","tokens_in":31936,"tokens_out":20506,"duration_ms":214083,"concrete_test":"Run the following continuous counterexample: let Ω=[0,1]^2 with U1,U2 independent uniform, and set X=Y=Z=U1+U2. Verify that (a) for every factorizing P and every z, P({Z=z})=0, so Definition 6.1 declares X⊥⊥_P Y | Z; and (b) H(X|z)=H(Y|z)={1,2}, since the line {U1+U2=z} is not a product set, so X and Y are not structurally independent given Z. If confirmed, Theorem 6.2 cannot be extended to continuous variables without changing Definition 6.1, and the manuscript must explicitly state the finite/discrete scope.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim, Theorem 6.2, is internally valid as stated: Definition 4.2 fixes a finite product Ω, and Section 4 states that all variables are discrete. The problem is the paper's framing that this 'generalizes' d-separation and covers examples such as temperature as a function of particle kinetic energies. Those examples are continuous, and the theorem cannot be extended to them under the paper's own Definition 6.1. There, conditional independence of events given C is defined by P(A∩C)P(B∩C)=P(A∩B∩C)P(C), with the convention that if P(C)=0 the statement holds vacuously. For any non-atomic continuous variable Z, P({Z=z})=0 for every z and every factorizing distribution, so X⊥⊥_P Y | Z is true for all X,Y,P. Concretely, take Ω=[0,1]^2, X=Y=Z=U1+U2 with uniform independent factors. Then X⊥⊥_P Y | Z holds for every factorizing P by the zero-probability convention, yet H(X|z)=H(Y|z)={1,2}, so structural independence fails. Thus the finite/discrete restriction is not a cosmetic gap; without it the stated equivalence is false. The manuscript should either develop a measure-theoretic conditional-independence notion or explicitly scope all claims and motivating examples to finite discrete systems.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"Factored space models (FSMs) represent a finite sample space as a product of independent factors and treat arbitrary variables as functions of the background factors. The paper defines histories H(X | C) and structural independence X ⊥_Ω Y | Z by disjointness of histories on each value of Z. Its central result, Theorem 6.2, states that for variables on a finite factored space, structural independence given Z holds if and only if conditional independence holds in every distribution that factorizes over the space. The proof is broken into soundness and completeness for events (Lemmas 6.3 and 6.4), with the completeness proof using 'cohistories' of probabilistically relevant factors, and a local-to-global strengthening (Proposition 6.6). The paper also proves that structural independence forms a compositional semigraphoid, constructs an FSM from any DAG, shows that node-level structural independence coincides with d-separation, and gives an example in which FSMs are strictly more expressive than DAGs as perfect maps.","tokens_in":32121,"tokens_out":20199,"duration_ms":207980,"significance":"If Theorem 6.2 holds, it gives a complete graph-free characterization of conditional independence in finite product spaces, including deterministic functions, and it recovers d-separation for DAG node variables as a special case. The proof in the appendix is detailed and largely self-contained, and the soundness/completeness argument via cohistories is convincing; I did not find a gap in the finite discrete setting. The local-to-global strengthening is a particularly clean addition. Two caveats temper significance: the theorem is confined to finite factored spaces, while several motivating examples are continuous; and the authors credit the original theorem to prior work [8], so the novelty lies in the proof and framing rather than in the statement itself.","major_comments":[{"comment":"Theorem 6.2 is proved only for finite factored spaces (finite index set I, finite factors Ω_i, finite value spaces), as stated in Definition 4.2 and the preamble to Section 4. The motivating examples, however, are continuous or only countably infinite: temperature as a function of kinetic energies of gas particles (Section 1 and Section 2.1), object positions in images, and kinetic energies are real-valued. This is not a cosmetic mismatch. Under the paper's own Definition 6.1, conditional independence of events is defined by P(A∩C)P(B∩C)=P(A∩B∩C)P(C), with the convention that P(C)=0 makes the statement vacuously true. If Z is a non-atomic continuous variable, then P({Z=z})=0 for every factorizing distribution, so X ⊥⊥_P Y | Z holds for all X, Y, P, while structural independence generally fails. For example, on Ω=[0,1]^2 with uniform independent factors and X=Y=Z=U1+U2, the zero-probability convention makes the probabilistic independence vacuous, whereas H(X | z) ∩ H(Y | z) is nonempty for every z in (0,2). Thus the equivalence in Theorem 6.2 cannot be extended to the motivating examples without a measure-theoretic treatment of conditioning. The manuscript should either restrict all claims and examples to finite discrete systems or develop a conditional-independence notion for continuous variables (e.g., via regular conditional distributions) and prove the corresponding theorem.","section":"§1, §2.1, §4 Definition 4.2, §6 Definition 6.1"},{"comment":"The scope statement 'When we speak of variables in this paper, we always mean discrete random variables' is broader than what is proved. Definition 4.2 requires a finite index set and finite factors, and the proof of Lemma 4.9 sums over Val(X) and uses finiteness of Val(X); the interpolation arguments in Appendix C likewise use finiteness of I. Consequently the equivalence in Theorem 6.2 is established only for variables with finite value spaces. Countable discrete variables (e.g., integer-valued functions of the factors) are not covered, so the abstract's claim to generalize the d-separation theorem should be qualified to finite factored spaces, or an extension to countable or measure-theoretic settings should be supplied.","section":"§4 (p. 4), Theorem 6.2"}],"minor_comments":[{"comment":"In the proof of Lemma A.2, the sentence 'we have ω2 = ω' should read 'we have ω2 = ω′'; with the printed equality the subsequent inference is not valid.","section":"Appendix A, Lemma A.2 proof"},{"comment":"In the proof of Lemma 6.5, 'Rλ → 0 as λ → 0+' should be 'Rλ → Q as λ → 0+'; as printed the convergence statement is nonsensical.","section":"Appendix C, proof of Lemma 6.5"},{"comment":"Definition 5.7(2) refers to independence of sets of variables X_W1, X_W2, X_W3 ⊆ X, but Section 4 defines structural and probabilistic independence only for (joint) random variables; the intended reduction to the joint variable (X_w)_{w∈W} should be stated explicitly.","section":"Definition 5.7"},{"comment":"There are spacing and typographical issues in the proof of Proposition 5.6 (e.g., '⇐ ⇒an' and the compressed equivalence chain); these should be cleaned up for readability.","section":"Proposition 5.6 proof"}],"recommendation":"major_revision","confidential_remarks":"The authors are transparent that the first proof of Theorem 6.2 appeared in [8] (Garrabrant 2021), so the main theorem is not entirely new. The present contribution is a self-contained proof, a cleaner formalization, and a detailed comparison with Bayesian networks. For a journal that prioritizes novel theorems, the editor should weigh whether the expositional contribution is sufficient; my recommendation of major revision is driven by the scope mismatch with the continuous motivating examples, not by concerns about the validity of the finite-discrete theorem."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Read this one if you work on causal abstraction, not if you want algorithms. The core theorem (Theorem 6.2) is real, and the proof in Appendix C is careful and self-contained: structural independence equals conditional independence across all factorizing distributions on a finite factored space. The genuinely new pieces are the DAG-to-FSM construction with equivalence of d-separation and structural independence for node variables, the perfect-map expressiveness comparison, and the semigraphoid analysis. Those are worth having. The paper is clearly written and credits the original FSM theorem to Garrabrant's earlier finite factored sets preprint; the novelty burden is real but not hidden.\n\nThe main soft spot is the gap between the formalism and the framing. The definitions and theorem only cover finite discrete variables and finite factored spaces. The motivating examples—temperature as a deterministic function of particle kinetic energies, pixels to object positions—are continuous, and the paper never says the theorem extends to them. The stress test is right about this: under the paper's own zero-probability convention, a continuous version would be vacuous, because every conditioning event like Z=z has probability zero for non-atomic variables. So this is not a cosmetic gap. The abstract's claim about generalizing d-separation needs a finite/discrete qualifier, and the examples need either discretization or an explicit limitation. It is still a sound finite discrete theorem, and the finite case is exactly where d-separation fails due to deterministic functions, so the central idea holds.\n\nMinor issues: some appendix notation is rough, and the structural time and self-reference speculation in Section 7 is not load-bearing. I would not reject over those.\n\nWho should read this: someone building formal foundations for causal representation learning or multiscale abstraction, provided they work in discrete settings or are willing to do the measure-theoretic extension themselves. The paper deserves a serious referee: the proof is technical enough to warrant checking, and the DAG comparison is likely to be used. I would accept it for review and ask the authors to scope the claims honestly in the abstract, especially the continuous examples and the extent to which the main theorem is a reformulation of prior work.","headline":"A careful finite-discrete theorem about structural vs conditional independence, with a real DAG-to-FSM construction; the continuous framing overstates the scope.","tokens_in":32732,"tokens_out":3068,"would_cite":true,"duration_ms":31729,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"One independence criterion is proven exact for deterministic variables.","keywords":["factored space models","structural independence","conditional independence","causal graphs","d-separation","deterministic relationships","levels of abstraction","Bayesian networks"],"falsifier":"A concrete refutation would be a finite factored space with events $A$, $B$, $C$ such that $A$ and $B$ are conditionally independent given $C$ under every factorizing distribution but $H(A \\mid C) \\cap H(B \\mid C)$ is nonempty; the proof of Lemmas C.7 and C.8 says no such triple exists, so exhibiting one would overturn Theorem 6.2.","tokens_in":31665,"feed_emoji":"🎲","tokens_out":6182,"duration_ms":60316,"temperature":0.7,"pith_summary":"Factored space models replace the causal graph with a Cartesian-product sample space whose axes are independent sources of randomness. The paper's central claim is that a purely structural relation—two variables are independent when they depend on disjoint axes—is exactly equivalent to statistical conditional independence holding in every distribution that factorizes over the product. This gives a sound and complete independence criterion that keeps working when variables are deterministic functions of other variables, the situation where d-separation fails and no causal graph can be a perfect map. The result generalizes the classical soundness and completeness theorem for d-separation and is intended as a step toward causal modeling across levels of abstraction, such as temperature as a function of particle kinetic energies.","feed_headline":"Exact independence criterion now covers deterministic variables","feed_subtitle":"Structural independence and statistical independence coincide for every factorizing distribution, generalizing d-separation.","key_machinery":"The central object is the history $H(X \\mid C)$, the unique minimal subset $J$ of the factor index set such that the background variables $U_J$ determine $X$ on $C$ and $J$ disintegrates $C$, meaning $C$ factorizes as $C_J \\times C_{I \\setminus J}$. Structural independence compares these histories: $X \\perp_\\Omega Y \\mid Z$ holds exactly when $H(X \\mid z) \\cap H(Y \\mid z) = \\emptyset$ for every value $z$. The completeness proof introduces the cohistory—the set of factors whose variation never changes $P(A \\mid C)$ under factorizing distributions—and establishes that cohistory equals history, so independence in all factorizing distributions forces the relevant factor sets to be disjoint. A local-to-global lemma treats each conditional-independence statement as a polynomial that vanishes on an open set of distributions and therefore vanishes everywhere, which also yields the paper's strong completeness result.","core_discovery":"The paper proves Theorem 6.2: for random variables $X$, $Y$, $Z$ on a finite factored space $\\Omega = \\times_{i \\in I} \\Omega_i$, $X$ and $Y$ are structurally independent given $Z$ if and only if $X$ and $Y$ are conditionally independent given $Z$ in every probability distribution $P$ that factorizes over $\\Omega$. Structural independence means that for every value $z$, the history $H(X \\mid z)$—the minimal set of factors needed to determine $X$ on the event $Z = z$—is disjoint from $H(Y \\mid z)$. The theorem is proved first for events, with the history shown to contain exactly those factors that are probabilistically relevant to the event given the conditioning set. When the factored space is constructed from a causal directed acyclic graph, structural independence of node variables is equivalent to d-separation, so the classical soundness and completeness theorem for d-separation follows as a special case; additionally, some distributions have a perfect-map factored space but no perfect-map DAG, showing the framework is strictly more expressive.","pith_inferences":["The theorem is stated for finite factored spaces only, so the motivating examples with continuous physical quantities, such as gas-particle kinetic energies, are not covered; a measure-theoretic or analytic extension is needed before those applications are justified.","The equivalence suggests a discovery strategy: search for factorizations whose structural independences match observed conditional independences, in the same way causal discovery searches for DAGs; strong completeness says a match on any open neighborhood of factorizing distributions already certifies the structure.","Structural time may give a formal handle on abstraction hierarchies, since a macro variable's history being a subset of a micro variable's history means the macro variable is determined no later than the micro variable in any process that reveals factors sequentially.","The paper's speculation about self-referencing systems suggests a testable direction: represent a model's summaries of its own internal states as variables and ask whether their histories align with the flow of influence, which could be probed in language-model experiments."],"forward_implications":["Any system modeled as independent sources of randomness gets a distribution-free independence test: two variables are independent in every factorizing distribution exactly when they read off disjoint sources, and this holds even when one variable is a deterministic function of another.","The classical soundness and completeness theorem for d-separation in Bayesian networks becomes a special case, so factored space models inherit the independence guarantees of causal graphs while adding coverage of deterministic relationships.","Factored space models are strictly more expressive than DAGs: there are distributions with a perfect-map factored space but no perfect-map causal graph, so the framework supports independence modeling where graph-based perfect maps do not exist.","Structural time, defined by history inclusion, reproduces the ancestor relation for node variables in a constructed Bayesian network, giving a way to compare variables at different levels of abstraction by their sources of randomness.","Because structural independence fails the intersection axiom while d-separation satisfies it, the framework's independence logic can represent deterministic constraints that causal graphs cannot."],"supporting_citations":[{"why":"Supplies the classical soundness and completeness theorem for d-separation that Theorem 6.2 generalizes.","marker":"[15]"},{"why":"The original finite factored sets framework whose first proof this paper reframes and extends.","marker":"[8]"},{"why":"Graphoids are used to show d-separation satisfies the intersection axiom, which is exactly why it cannot represent deterministic relationships.","marker":"[19]"},{"why":"Source for statistical independence forming a semigraphoid, used in the proof that structural independence is a compositional semigraphoid.","marker":"[18]"},{"why":"The D-separation criterion is compared against and shown insufficient for the paper's deterministic-relationship examples.","marker":"[11]"},{"why":"Defines perfect maps, the notion used to compare the expressiveness of DAGs and factored spaces.","marker":"[25]"},{"why":"The factor-graph framing motivates viewing factors as independent mechanisms in the construction from Bayesian networks.","marker":"[26]"}],"fun_headline_variants":["Independence criterion that handles deterministic variables","Causal independence across abstraction levels","Factored space models beat DAGs for deterministic causality","New theorem: structural independence means statistical independence"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The theorem assumes finiteness throughout—finite index set, finite factors, and finite value spaces—and the proof uses that finiteness directly, so continuous and infinite settings are outside the stated result.","fun_headline_variants_meta":{"raw":{"variants":["Independence criterion that handles deterministic variables","Causal independence across abstraction levels","Factored space models beat DAGs for deterministic causality","New theorem: structural independence means statistical independence"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000222,"raw_usage":{"total_tokens":1470,"prompt_tokens":978,"completion_tokens":492,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":594,"completion_tokens_details":{"reasoning_tokens":437}},"tokens_in":594,"tokens_out":492,"duration_ms":5575,"temperature":1.0,"reasoning_tokens":437,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T23:19:02.052725+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A concrete refutation would be a finite factored space with events $A$, $B$, $C$ such that $A$ and $B$ are conditionally independent given $C$ under every factorizing distribution but $H(A \\mid C) \\cap H(B \\mid C)$ is nonempty; the proof of Lemmas C.7 and C.8 says no such triple exists, so exhibiting one would overturn Theorem 6.2.","supporting_citations":[{"cited_title":"Probabilistic graphical models: principles and techniques","cited_arxiv_id":null,"evidence_quote":"Supplies the classical soundness and completeness theorem for d-separation that Theorem 6.2 generalizes."},{"cited_title":"Temporal Inference with Finite Factored Sets","cited_arxiv_id":"2109.11513","evidence_quote":"The original finite factored sets framework whose first proof this paper reframes and extends."},{"cited_title":"Graphoids: Graph-Based Logic for Reasoning about Relevance Relations or When would x tell you more about y if you already know z?","cited_arxiv_id":null,"evidence_quote":"Graphoids are used to show d-separation satisfies the intersection axiom, which is exactly why it cannot represent deterministic relationships."},{"cited_title":"Probabilistic Reasoning in Intelligent Systems: Networks of Plausible Inference","cited_arxiv_id":null,"evidence_quote":"Source for statistical independence forming a semigraphoid, used in the proof that structural independence is a compositional semigraphoid."},{"cited_title":"Identifying independence in Bayesian networks","cited_arxiv_id":null,"evidence_quote":"The D-separation criterion is compared against and shown insufficient for the paper's deterministic-relationship examples."},{"cited_title":"Causation, prediction, and search","cited_arxiv_id":null,"evidence_quote":"Defines perfect maps, the notion used to compare the expressiveness of DAGs and factored spaces."},{"cited_title":"Data mining: practical machine learning tools and techniques with Java implemen- tations","cited_arxiv_id":null,"evidence_quote":"The factor-graph framing motivates viewing factors as independent mechanisms in the construction from Bayesian networks."}],"review_version":1}