REVIEW 2 major objections 5 minor 13 references
A Theory of Structural Independence
T0 review · 2 major / 5 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read The paper proves that an independence is forced by the independence of background variables, in every compatible product distribution, exactly when the two variables' histories are almost surely disjoint.
desk verdict The infinite-index generalization of structural independence is a real idea, but the completeness proof has a demonstrable error in the interpolation lemma, so the main theorem is unproven as written. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the history $\mathcal{H}(X\mid Z)$, a $\sigma(Z)$-measurable random index set defined as the almost surely minimal $J(\omega)\subseteq I$ such that $\sigma(X)\subseteq\sigma(U_J,Z)$ and $U_J\perp_P U_{J^c}\mid Z$ for all $P\in\Delta^\times$. Its dual, the irrelevance $\mathcal{I}(X\mid Z)$, is the almost surely maximal set of indices whose one-coordinate perturbations cannot change the conditional distribution of $X$ given $Z$. The proofs of completeness and of the density and rarity results run through the equality $\mathcal{H}=\mathcal{I}$ and through a polynomial interpolation argument (Theorem 6.2.1) that connects equivalent product measures by a curve of product densities; the counterexample in Section 8 shows why the finite-case rectangle criterion for disintegration cannot be used in general.
What would settle it
Check the finite-dimensional curve in Theorem 6.2.1 directly: the displayed polynomial $p(\lambda,x)=\prod_{n=0}^{m}(\lambda x_n+\lambda)$ vanishes at $\lambda=0$, and at $\lambda=1$ the proposed density $\varphi_\lambda$ has total mass $2^{m+1}$ rather than $1$, so the displayed construction is not a probability density; finding a corrected curve that is a probability density for every $\lambda$ and keeps the conditional-expectation inequality alive for a dense set would repair the proof, while proving no such curve exists would show that the completeness direction is unsupported.
Extended reading notes
Core claim
Theorem 6.4.1 establishes the fundamental theorem of structural independence: for an independent family $U=(U_i)_{i\in I}$ on $(\Omega,\mathcal{A},\mathbb{P})$, arbitrary $\sigma(U)$-measurable random elements $X,Y,Z$, and $\Delta^\times = \{P : P\sim\mathbb{P},\ U\text{ independent under }P\}$, one has $X\perp_P Y\mid Z$ for all $P\in\Delta^\times$ if and only if $\mathcal{H}(X\mid Z)\cap\mathcal{H}(Y\mid Z)=\varnothing$ $\mathbb{P}$-almost surely. Here $\mathcal{H}(X\mid Z)$ is the almost surely minimal $\sigma(Z)$-measurable random index set that generates $X$ from $U$ given $Z$ and disintegrates $Z$, meaning $U_{\mathcal{H}}$ and $U_{\mathcal{H}^c}$ are conditionally independent given $Z$. The paper proves both directions: soundness follows directly from generation, while completeness is obtained by introducing the dual random index set of irrelevance $\mathcal{I}(X\mid Z)$, proving $\mathcal{H}=\mathcal{I}$, and showing that overlapping histories force a dense set of distributions that violate the independence.
Load-bearing premise
The completeness half of the fundamental theorem depends on the interpolation lemma in Theorem 6.2.1, which claims that any two equivalent product measures can be joined by a continuous curve of product probability densities while preserving a conditional-probability inequality; if this lemma cannot be repaired, the proof that overlapping histories force failure of independence in some distribution is incomplete.
Editorial extensions
If this is right
- Distribution-free conditional independence for variables built from $U$ reduces to checking whether two random index sets are almost surely disjoint, making a quantified-over-all-distributions statement purely combinatorial.
- Structural independence is a compositional semigraphoid: it satisfies symmetry, decomposition, weak union, contraction, and composition, so pairwise structural independence of a vector implies joint structural independence.
- In a structural causal model, structural independence is a $d$-separation-like criterion for arbitrary functions of exogenous variables; in the toy setting, $V_1\perp (V_1+V_2)$ in all compatible models implies $V_1$ is an ancestor of $V_2$, and a parent when $V=(V_1,V_2)$.
- If $X$ and $Y$ are not structurally independent given $Z$, then the set of distributions in $\Delta^\times$ where $X\perp_P Y\mid Z$ holds is closed and nowhere dense in the $d_1$ metric, so accidental, non-structural independences are topologically rare.
- The history map is the almost surely unique maximal map satisfying the four desiderata, so it provides a canonical invariant for the structure of an independent family rather than a choice-dependent construction.
Reading between the lines
- The Section 8 counterexample implies that in continuous settings structural independence cannot be read off from the atoms of the conditioning variable; a practical implementation would need reference measures and conditional expectations, not support geometry.
- If the interpolation lemma is repaired, the history and irrelevance duality suggests a template for non-product constraint classes: define structural dependence relative to any set of distributions closed under single-coordinate perturbations, and the same minimal and maximal random index set argument might carry through.
- The author leaves implicit a statistical reading: because non-structural independences are nowhere dense, random perturbations of an observed distribution should destroy them, which could be turned into a finite-sample test that distinguishes structural from accidental independence in causal discovery.
- The equality $\mathcal{H}=\mathcal{I}$ indicates that minimal generating set and maximal irrelevant set are two faces of the same quantity; this duality may transfer to other settings where one wants to identify the footprint of a set of variables on an independent noise family.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper develops a theory of structural independence for an independent family U=(U_i)_{i∈I}. For σ(U)-measurable random elements X, Y, Z, it defines the history H(X|Z), a random index set measuring dependence on U given Z, and claims Theorem 6.4.1: X and Y are independent given Z in every distribution P that renders U independent and is mutually absolutely continuous with respect to the reference measure if and only if H(X|Z)∩H(Y|Z)=∅ almost surely. Sections 4–5 build the machinery of random index sets, disintegration, and generation; Section 6 introduces the dual notion of irrelevance and attempts to prove completeness; Section 7 derives semigraphoid properties and a maximality/uniqueness theorem for the history; Sections 8–9 present a counterexample about disintegration and an application to structural causal models.
Significance. If the completeness proof can be repaired, this is a valuable contribution: it gives a clean random-index-set calculus for distribution-free conditional independence, a simple soundness direction, a uniqueness/maximality theorem, and a semigraphoid structure. The use of Kakutani's theorem in Appendix A is appropriate, and the paper is largely self-contained. The central characterization is attractive. However, the completeness direction is currently unsupported because the key approximation argument in Section 6.2 is invalid as written; this blocks the main theorem.
major comments (2)
- [§6.2, Theorem 6.2.1] The interpolation polynomial p(λ,x)=∏_{n=0}^m(λ x_n+λ) does not define the claimed curve. Since p(0,x)=0, one has R_0=0·P, not P, and R'_0=0·P, not Q; moreover ∫φ_λ dP=(2λ)^{m+1}, so R_λ is not a probability measure for λ≠1/2. Consequently the assertion p'(0,ω)≠0 on C is impossible: with this p, p'(0,ω)=0 identically. The approximating sequence (P_n,Q_n) used in Theorem 6.2.2 is therefore not constructed, and the completeness direction of Theorems 6.3.5 and 6.4.1 is unproven as written. A convex interpolation such as p(λ,x)=∏_{n=0}^m((1−λ)+λ x_n) would repair the endpoint and normalization, but this needs to be carried out explicitly and the subsequent polynomial argument re-examined.
- [§6.2, finite-dimensional step] The equivalence between R_λ(A|Z)=R'_λ(A|Z) and p'(λ,ω)=0 uses the identity E(p(λ,Φ)1_A|Z)=p(λ,(E(φ_n1_A|Z))_n), which is not an instance of linearity of conditional expectation. It would require the factors φ_n to be conditionally independent given Z, which is not assumed and is false in general for σ(U)-measurable Z. As written, the polynomial method does not locate the set where the conditional probabilities differ. The proof of Theorem 6.4.3 in §6.4 appears to rely on the same problematic polynomial reduction.
minor comments (5)
- [§5, Lemma 5.9] The statement says that ⋂ J_s 'generates Z given X', and the proof contains the inclusion σ(U_{J_s},Z) ⊆ σ(U_J,Z), which is opposite to the inclusion one would expect from J⊆J_s; this needs correction or clarification.
- [§6.1, Theorem 6.1.7] In the final sentence, 'Q(A|Z)' should read 'Q(B|Z)'.
- [§6.2] The symbol p' is used both for the polynomial difference and for a derivative in the phrase p'(0,ω); this is confusing and should be renamed.
- [§6.4, Theorem 6.4.3] The displayed equivalence for A⫫_{R_λ} B|Z has unbalanced parentheses and stray plus signs; as printed it is not correct. The intended algebraic identity should be written out cleanly.
- [§9, Lemma 9.4] The step 'i=ℋ(U_i)⊆ℋ(V_i)' conflates an index with a random index set; the proof of the ancestor conclusion is only sketched and should be made precise.
Circularity Check
No circularity: the history is constructed from first principles via minimal generating random index sets, and the fundamental theorem is proved from that construction rather than assumed.
full rationale
The derivation chain is self-contained. The paper defines the history H(X|Z) as the almost surely minimal random index set that generates X given Z (Definition 5.11), where generation requires sigma(X) subset sigma(U_J, Z) and that J disintegrates Z, i.e. U_J is conditionally independent of its complement given Z for all P in the reference class (Definition 5.2). Existence is proved by closure under intersections and chains (Lemmas 5.5, 5.7, 5.8, 5.9 and Theorem 5.10), without invoking the target characterization. Soundness (Theorem 6.1) follows directly from the generation property. Completeness (Theorem 6.3.5) is obtained through the dual notion of irrelevance, with the duality H = I established in Theorem 6.3.4; none of these steps reduces to the desired iff statement. The Desiderata 5.1 are stated as a specification before the construction, but the proof of Theorem 6.4.1 does not use them as an input; Theorem 7.16 shows maximality and uniqueness under those desiderata after the fundamental theorem has been proved. The citations to the author's prior finite theory [5], [6], and [8] are contextual or motivational; Section 3 explicitly says proofs of the finite theory follow from the general theory, so prior work is not load-bearing. The skeptical concern about Theorem 6.2.1 -- that the interpolation polynomial p(lambda,x) = prod(lambda x_n + lambda) is zero at lambda = 0 and not normalized -- is a technical correctness gap in the completeness argument, not a circularity: the target result is not assumed as an input, and the gap is in principle repairable by a different interpolation. Unsupported conjectures in Section 10 are clearly labeled as conjectures or future work and are not used in the central derivation. Therefore, on the circularity axis, there is no significant circularity.
Assumptions & free parameters
assumptions (5)
- standard math Kakutani's theorem on equivalence of infinite product measures (Theorem A.5)
- standard math Zorn's lemma applied to random index sets modulo almost sure equality (Lemma 4.2.5)
- domain assumption All σ-algebras are completed with respect to the reference measure ℙ (Notation 4.3.1)
- domain assumption Random elements X, Y, Z are σ(U)-measurable (Section 5)
- domain assumption Causal models are acyclic and purely probabilistic in the application (Section 9, Definition 9.1, Lemma 9.4)
invented entities (2)
-
History H(X|Z)
independent evidence
-
Random index set of irrelevance I(X|Z)
independent evidence
Cite this review
Pith. "Pith review of A Theory of Structural Independence." pith.science (2026). https://pith.science/paper/QEHW5GUS
@misc{pith2026241200847,
author = {Pith},
title = {Pith review of: A Theory of Structural Independence},
year = {2026},
howpublished = {\url{https://pith.science/paper/QEHW5GUS}},
note = {Machine review of arXiv:2412.00847}
}
abstract
Structural independence is the (conditional) independence that arises from the structure rather than the precise numerical values of a distribution. We develop this concept and relate it to $d$-separation and structural causal models. Formally, let $U = (U_i)_{i \in I}$ be an independent family of random elements on a probability space $(\Omega, \mathcal{A}, \mathbb{P})$. Let $X$, $Y$, and $Z$ be arbitrary $\sigma(U)$-measurable random elements. We characterize all independences $X \perp Y \mid Z$ implied by the independence of $U$ and call these independences \textit{structural}. Formally, these are the independences which hold in all probability measures $P$ that render $U$ independent and are absolutely continuous with respect to $\mathbb{P}$; i.e., for all such $P$, it must hold that $X \perp_P Y \mid Z$. We introduce the history $\mathcal{H}(X \mid Z) : \Omega \to \mathcal{P}(I)$, a combinatorial object that measures the dependence of $X$ on $U_i$ for each $i \in I$ given $Z$. The independence of $X$ and $Y$ given $Z$ is implied by the independence of $U$ if and only if $\mathcal{H}(X \mid Z) \cap \mathcal{H}(Y \mid Z) = \emptyset$ almost surely with respect to $\mathbb{P}$. Finally, we apply this $d$-separation-like criterion in structural causal models to discover a causal direction in a toy setting.
Figures
Reference graph
Works this paper leans on
-
[1]
On equivalence of infinite product measures,
S. Kakutani, “On equivalence of infinite product measures,” Annals of Mathematics, vol. 49, no. 1, pp. 214–224, 1948
work page 1948
-
[2]
Pearl, Causality
J. Pearl, Causality. Cambridge university press, 2009
2009
-
[3]
Causal networks: Semantics and expressiveness,
T. Verma and J. Pearl, “Causal networks: Semantics and expressiveness,” Machine intelligence and pattern recognition, vol. 9. Elsevier, pp. 69–76, 1990
work page 1990
-
[4]
An introduction to causal discovery,
M. Huber, “An introduction to causal discovery,” Swiss Journal of Economics and Statistics, vol. 160, no. 1, p. 14, 2024
work page 2024
-
[5]
Factored Space Models: Towards Causality between levels of abstraction ,
S. Garrabrant, M. G. Mayer, M. Wache, L. Lang, S. Eisenstat, and H. Dell, “Factored Space Models: Towards Causality between levels of abstraction ,” submitted to the Journal of Causal Inference
-
[6]
Temporal Inference with Finite Factored Sets,
S. Garrabrant, “Temporal Inference with Finite Factored Sets,” arXiv preprint arXiv:2109.11513, 2021
arXiv 2021
-
[7]
Identifying independence in Bayesian networks,
D. Geiger, T. Verma, and J. Pearl, “Identifying independence in Bayesian networks,” Networks, vol. 20, no. 5, pp. 507–534, 1990
work page 1990
-
[8]
Causality with Deterministic Relationships,
M. G. Mayer, “Causality with Deterministic Relationships,” Bachelor Thesis, University of Kaiserslautern, 2023
work page 2023
Show all 13 references
-
[9]
Graphoids: Graph-Based Logic for Reasoning about Relevance Relations or When would x tell you more about y if you already know z?,
J. Pearl and A. Paz, “Graphoids: Graph-Based Logic for Reasoning about Relevance Relations or When would x tell you more about y if you already know z?,” Probabilistic and Causal Inference: The Works of Judea Pearl. pp. 189–200, 1985
1985
-
[10]
R. L. Wheeden and A. Zygmund, Measure and integral, vol. 26. Dekker New York, 1977
1977
-
[11]
Intervention and Conditioning in Causal Bayesian Networks,
S. Galhotra and J. Y. Halpern, “Intervention and Conditioning in Causal Bayesian Networks,” in Advances in Neural Information Processing Systems , A. Globerson, L. Mackey, D. Belgrave, A. Fan, U. Paquet, J. Tomczak, and C. Zhang, Eds., Curran Associates, Inc., 2024, pp. 89019–...
2024
-
[12]
Qualitative Mechanism Independence,
O. E. Richardson, S. J. Peters, and J. Halpern, “Qualitative Mechanism Independence,” in The Thirty-eighth Annual Conference on Neural Information Processing Systems,
-
[13]
Durrett, Probability: theory and examples
R. Durrett, Probability: theory and examples. Cambridge university press, 2019. 39
2019
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.