Pith. sign in

REVIEW 2 major objections 5 minor 13 references

A Theory of Structural Independence

T0 review · 2 major / 5 minor · reviewed 2026-08-12 · deepseek-v4-flash

Pith's one-line read The paper proves that an independence is forced by the independence of background variables, in every compatible product distribution, exactly when the two variables' histories are almost surely disjoint.

desk verdict The infinite-index generalization of structural independence is a real idea, but the completeness proof has a demonstrable error in the interpolation lemma, so the main theorem is unproven as written. read the letter →

arxiv 2412.00847 v2 pith:QEHW5GUS submitted 2024-12-01 math.PR

classification math.PR MSC 60A99
keywords structuralindependenceconditionalrandomindexsethistoryd-separationcausalmodelsemigraphoidinfiniteproductmeasures
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper asks which conditional independences are forced by the bare structural assumption that a family $U=(U_i)_{i\in I}$ of background random elements is independent, regardless of the numerical distribution they receive. It proves that for any $\sigma(U)$-measurable random elements $X,Y,Z$, the independence $X\perp_P Y\mid Z$ holds for every probability $P$ that makes $U$ independent and is mutually absolutely continuous with the reference measure exactly when the histories $\mathcal{H}(X\mid Z)$ and $\mathcal{H}(Y\mid Z)$ are almost surely disjoint. The history is a random index set: it picks, for each outcome, the minimal set of coordinates of $U$ on which $X$ still depends after conditioning on $Z$. This gives a $d$-separation-like criterion for structural causal models and settles which independences are structural rather than accidental. A reader should care because it turns a quantification over all distributions into a combinatorial check on random sets.

What carries the argument

The central object is the history $\mathcal{H}(X\mid Z)$, a $\sigma(Z)$-measurable random index set defined as the almost surely minimal $J(\omega)\subseteq I$ such that $\sigma(X)\subseteq\sigma(U_J,Z)$ and $U_J\perp_P U_{J^c}\mid Z$ for all $P\in\Delta^\times$. Its dual, the irrelevance $\mathcal{I}(X\mid Z)$, is the almost surely maximal set of indices whose one-coordinate perturbations cannot change the conditional distribution of $X$ given $Z$. The proofs of completeness and of the density and rarity results run through the equality $\mathcal{H}=\mathcal{I}$ and through a polynomial interpolation argument (Theorem 6.2.1) that connects equivalent product measures by a curve of product densities; the counterexample in Section 8 shows why the finite-case rectangle criterion for disintegration cannot be used in general.

What would settle it

Check the finite-dimensional curve in Theorem 6.2.1 directly: the displayed polynomial $p(\lambda,x)=\prod_{n=0}^{m}(\lambda x_n+\lambda)$ vanishes at $\lambda=0$, and at $\lambda=1$ the proposed density $\varphi_\lambda$ has total mass $2^{m+1}$ rather than $1$, so the displayed construction is not a probability density; finding a corrected curve that is a probability density for every $\lambda$ and keeps the conditional-expectation inequality alive for a dense set would repair the proof, while proving no such curve exists would show that the completeness direction is unsupported.

Watch

Extended reading notes

Core claim

Theorem 6.4.1 establishes the fundamental theorem of structural independence: for an independent family $U=(U_i)_{i\in I}$ on $(\Omega,\mathcal{A},\mathbb{P})$, arbitrary $\sigma(U)$-measurable random elements $X,Y,Z$, and $\Delta^\times = \{P : P\sim\mathbb{P},\ U\text{ independent under }P\}$, one has $X\perp_P Y\mid Z$ for all $P\in\Delta^\times$ if and only if $\mathcal{H}(X\mid Z)\cap\mathcal{H}(Y\mid Z)=\varnothing$ $\mathbb{P}$-almost surely. Here $\mathcal{H}(X\mid Z)$ is the almost surely minimal $\sigma(Z)$-measurable random index set that generates $X$ from $U$ given $Z$ and disintegrates $Z$, meaning $U_{\mathcal{H}}$ and $U_{\mathcal{H}^c}$ are conditionally independent given $Z$. The paper proves both directions: soundness follows directly from generation, while completeness is obtained by introducing the dual random index set of irrelevance $\mathcal{I}(X\mid Z)$, proving $\mathcal{H}=\mathcal{I}$, and showing that overlapping histories force a dense set of distributions that violate the independence.

Load-bearing premise

The completeness half of the fundamental theorem depends on the interpolation lemma in Theorem 6.2.1, which claims that any two equivalent product measures can be joined by a continuous curve of product probability densities while preserving a conditional-probability inequality; if this lemma cannot be repaired, the proof that overlapping histories force failure of independence in some distribution is incomplete.

Editorial extensions

If this is right

  • Distribution-free conditional independence for variables built from $U$ reduces to checking whether two random index sets are almost surely disjoint, making a quantified-over-all-distributions statement purely combinatorial.
  • Structural independence is a compositional semigraphoid: it satisfies symmetry, decomposition, weak union, contraction, and composition, so pairwise structural independence of a vector implies joint structural independence.
  • In a structural causal model, structural independence is a $d$-separation-like criterion for arbitrary functions of exogenous variables; in the toy setting, $V_1\perp (V_1+V_2)$ in all compatible models implies $V_1$ is an ancestor of $V_2$, and a parent when $V=(V_1,V_2)$.
  • If $X$ and $Y$ are not structurally independent given $Z$, then the set of distributions in $\Delta^\times$ where $X\perp_P Y\mid Z$ holds is closed and nowhere dense in the $d_1$ metric, so accidental, non-structural independences are topologically rare.
  • The history map is the almost surely unique maximal map satisfying the four desiderata, so it provides a canonical invariant for the structure of an independent family rather than a choice-dependent construction.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The Section 8 counterexample implies that in continuous settings structural independence cannot be read off from the atoms of the conditioning variable; a practical implementation would need reference measures and conditional expectations, not support geometry.
  • If the interpolation lemma is repaired, the history and irrelevance duality suggests a template for non-product constraint classes: define structural dependence relative to any set of distributions closed under single-coordinate perturbations, and the same minimal and maximal random index set argument might carry through.
  • The author leaves implicit a statistical reading: because non-structural independences are nowhere dense, random perturbations of an observed distribution should destroy them, which could be turned into a finite-sample test that distinguishes structural from accidental independence in causal discovery.
  • The equality $\mathcal{H}=\mathcal{I}$ indicates that minimal generating set and maximal irrelevant set are two faces of the same quantity; this duality may transfer to other settings where one wants to identify the footprint of a set of variables on an independent noise family.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 5 minor

Summary. The paper develops a theory of structural independence for an independent family U=(U_i)_{i∈I}. For σ(U)-measurable random elements X, Y, Z, it defines the history H(X|Z), a random index set measuring dependence on U given Z, and claims Theorem 6.4.1: X and Y are independent given Z in every distribution P that renders U independent and is mutually absolutely continuous with respect to the reference measure if and only if H(X|Z)∩H(Y|Z)=∅ almost surely. Sections 4–5 build the machinery of random index sets, disintegration, and generation; Section 6 introduces the dual notion of irrelevance and attempts to prove completeness; Section 7 derives semigraphoid properties and a maximality/uniqueness theorem for the history; Sections 8–9 present a counterexample about disintegration and an application to structural causal models.

Significance. If the completeness proof can be repaired, this is a valuable contribution: it gives a clean random-index-set calculus for distribution-free conditional independence, a simple soundness direction, a uniqueness/maximality theorem, and a semigraphoid structure. The use of Kakutani's theorem in Appendix A is appropriate, and the paper is largely self-contained. The central characterization is attractive. However, the completeness direction is currently unsupported because the key approximation argument in Section 6.2 is invalid as written; this blocks the main theorem.

major comments (2)
  1. [§6.2, Theorem 6.2.1] The interpolation polynomial p(λ,x)=∏_{n=0}^m(λ x_n+λ) does not define the claimed curve. Since p(0,x)=0, one has R_0=0·P, not P, and R'_0=0·P, not Q; moreover ∫φ_λ dP=(2λ)^{m+1}, so R_λ is not a probability measure for λ≠1/2. Consequently the assertion p'(0,ω)≠0 on C is impossible: with this p, p'(0,ω)=0 identically. The approximating sequence (P_n,Q_n) used in Theorem 6.2.2 is therefore not constructed, and the completeness direction of Theorems 6.3.5 and 6.4.1 is unproven as written. A convex interpolation such as p(λ,x)=∏_{n=0}^m((1−λ)+λ x_n) would repair the endpoint and normalization, but this needs to be carried out explicitly and the subsequent polynomial argument re-examined.
  2. [§6.2, finite-dimensional step] The equivalence between R_λ(A|Z)=R'_λ(A|Z) and p'(λ,ω)=0 uses the identity E(p(λ,Φ)1_A|Z)=p(λ,(E(φ_n1_A|Z))_n), which is not an instance of linearity of conditional expectation. It would require the factors φ_n to be conditionally independent given Z, which is not assumed and is false in general for σ(U)-measurable Z. As written, the polynomial method does not locate the set where the conditional probabilities differ. The proof of Theorem 6.4.3 in §6.4 appears to rely on the same problematic polynomial reduction.
minor comments (5)
  1. [§5, Lemma 5.9] The statement says that ⋂ J_s 'generates Z given X', and the proof contains the inclusion σ(U_{J_s},Z) ⊆ σ(U_J,Z), which is opposite to the inclusion one would expect from J⊆J_s; this needs correction or clarification.
  2. [§6.1, Theorem 6.1.7] In the final sentence, 'Q(A|Z)' should read 'Q(B|Z)'.
  3. [§6.2] The symbol p' is used both for the polynomial difference and for a derivative in the phrase p'(0,ω); this is confusing and should be renamed.
  4. [§6.4, Theorem 6.4.3] The displayed equivalence for A⫫_{R_λ} B|Z has unbalanced parentheses and stray plus signs; as printed it is not correct. The intended algebraic identity should be written out cleanly.
  5. [§9, Lemma 9.4] The step 'i=ℋ(U_i)⊆ℋ(V_i)' conflates an index with a random index set; the proof of the ancestor conclusion is only sketched and should be made precise.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the history is constructed from first principles via minimal generating random index sets, and the fundamental theorem is proved from that construction rather than assumed.

full rationale

The derivation chain is self-contained. The paper defines the history H(X|Z) as the almost surely minimal random index set that generates X given Z (Definition 5.11), where generation requires sigma(X) subset sigma(U_J, Z) and that J disintegrates Z, i.e. U_J is conditionally independent of its complement given Z for all P in the reference class (Definition 5.2). Existence is proved by closure under intersections and chains (Lemmas 5.5, 5.7, 5.8, 5.9 and Theorem 5.10), without invoking the target characterization. Soundness (Theorem 6.1) follows directly from the generation property. Completeness (Theorem 6.3.5) is obtained through the dual notion of irrelevance, with the duality H = I established in Theorem 6.3.4; none of these steps reduces to the desired iff statement. The Desiderata 5.1 are stated as a specification before the construction, but the proof of Theorem 6.4.1 does not use them as an input; Theorem 7.16 shows maximality and uniqueness under those desiderata after the fundamental theorem has been proved. The citations to the author's prior finite theory [5], [6], and [8] are contextual or motivational; Section 3 explicitly says proofs of the finite theory follow from the general theory, so prior work is not load-bearing. The skeptical concern about Theorem 6.2.1 -- that the interpolation polynomial p(lambda,x) = prod(lambda x_n + lambda) is zero at lambda = 0 and not normalized -- is a technical correctness gap in the completeness argument, not a circularity: the target result is not assumed as an input, and the gap is in principle repairable by a different interpolation. Unsupported conjectures in Section 10 are clearly labeled as conjectures or future work and are not used in the central derivation. Therefore, on the circularity axis, there is no significant circularity.

Assumptions & free parameters 0 free parameters · 5 assumptions · 2 invented entities

The paper introduces two constructed objects (history and irrelevance) rather than fitted parameters. There are no free parameters. The main axioms are standard measure theory, Zorn's lemma, completed σ-algebras, and the domain restrictions that X,Y,Z are σ(U)-measurable and that causal models are acyclic/purely probabilistic in the application.

assumptions (5)
  • standard math Kakutani's theorem on equivalence of infinite product measures (Theorem A.5)
    Used in Lemma A.8 to factor densities of equivalent product measures into one-dimensional components; essential for the structure of Δ×.
  • standard math Zorn's lemma applied to random index sets modulo almost sure equality (Lemma 4.2.5)
    Used to prove existence and uniqueness of minimal generating random index sets and hence the history.
  • domain assumption All σ-algebras are completed with respect to the reference measure ℙ (Notation 4.3.1)
    The theory operates with completed σ-algebras and almost sure relations; this is a convention throughout.
  • domain assumption Random elements X, Y, Z are σ(U)-measurable (Section 5)
    The fundamental theorem only applies to random elements determined by the independent family U.
  • domain assumption Causal models are acyclic and purely probabilistic in the application (Section 9, Definition 9.1, Lemma 9.4)
    Lemma 9.4 assumes V_i is not independent of U_i ('purely probabilistic') and the graph is acyclic to conclude ancestor relations from history inclusions.
invented entities (2)
  • History H(X|Z) independent evidence
    purpose: A σ(Z)-measurable random index set measuring the dependence of X on the independent family U given Z; used to define structural independence.
    The fundamental theorem provides an external handle: disjointness of histories is equivalent to conditional independence in every P ∈ Δ×, which is testable in principle.
  • Random index set of irrelevance I(X|Z) independent evidence
    purpose: Dual object to the history, used in the completeness proof to prove that non-irrelevance forces independence to fail for some P.
    Shown to equal the history complement (Theorem 6.3.4), inheriting the same external characterization.

how reviews work

0 comments
Cite this review

Pith. "Pith review of A Theory of Structural Independence." pith.science (2026). https://pith.science/paper/QEHW5GUS

@misc{pith2026241200847,
  author       = {Pith},
  title        = {Pith review of: A Theory of Structural Independence},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/QEHW5GUS}},
  note         = {Machine review of arXiv:2412.00847}
}
abstract

Structural independence is the (conditional) independence that arises from the structure rather than the precise numerical values of a distribution. We develop this concept and relate it to $d$-separation and structural causal models. Formally, let $U = (U_i)_{i \in I}$ be an independent family of random elements on a probability space $(\Omega, \mathcal{A}, \mathbb{P})$. Let $X$, $Y$, and $Z$ be arbitrary $\sigma(U)$-measurable random elements. We characterize all independences $X \perp Y \mid Z$ implied by the independence of $U$ and call these independences \textit{structural}. Formally, these are the independences which hold in all probability measures $P$ that render $U$ independent and are absolutely continuous with respect to $\mathbb{P}$; i.e., for all such $P$, it must hold that $X \perp_P Y \mid Z$. We introduce the history $\mathcal{H}(X \mid Z) : \Omega \to \mathcal{P}(I)$, a combinatorial object that measures the dependence of $X$ on $U_i$ for each $i \in I$ given $Z$. The independence of $X$ and $Y$ given $Z$ is implied by the independence of $U$ if and only if $\mathcal{H}(X \mid Z) \cap \mathcal{H}(Y \mid Z) = \emptyset$ almost surely with respect to $\mathbb{P}$. Finally, we apply this $d$-separation-like criterion in structural causal models to discover a causal direction in a toy setting.

Figures

Figures reproduced from arXiv: 2412.00847 by the authors.

Figure 1
Figure 1. An illustration of Ω = Ω1 × Ω2 We now construct the random element 𝑍 : Ω → 𝑆 2 on which we will condition. For 𝑖, 𝑗 ∈ 𝐼, set 𝑆𝑖𝑗 = 𝑆𝑖 × 𝑆𝑗 . Then Ω = ⋃𝑖,𝑗∈𝐼 𝑆𝑖𝑗 . Therefore it suffices to define 𝑍 on each of 𝑆𝑖𝑗 , we write 𝑍𝑖𝑗 for 𝑍|𝑆𝑖𝑗 . Let 𝛼 ∈ (0, 1) and 𝛽 ∈ (0, 1). Let 𝐸 = {(𝑎, 𝑏) ∈ 𝑆 2 : 𝑎 + 𝛽 ⋅ 𝑏, 𝛼𝛼 + 𝛽 ∈ [0, 1]}. Let for (𝑎, 𝑏) ∈ 𝐸 𝑍 −1 12 ( 𝑎 𝑏 ) = ( 1 𝛼 0 1 )( 𝑎 𝑏 ) = ( 𝑎 𝛼⋅𝑎+𝑏 ) ∈ 𝑆12 𝑍 −1 11 ( 𝑎 𝑏 ) = ( … view at source ↗
Figure 2
Figure 2. An illustration of 𝑍 for 𝛼 = 𝛽 = 1 4 in the region 𝑍 −1 [0, 3 4 ] 2 . The cells correspond to the partition {𝑍 ∈ [𝑎 − 𝜀, 𝑎 + 𝜀) × [𝑏 − 𝜀, 𝑏 + 𝜀) : 𝑎, 𝑏 ∈ {𝜀(𝑛 + 1 2 ) : 𝑛 ∈ {0, 1, 2}}} for 𝜀 = 1 4 . The numbers inscribed in the cells illustrate which cells are in the same part of the partition. Note that 𝑍 has rectangular atoms, since for (𝑎, 𝑏) ∈ 𝐸, 𝑍 −1 (𝑎, 𝑏) = {𝑎, 𝑎 + 𝛽 ⋅ 𝑏} × {𝑏, 𝛼 ⋅ 𝑎 + 𝑏} and otherwise, 𝑍 −1 … view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

13 extracted references · 11 canonical work pages

  1. [1]

    On equivalence of infinite product measures,

    S. Kakutani, “On equivalence of infinite product measures,” Annals of Mathematics, vol. 49, no. 1, pp. 214–224, 1948

  2. [2]

    Pearl, Causality

    J. Pearl, Causality. Cambridge university press, 2009

  3. [3]

    Causal networks: Semantics and expressiveness,

    T. Verma and J. Pearl, “Causal networks: Semantics and expressiveness,” Machine intelligence and pattern recognition, vol. 9. Elsevier, pp. 69–76, 1990

  4. [4]

    An introduction to causal discovery,

    M. Huber, “An introduction to causal discovery,” Swiss Journal of Economics and Statistics, vol. 160, no. 1, p. 14, 2024

  5. [5]

    Factored Space Models: Towards Causality between levels of abstraction ,

    S. Garrabrant, M. G. Mayer, M. Wache, L. Lang, S. Eisenstat, and H. Dell, “Factored Space Models: Towards Causality between levels of abstraction ,” submitted to the Journal of Causal Inference

  6. [6]

    Temporal Inference with Finite Factored Sets,

    S. Garrabrant, “Temporal Inference with Finite Factored Sets,” arXiv preprint arXiv:2109.11513, 2021

  7. [7]

    Identifying independence in Bayesian networks,

    D. Geiger, T. Verma, and J. Pearl, “Identifying independence in Bayesian networks,” Networks, vol. 20, no. 5, pp. 507–534, 1990

  8. [8]

    Causality with Deterministic Relationships,

    M. G. Mayer, “Causality with Deterministic Relationships,” Bachelor Thesis, University of Kaiserslautern, 2023

Show all 13 references
  1. [9]

    Graphoids: Graph-Based Logic for Reasoning about Relevance Relations or When would x tell you more about y if you already know z?,

    J. Pearl and A. Paz, “Graphoids: Graph-Based Logic for Reasoning about Relevance Relations or When would x tell you more about y if you already know z?,” Probabilistic and Causal Inference: The Works of Judea Pearl. pp. 189–200, 1985

  2. [10]

    R. L. Wheeden and A. Zygmund, Measure and integral, vol. 26. Dekker New York, 1977

  3. [11]

    Intervention and Conditioning in Causal Bayesian Networks,

    S. Galhotra and J. Y. Halpern, “Intervention and Conditioning in Causal Bayesian Networks,” in Advances in Neural Information Processing Systems , A. Globerson, L. Mackey, D. Belgrave, A. Fan, U. Paquet, J. Tomczak, and C. Zhang, Eds., Curran Associates, Inc., 2024, pp. 89019–...

  4. [12]

    Qualitative Mechanism Independence,

    O. E. Richardson, S. J. Peters, and J. Halpern, “Qualitative Mechanism Independence,” in The Thirty-eighth Annual Conference on Neural Information Processing Systems,

  5. [13]

    Durrett, Probability: theory and examples

    R. Durrett, Probability: theory and examples. Cambridge university press, 2019. 39

Pith tools

Reviewed August 12, 2026 · model on record in the stance chip above.