REVIEW 3 major objections 6 minor 44 references
This paper proves an exact identification threshold for cyclic latent states behind unknown sensors: d−1 mechanism targets suffice precisely when the one untouched node directly drives every other node; otherwise all d are required.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-01 12:28 UTC pith:RBXWBQ67
load-bearing objection Strong cyclic latent-variable identifiability results with honest scope, but the headline K=d−1 threshold rests on calibrated shift probes that may be as hard to guarantee as the identification itself. the 3 major comments →
Equilibrium Causal Games: Separation, Identification, and the Identifiability of Cyclic Latent States
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
Central claim: with aligned response maps M0^(e)=H(I−B^(e))^{-1}, a well-posed mechanism intervention on node j changes M0 by a rank-one matrix whose left factor is the known column M0e_j; the outer-product solve returns row j of the resolvent P=(I−B)^{-1}, and stacking rows gives P, then B=I−P^{-1} and H=M0P^{-1}, up to permutation and scale (the declared ≃, residual rotation collapsed by intervention). With one node r untouched, the final row is pinned by the zero-diagonal constraints through the completion matrix L=−diag(Ce_r)C^⊤, in closed form: det L=−(∏_{k≠r}B_kr)det C, corank L=#{k≠r:B_kr=0}. Thus K=d−1 mechanism targets identify the latent system, for every d, exactly when the sole u
What carries the argument
Two objects carry the argument. The first is the equilibrium resolvent P=(I−B)^{-1} of the linear latent model V=BV+U: with sensor map X=HV, the observed response is M0=H(I−B)^{-1}, and feedback contributions enter through this inverse rather than acyclic path sums. The second is the rank-one intervention identity of Lemma 1: zeroing row j of B by a do(j) edit makes the response difference D_j=M0^(j)−M0 rank one with known left factor M0e_j, and the least-squares solve w_j^⊤=−(M0e_j)^⊤D_j/‖M0e_j‖² returns (P−I)_{j,:}/P_jj — row j of P. The threshold rests on Proposition 5's completion matrix L=−diag(Ce_r)C^⊤ (C=I−B), the differential of the zero-diagonal constraints with row r missing; coran
Load-bearing premise
The load-bearing premise is that the experimenter can acquire aligned population response maps M0^(e)=H(I−B^(e))^{-1} in every environment, which requires identical full-rank labelled latent-coordinate shift probes, actuation that reaches the state only through the latent system, and sensor-noise conditional means that are equal (or known-corrected) across baseline and probe arms — the paper's own collision argument shows that without these, the response maps are unrecoverabl
What would settle it
Compute L=−diag(Ce_r)C^⊤ on a d=3 graph where the untargeted node has a missing child; Proposition 5 predicts corank L = #missing children and a positive-dimensional family of alternative factorizations reproducing every environment response. A symbolic or high-precision evaluation where corank L differs, or a numeric collision search that finds no such family, would refute the paper's central threshold.
If this is right
- Experiment design for cyclic latent recovery reduces to a closed-form graph check: leave one node untouched, count the children it fails to parent, and you know whether K=d−1 mechanism targets will identify (H,B) up to the declared equivalence or whether all d targets are needed.
- Shift-only interventions (adding constants to equations while keeping B fixed) are provably insufficient for d≥2 in the ambient unknown-support, unknown-H class: they recover at most M0=H(I−B)^{-1}, never the interaction matrix B separately from the sensor map H.
- Passive equilibrium data from a stable cyclic system with unknown sensing are maximally uninformative about the wiring: for any legal zero-diagonal B′ there exists a legal sensor map reproducing the identical law, so targeted interventions, not merely more observations, are the necessary cure.
- Non-Gaussian (LiNG) sources collapse the passive source-frame rotation, but they neither align environments nor separate H from B; the two ambiguities are transverse and require different information, so non-Gaussianity cannot substitute for aligned mechanism responses.
- In the nonlinear regime, isotropic Gaussian source blocks admit hidden twists within and across blocks that preserve every labelled law, leaving within-block factors and even the block partition unidentified; under score-rank and irreducibility conditions, only the finest independent source-block representation is identified, up to block permutation and blockwise diffeomorphisms.
Where Pith is reading between the lines
- The rank-one structure of Lemma 1 suggests a cheap data-quality diagnostic for real deployments: before trusting a recovery, check that the estimated response difference at each targeted node is numerically rank one; systematic failure at a declared target indicates acquisition violations (shifts not aligned, sensor drift, or actuator side-effects) rather than estimator noise.
- The K=d−1 iff reads like an observability condition: the untargeted node's missing out-edges are exactly the unreachable directions of the identification map. The same completion-matrix technique may transfer to other latent-variable problems where a structural constraint class (here, zero diagonal) supplies the information an extra intervention would otherwise provide.
- For practitioners in energy markets or epidemiology, the paper implies a division of labour: run the gauge analysis first to certify which conclusions passive equilibrium data cannot support, then choose mechanism targets by the corank criterion rather than by graph heuristics — the cost of one wrong omitted target can be the jump from K=d−1 to K=d.
- The nonlinear twist collisions suggest that nonlinear block-identification claims in practice should be treated with caution unless the sensor or the interventions break isotropy; whether realistic departures from isotropic sources, or block-separable sensors, close the gap is an open route the paper explicitly leaves.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces the Equilibrium Causal Game (ECG), a cyclic-SCM-with-game-layer framework, and studies separation and identification of latent equilibrium states from observed data. It contains three families of results: (1) separation theory for cyclic equilibrium graphs, including sound-but-incomplete σ-separation, a trek-separation theorem for linear ECGs, and a functional-separation sufficient condition; (2) an observed-layer identification analysis with a potential-game specification test; and (3) the main latent-layer results. The central technical claim is Theorem 16 with Proposition 5: given aligned population response matrices M0^(e)=H(I−B^(e))^{-1} and well-posed single-target mechanism interventions, full target coverage identifies (H,B) up to ≃; with one untargeted node r, K=d−1 targets suffice exactly when r directly parents every other node, and otherwise K=d targets are necessary. Nonlinear results provide isotropy-preserving collisions (Theorem 17) and a scoped source-block identification theorem (Theorem 18). The paper is unusually explicit about assumptions and limitations, and it marks experiments as finite-instance evidence rather than proof.
Significance. If the central claims hold, the paper gives an exact, graph-global intervention-count threshold for cyclic latent linear models, extending acyclic interventional causal representation learning to equilibrium systems. The closed-form algebra in Lemma 1 and Proposition 5 (Sherman-Morrison rank-one differences and the cofactor completion matrix) is a substantive and checkable contribution. The paper also provides useful negative results (Theorem 15, Theorem 17) that prevent overclaiming, and it is honest about the scope of its positive results. I agree with the reader's assessment that the identification algebra is internally coherent and not circular: the equivalence ≃ is not assumed into the conclusion, and the residual gauges are independently shown to be genuine ambiguities. The main risk is upstream of the algebra: the theorem's input, the aligned response matrices, is available only through a probe design that itself requires calibrated access to latent coordinates, and the equivalence relation ≃ is not fully specified. These issues affect the precise statement and practical scope of the central identification result.
major comments (3)
- [Theorem 16, acquisition paragraph] The theorem's key input M0^(e)=H(I−B^(e))^{-1} is obtained, under the only theorem-backed acquisition route, by applying 'known latent-coordinate shift probes' with design D. If the latent coordinate frame is not known—precisely the quantity the theorem aims to recover—then a probe declared as δ is Aδ in the true coordinates, and the formula M0=R D^T(DD^T)^{-1} returns M0 A instead of M0. Lemma 1 then fails: its rank-one recovery relies on the known left factor M0 e_j, which under A becomes a mixture of columns of M0. The text discusses the additive collision Re=M0D+N, but the coordinate-frame issue is a multiplicative gauge, not a sensor-mean effect. This is load-bearing for the headline K=d−1 iff. Please either provide an acquisition argument that does not presuppose the latent coordinate frame, or state explicitly in the main theorems and abstract that the result assumes calibrated ac
- [Definition 5 and Section 7.2] The equivalence relation ≃ is not fully specified. The definition states H'=H Pπ Λ and η'=Λ^{-1} Pπ^T η, but it does not give the corresponding transformations of B' and Ω', nor the precise action of the residual SO(|S|) gauge on B and Ω. Since the central conclusions are 'identified up to ≃', and proofs of non-≃-equivalence (e.g., Proposition 5 proof and Theorem 16(b)) argue only from the form of H', the relation is not a well-defined equivalence relation on the full parameter space. Please complete the definition, state how B and Ω transform under permutation, scaling, and the residual rotation, and verify that ≃ is indeed an equivalence relation. This is necessary for the precise formulation of all latent-layer identification claims.
- [Theorem 16(d) and Exp. 9] The sample-complexity claim 'plug-in recovery is √n-consistent' is stated for a 'regular instance' satisfying unknown population conditions such as σmin(L)>0, Pjj≠0, and fixed probe-design rank. The theorem does not supply a data-dependent way to verify that the instance is in the identified regime before committing to the estimator. This is a limitation rather than a mathematical error, but the practical message of a √n rate should be more carefully scoped: the rate is conditional on non-degeneracy conditions that are part of the inference target. The experiment description already notes that Exp. 9 only instantiates the calibrated design; please make this caveat prominent in the theorem statement as well.
minor comments (6)
- [Abstract and general text] There are several typos and punctuation issues, e.g., 'Power grids, markets, and interacting populations, settle into' and inconsistent capitalization such as 'proposition 7' vs. 'Theorem 16'. A careful copyedit is needed.
- [Section 7.2, proof of Theorem 16(b)] The construction E_t = t(e_s − C_rs e_r)e_r^T is terse. A one-sentence explanation that I+E_t has identity rows on all target rows j ∉ {r,s} would help the reader verify the reproduction of every targeted environment.
- [Section 6, Definition 4] The phrase 'at most one Gaussian (the LiNG condition)' should clarify whether 'one' means one per source component or one globally. This affects the interpretation of the hedge-collapse statement.
- [Section A, Lemma 3 and Theorem 6] The no-cancellation Lemma 3 is central to the trek-separation proof but is stated only in the proof appendix. Moving its statement (and a short explanation) to the main text would make Theorem 6 much easier to verify.
- [Section 9, experimental reproducibility] The code and data are said to be 'available from the author upon reasonable request'. For a paper with this many claims, a permanent repository or artifact link is strongly preferable.
- [Section 10, Limitations] The limitations section is thorough and honest. However, the specific issue of the acquisition probes requiring known latent coordinates is not listed among the limitations; it is only implicit in the 'active route uses calibrated full-rank shift probes' sentence. Please add an explicit bullet.
Circularity Check
No significant circularity: identification theorems are algebraic derivations from declared aligned response maps; restrictive acquisition assumptions are scoping, not self-referential.
full rationale
The central claims are derived by explicit algebra from declared population-level inputs. Theorem 16's input is the aligned response map M0(e)=H(I-B(e))^-1, obtained via the exact identity M0(e)=Re D^T(DD^T)^-1 under the stated no-collision and known-probe assumptions; this is an observational reduction using a declared intervention design, not a fitted parameter. Lemma 1 then uses Sherman-Morrison to write the intervention difference D_j = -(M0 e_j)(b_j^T P)/P_jj and solves for row j of P from the known left factor M0 e_j; Proposition 5 completes the missing row through the closed-form L = -diag(Ce_r)C^T. None of these steps assumes B, H, or the target row. The remaining ambiguity in the declared equivalence ≃ is independently shown to be genuine (Theorems 13, 15, 17 and Proposition 8), so the 'up to ≃' claim is not an equivalence-by-definition. There are no self-citations and no fitted values recycled as predictions. The upstream acquisition assumptions - known latent-coordinate shift probes, mean-stable sensing, no direct actuator-to-sensor effect - are restrictive and acknowledged in the acquisition paragraph via the collision Re = M0(e)D + N, and the Limitations section (item 4) explicitly states that the active route requires calibrated full-rank shift probes; item 7 disclaims external validation. These are scoping conditions rather than circular reductions: M0(e) is not the target (H,B), and the theorem does not claim recovery from unaligned, uncalibrated, or probe-biased data. The algebra of identification is therefore self-contained relative to its declared assumptions.
Axiom & Free-Parameter Ledger
axioms (7)
- domain assumption Represented exogeneity: every shared exogenous source is represented as an explicit latent common parent in the augmented graph.
- domain assumption Loop-solvability premise: the ECG is uniquely solvable with respect to every strongly connected subset.
- domain assumption Stable selection class: contraction ρ(B)<1 on each cyclic component defines the selected equilibria.
- domain assumption LiNG source condition: independent nondegenerate source components with at most one Gaussian, plus Comon's ICA identifiability to monomial ambiguity.
- domain assumption Acquisition assumptions: known labelled latent-coordinate shift probes of rank D=d, actuator affects X only through V, and probe-arm-invariant sensor conditional mean.
- domain assumption Zero-diagonal B and full-column-rank environment-invariant H in the ambient class.
- standard math External mathematical theorems: Comon's ICA, Craig–Sakamoto, Jacobi/Coates/all-minors expansion, Cauchy–Binet, Menger/linkage, and the half-trek criterion.
read the original abstract
Power grids, markets, and interacting populations, settle into feedback driven equilibria observed through unknown sensors. Our Equilibrium Causal Game (ECG) joins a game to its cyclic causal model, hidden inputs, sensor map, and rules for interventions and equilibrium selection; interventions edit declared objects and recompute equilibrium. Under stated conditions, ECG-separation is sound but incomplete in our examples. Back-door/half-trek routes identify observed queries. Yet for an untouched rotationally symmetric Gaussian block, second moments determine only a source-frame rotation, across which distinct-variable effects generically change. Unknown sensing creates a separate ambiguity. In passive stable linear models without self-effects, unknown wiring and full-rank unknown sensing leave $B$ completely unidentified for $d\ge2$. Under LiNG, non-Gaussianity removes the source rotation; mechanism interventions separate sensing from interactions. With unknown support, invariant sensing, aligned responses, and well-posed single-target interventions identify $(H,B)$ up to declared equivalence. Of $d$ targets, $d-1$ suffice exactly when the sole untargeted node directly parents all others; otherwise $d$ are needed. Acquisition probes are excluded; known wiring gives no universal count. With nonlinear sensing, isotropic Gaussian source blocks admit hidden twists within and across blocks in labelled environments preserving required radial laws. Conversely, under stated positivity, informative one-block changes, rank, and irreducibility conditions, the finest independent source-block representation is identified within the stated alternative class up to block permutation and blockwise coordinate changes, but not downstream mechanisms or the sensor/interaction split. Together, these results show which causal conclusions equilibrium data support and which require targeted experiments.
Figures
Reference graph
Works this paper leans on
-
[1]
A Simple Proof of the
Li, Chi-Kwong , journal=. A Simple Proof of the. 2000 , doi=
2000
-
[2]
Proceedings of the 40th International Conference on Machine Learning (ICML) , series =
Ahuja, Kartik and Mahajan, Divyat and Wang, Yixin and Bengio, Yoshua , title =. Proceedings of the 40th International Conference on Machine Learning (ICML) , series =. 2023 , eprint =
2023
-
[3]
, title =
Blom, Tineke and Bongers, Stephan and Mooij, Joris M. , title =. Proceedings of the 35th Conference on Uncertainty in Artificial Intelligence (UAI) , year =
-
[4]
Theoretical Aspects of Cyclic Structural Causal Models , journal =
Bongers, Stephan and Peters, Jonas and Sch. Theoretical Aspects of Cyclic Structural Causal Models , journal =
-
[5]
Foundations of Structural Causal Models with Cycles and Latent Variables , journal =
Bongers, Stephan and Forr. Foundations of Structural Causal Models with Cycles and Latent Variables , journal =. 2021 , doi =
2021
-
[6]
Journal of Industrial Economics , volume =
Borenstein, Severin and Bushnell, James , title =. Journal of Industrial Economics , volume =
-
[7]
Strategic Interaction and Networks , journal =
Bramoull. Strategic Interaction and Networks , journal =
-
[8]
Advances in Neural Information Processing Systems , volume =
Brehmer, Johann and De Haan, Pim and Lippe, Phillip and Cohen, Taco , title =. Advances in Neural Information Processing Systems , volume =
-
[9]
Learning Linear Causal Representations from Interventions under General Nonlinear Mixing , booktitle =
Buchholz, Simon and Rajendran, Goutham and Rosenfeld, Elan and Aragam, Bryon and Sch. Learning Linear Causal Representations from Interventions under General Nonlinear Mixing , booktitle =
-
[10]
Neural Computation , volume =
Chalak, Karim and White, Halbert , title =. Neural Computation , volume =. 2012 , doi =
2012
-
[11]
1976 , edition =
Chen, Wai-Kai , title =. 1976 , edition =
1976
-
[12]
, title =
Coates, Clarence L. , title =. IRE Transactions on Circuit Theory , volume =
-
[13]
Signal Processing , volume =
Comon, Pierre , title =. Signal Processing , volume =. 1994 , doi =
1994
-
[14]
arXiv preprint arXiv:2603.04780 , year =
Dai, Haoyue and Albrecht, Immanuel and Spirtes, Peter and Zhang, Kun , title =. arXiv preprint arXiv:2603.04780 , year =
-
[15]
Ferradini, Carla and Gitton, Victor and Vilasini, V. , title =. arXiv preprint arXiv:2502.04171 , year =
-
[16]
Constraint-Based Causal Discovery for Non-Linear Structural Causal Models with Cycles and Latent Confounders , booktitle =
Forr. Constraint-Based Causal Discovery for Non-Linear Structural Causal Models with Cycles and Latent Confounders , booktitle =
-
[17]
Causal Calculus in the Presence of Cycles, Latent Confounders and Selection Bias , booktitle =
Forr. Causal Calculus in the Presence of Cycles, Latent Confounders and Selection Bias , booktitle =. 2019 , eprint =
2019
-
[18]
The Annals of Statistics , volume =
Foygel, Rina and Draisma, Jan and Drton, Mathias , title =. The Annals of Statistics , volume =
-
[19]
Artificial Intelligence , volume =
Hammond, Lewis and Fox, James and Everitt, Tom and Carey, Ryan and Abate, Alessandro and Wooldridge, Michael , title =. Artificial Intelligence , volume =. 2023 , doi =
2023
-
[20]
and Johnson, Charles R
Horn, Roger A. and Johnson, Charles R. , title =
-
[21]
, title =
Hyttinen, Antti and Eberhardt, Frederick and Hoyer, Patrik O. , title =. Journal of Machine Learning Research , volume =
-
[22]
Nonlinear
Hyv. Nonlinear. Proceedings of the 22nd International Conference on Artificial Intelligence and Statistics (AISTATS) , pages =
-
[23]
and Monti, Ricardo Pio and Hyv
Khemakhem, Ilyes and Kingma, Diederik P. and Monti, Ricardo Pio and Hyv. Variational Autoencoders and Nonlinear. Proceedings of the 23rd International Conference on Artificial Intelligence and Statistics (AISTATS) , pages =
-
[24]
, title =
Lacerda, Gustavo and Spirtes, Peter and Ramsey, Joseph and Hoyer, Patrik O. , title =. Proceedings of the 24th Conference on Uncertainty in Artificial Intelligence (UAI) , pages =. 2008 , eprint =
2008
-
[25]
Additive Decoders for Latent Variables Identification and Cartesian-Product Extrapolation , booktitle =
Lachapelle, S. Additive Decoders for Latent Variables Identification and Cartesian-Product Extrapolation , booktitle =
-
[26]
and Richardson, Thomas S
Lauritzen, Steffen L. and Richardson, Thomas S. , title =. Journal of the Royal Statistical Society: Series B (Statistical Methodology) , volume =. 2002 , doi =
2002
-
[27]
arXiv preprint arXiv:2601.12707 , year =
Liao, Junyi and Zhu, Zihan and Fang, Ethan and Yang, Zhuoran and Tarokh, Vahid , title =. arXiv preprint arXiv:2601.12707 , year =
-
[28]
Proceedings of the 39th International Conference on Machine Learning (ICML) , year =
Lippe, Phillip and Magliacane, Sara and L. Proceedings of the 39th International Conference on Machine Learning (ICML) , year =. 2202.03169 , archivePrefix =
-
[29]
Causal Representation Learning for Instantaneous and Temporal Effects in Interactive Systems , booktitle =
Lippe, Phillip and Magliacane, Sara and L. Causal Representation Learning for Instantaneous and Temporal Effects in Interactive Systems , booktitle =
-
[30]
, title =
Monderer, Dov and Shapley, Lloyd S. , title =. Games and Economic Behavior , volume =
-
[31]
Causal Representation Learning Made Identifiable by Grouping of Observational Variables , journal =
Morioka, Hiroshi and Hyv. Causal Representation Learning Made Identifiable by Grouping of Observational Variables , journal =
-
[32]
Games and Economic Behavior , volume =
Parise, Francesca and Ozdaglar, Asuman , title =. Games and Economic Behavior , volume =
-
[33]
, title =
Peters, Spencer and Halpern, Joseph Y. , title =. Journal of Artificial Intelligence Research , volume =. 2025 , doi =
2025
-
[34]
Advances in Neural Information Processing Systems , volume =
Rothenh. Advances in Neural Information Processing Systems , volume =. 2015 , doi =
2015
-
[35]
Shimizu, Shohei and Hoyer, Patrik O. and Hyv. A Linear Non-. Journal of Machine Learning Research , volume =
-
[36]
Shpitser, Ilya and Pearl, Judea , title =. Proc. 21st AAAI Conf. on Artificial Intelligence (AAAI) , pages =
-
[37]
Proceedings of the 29th Conference on Uncertainty in Artificial Intelligence (UAI) , year =
Spirtes, Peter , title =. Proceedings of the 29th Conference on Uncertainty in Artificial Intelligence (UAI) , year =. 1309.7004 , archivePrefix =
-
[38]
Proceedings of the 40th International Conference on Machine Learning (ICML) , series =
Squires, Chandler and Seigal, Anna and Bhate, Salil and Uhler, Caroline , title =. Proceedings of the 40th International Conference on Machine Learning (ICML) , series =. 2023 , eprint =
2023
-
[39]
The Annals of Statistics , volume =
Sullivant, Seth and Talaska, Kelli and Draisma, Jan , title =. The Annals of Statistics , volume =
-
[40]
Journal of Machine Learning Research , volume =
White, Halbert and Chalak, Karim , title =. Journal of Machine Learning Research , volume =
-
[41]
Proceedings of the 36th International Conference on Machine Learning (ICML) , pages =
Yu, Lantao and Song, Jiaming and Ermon, Stefano , title =. Proceedings of the 36th International Conference on Machine Learning (ICML) , pages =
-
[42]
Advances in Neural Information Processing Systems , volume =
Zhang, Junzhe and Kumor, Daniel and Bareinboim, Elias , title =. Advances in Neural Information Processing Systems , volume =
-
[43]
Advances in Neural Information Processing Systems (NeurIPS) , volume =
Zhang, Jiaqi and Squires, Chandler and Greenewald, Kristjan and Srivastava, Akash and Shanmugam, Karthikeyan and Uhler, Caroline , title =. Advances in Neural Information Processing Systems (NeurIPS) , volume =. 2023 , eprint =
2023
-
[44]
Nonparametric Identifiability of Causal Representations from Unknown Interventions , booktitle =
von K. Nonparametric Identifiability of Causal Representations from Unknown Interventions , booktitle =
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.