REVIEW 4 major objections 6 minor 1 cited by
Persistent change in an inferred time-series causal graph as conditioning depth grows signals that the observed state is incomplete — a warning of latent confounding or hidden memory, not a proof of either.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
Depth-dependent instability of inferred causal graphs is proposed as a non-specific warning sign of hidden memory in time-series causal discovery, demonstrated on simulations and four zebrafish calcium-imaging recordings.
T0 review reviewed 2026-08-03 challenge →
load-bearing objection Plausible, cheap model-checking idea with an honest body, but the abstract overclaims a calibration the body denies, and the real-data signature may just be finite-sample power loss. the 4 major comments →
Conditioning-Depth Diagnostics for Hidden Memory in Temporal Causal Discovery
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
Core claim
The central claim is the Heuristic principle of Section 3.2: for an observed process of finite order τ, once the conditioning depth p reaches τ, increasing p to p+1 should not systematically change the population-level graph recovered from the observed process. Persistent changes in the inferred graph as p increases therefore indicate that the chosen observed state representation is inadequate — a warning consistent with latent confounding or hidden memory, though not a proof that a particular latent variable exists. The paper implements this as a depth sweep of a constraint-based learner, with its c-GC and c-GC* variants as primary because their depth parameter is a matched fixed-horizon in
What carries the argument
The workhorse is the conditioning-depth sweep together with graph instability statistics defined on the inferred adjacency matrices. The Heuristic principle (Section 3.2) states that for p ≥ τ, enlarging the conditioning depth from p to p+1 should not induce systematic changes in the population graph; Remark 1 draws the diagnostic converse — persistent changes signal an inadequate observed state representation. The primary statistics are the normalized instability D_p = ||Â(p) − Â(p−1)||_0 / m at each adjacent step, its deletion and addition components D⁻_p and D⁺_p, the largest transition T_obs, and cumulative instability S_obs. c-GC and c-GC* carry the analysis because their depth paramete
Load-bearing premise
The interpretation assumes that the observed depth-sensitivity signature reflects genuine state inadequacy rather than the finite-sample behavior of the conditional-independence tests as the conditioning set grows; the paper lists finite-sample power loss as a possible contributor (Section 4.2) but provides no power-controlled baseline or surrogate-null calibration to rule it out.
What would settle it
Simulate an order-1 Markovian system at the sample sizes of the calcium-imaging recordings, run the same depth sweep (p=1 to 7) with c-GC and c-GC*, and check whether the deletion-dominated drop from p=1 to p=2 appears as often under this adequate-state null as in the non-Markovian simulations. Alternatively, run the paper's own bootstrap calibration (B=200 surrogates from a fitted order-1 null) on the real recordings: if T_obs falls inside the null distribution, the flagging signature is not evidence of hidden memory.
If this is right
- Stable graph trajectories across a depth sweep support the adequacy of the observed Markov state used by the learner; depth-sensitive trajectories warn that causal claims from shallow conditioning are not well supported.
- The diagnostic does not recover a latent graph or prove a latent confounder exists; instability alone cannot separate latent confounding from lag-order misspecification, nonstationarity, measurement error, or finite-sample test-power loss.
- In the calcium-imaging recordings, the sharp, deletion-dominated drop between p=1 and p=2 followed by leveling-off is compatible with unobserved neural drivers among the many unrecorded v2a-RSNs, and the saturation gives an operational cue for picking a conditioning depth.
- The bootstrap calibration protocol — surrogate-null bands, a global p-value for T_obs, an estimated onset depth, and edgewise FDR-controlled localization — is specified in the paper but deliberately not applied, so the reported real-data evidence is descriptive sensitivity analysis rather than a calibrated rejection of a Markovian null.
- Practically, the method implies that a depth sweep should accompany constraint-based time-series causal discovery: if conditioning on one more lag substantially changes the graph, the shallow graph should not be interpreted as causal.
Where Pith is reading between the lines
- Testable extension: apply the paper's own surrogate-null calibration (B bootstrapped series from a fitted order-1 null) to the calcium-imaging recordings. If the observed T_obs falls inside the null band, the p=1-to-p=2 drop is as consistent with power loss from larger conditioning sets as with hidden memory; the abstract reports a B=200 calibration that does not reject the order-1 null while the
- The diagnostic is portable: any learner with a non-adaptive, explicit depth parameter could host the same sweep. A power-controlled baseline — an adequately specified order-1 system subjected to increasing conditioning sets — would map the expected instability under the null and make the signature interpretable.
- Because instability is nonspecific, the sweep would gain discriminative power in combination with other signals, such as the direction of change (deletion-dominated transitions suggest spurious-edge pruning) or the consistency of the onset depth across multiple recordings.
- The clean mapping 'stable means adequate state, unstable means inadequate state' assumes the learner's conditional-independence tests are calibrated at every depth; if tests lose power as conditioning sets grow, even a correctly represented system could show a drop-and-level pattern, which is the premise most in need of a controlled check.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes a conditioning-depth diagnostic for temporal causal discovery: run a constraint-based learner across a grid of lag depths and measure instability of the inferred graph. Under an adequate finite-order Markov representation, the population graph should stabilize once p exceeds the true memory order (Heuristic principle, Section 3.2); persistent depth sensitivity is taken as evidence that the observed state representation is inadequate, possibly due to latent confounding or other hidden memory. The diagnostics are normalized instability D_p, its deletion/addition decomposition D^-_p/D^+_p, and summary statistics T_obs and S_obs (Section 3.6). Synthetic experiments compare c-GC, c-GC*, PCMCI+, and JPCMCI+ on Markovian and non-Markovian systems; real-data analyses on four calcium-imaging datasets show a sharp edge-count drop from p=1 to p=2 with deletions dominating, followed by relative stabilization. The paper explicitly disclaims that the method identifies latent confounders or separates alternative mechanisms, and states that the reported statistics are descriptive rather than calibrated.
Significance. If the diagnostic is valid, it addresses a real need: a low-cost model-checking warning for temporal causal discovery, especially for practitioners who must choose a conditioning depth for Granger-style or constraint-based methods. The manuscript is honest about its limitations, explicitly disclaims latent-graph recovery and mechanism identification, and provides reproducible code/notebooks. The population-level intuition is plausible, and the synthetic results show a separation between the Markovian and non-Markovian regimes for the c-GC variants. However, the load-bearing real-data interpretation is not statistically calibrated, and the abstract asserts calibration results that the body disavows. The contribution is modest but potentially useful; its current support is incomplete.
major comments (4)
- [Abstract vs. Sections 3.6 and 4.2] The arXiv abstract states: 'B=200 bootstrap calibration does not reject the fitted order-1 null' and reports repeat-level instability counts ('stable in every repeat' / 'positive instability in all c-GC repeats and eight of ten c-GC* repeats'). Section 3.6 explicitly says 'This calibration step is not an empirical component of the experiments reported below' and Section 4.2 says 'No surrogate null calibration, stationarity test, or finite sample power analysis is applied to these recordings.' The body also does not report per-repeat instability counts for the synthetic scenario. This is a direct factual inconsistency. If the abstract were accurate, the B=200 non-rejection would be exactly what a power-loss artifact predicts for an order-1 system, so the abstract's claim would not rescue the hidden-memory interpretation. The abstract must be corrected to match the body, or the missing res
- [Section 4.2, Table 2, Figs. 6-7] The central real-data evidence is the deletion-dominated drop from p=1 to p=2, interpreted as 'evidence of observed state inadequacy.' The paper lists 'finite-sample power loss from larger conditioning sets' as a possible contributor but provides no null distribution for D_p, T_obs, or D^-_2 under an adequate order-1 model at the real-data dimensionality (d=92-165). The Heuristic principle in Section 3.2 is population-level; applying it directly to finite-sample permutation tests requires a power-controlled baseline. The simulation baseline (Fig. 2) does not close this gap because the simulation dimensions and sample sizes are unspecified (see next comment). Without a surrogate-null or power analysis, the real-data signature cannot be distinguished from the finite-sample behavior of an actually adequate order-1 system. This is the load-bearing weakness of the empirical claim.
- [Section 4.1] The synthetic setup is too underspecified to support the transfer to real data. 'In each repetition, the number of time series and the length are sampled randomly' — no distributions, ranges, or typical values are given. The non-Markovian 'smooth unobserved driver term generated independently of the observed series with random amplitudes, centers, and widths' is not a formal DGP; it is unclear whether the driver causally affects at least two observed variables (i.e., whether it is a latent confounder or only a shared input). Without explicit equations, sample sizes, dimensions, and signal-to-noise settings, the separation shown in Figs. 2-5 cannot be reproduced, and no power comparison to the real-data regime (d=92-165, T≈3600) is possible. This is load-bearing because the real-data interpretation relies on the simulation showing that a latent driver produces depth sensitivity while a Ma
- [Sections 3.5 and 5] The primary real-data analysis uses the author's own c-GC and c-GC* learners, from an unreviewed preprint (Adedayo 2025). The simulation results show that these variants give the 'clearest separation,' while PCMCI+ and JPCMCI+ are weaker. Since the diagnostic's behavior depends on the conditional-independence test's power and calibration, it is unclear whether the claimed diagnostic property belongs to the conditioning-depth idea or to the specific testing procedure of c-GC/c-GC*. The manuscript should either provide independent validation of these learners or systematically analyze how test power and calibration affect the instability statistics, especially for the null case. Otherwise, the generality of the proposed diagnostic is not established.
minor comments (6)
- [Title] The full text title is 'Markovianity-Based Conditioning Depth Diagnostics for Hidden Confounding in Observational Datasets,' while the arXiv metadata gives 'Conditioning-Depth Diagnostics for Hidden Memory in Temporal Causal Discovery.' These should be reconciled.
- [Abstract] Typos: 'promnient' (should be 'prominent') and 'caustiosly' (should be 'cautiously') in the full-text abstract.
- [Section 3.6] The notation ||·||_0 for D_p is nonstandard for a normalized count; the text defines D_p as a fraction, so writing it as (1/m) * sum of indicators would be clearer.
- [Section 4.1, captions of Figs. 2-5] The figure captions write τ={0,1} for single-lag and τ={0,1,2} for multi-lag; the text elsewhere says τ=1 or τ∈{1,2}. The zero-lag notation is confusing; please align.
- [Section 5] Ungrammatical sentence: 'In the both simulations with cases, both variants separate stable Markovian data...' should be rewritten.
- [Section 3.7] The interpretation rule uses 'substantially' without an operational threshold. A predefined cutoff or a calibrated null would make the rule actionable; as written, it can only be applied post hoc.
Circularity Check
No significant circularity: the instability statistics are defined directly from adjacency outputs, the heuristic is an explicit assumption tested by simulation, and the self-cited learners are operationally specified and externally compared. The abstract/body bootstrap contradiction is a validity issue, not a circularity.
full rationale
The paper's derivation chain is not circular. The graph-instability statistics D_p, D^-_p, D^+_p, T_obs, and S_obs are defined directly from the sequence of inferred adjacency matrices (Sec. 3.6), not from any fitted parameter that is later called a prediction. The central 'Heuristic principle' (Sec. 3.2) is an explicitly stated assumption about population-level conditional tests, not a derived theorem, and the paper does not claim to derive it from first principles. Synthetic data are generated from explicit VAR and latent-driver DGPs (Sec. 4.1) rather than fitted to the instability diagnostics; the Markovian versus non-Markovian contrast is a simulation experiment, not an identity. Real-data results are explicitly described as descriptive model-checking summaries, and the paper repeatedly disclaims that the signature does not uniquely identify latent confounding (Secs. 3.6, 4.2, 5). The c-GC and c-GC* learners come from a self-citation (Adedayo 2025), but Section 3.5 gives a self-contained operational description of their two-stage tests, and the paper includes PCMCI+ and JPCMCI+ as independent external comparisons. Thus the self-citation is not load-bearing for the definition of the diagnostic or for the instability statistics. The main serious problem is the abstract's statement that 'B=200 bootstrap calibration does not reject the fitted order-1 null,' while Section 3.6 states 'This calibration step is not an empirical component of the experiments reported below' and Section 4.2 says 'No surrogate null calibration, stationarity test, or finite sample power analysis is applied to these recordings.' This is an internal inconsistency and a correctness/validity concern, not a circular reduction: no claimed prediction is equivalent by construction to an input. Likewise, the acknowledged finite-sample power-loss confound (Sec. 4.2) is a threat to the real-data interpretation, but it does not make the derivation circular. Under the stated hard rules, these concerns should not raise the circularity score.
Axiom & Free-Parameter Ledger
free parameters (4)
- c-GC / c-GC* significance thresholds α and β =
α=0.01, β=0.001
- Non-Markovian smooth driver parameters =
random amplitudes, centers, widths (unspecified)
- Conditioning-depth grid =
npasts = 1..7 (p = 2..7 for multi-lag)
- fish-3 ROI subset size =
130 of 420 ROIs
axioms (5)
- domain assumption Causal Markov assumption: each variable is independent of its non-effects given its direct causes.
- domain assumption Faithfulness: observed conditional independencies arise from graphical separation, not parameter cancellations.
- ad hoc to paper Heuristic principle: for p ≥ τ, enlarging conditioning depth from p to p+1 should not systematically change the population graph.
- domain assumption A stabilizing depth exists within the tested grid (hidden memory shorter than p_max).
- ad hoc to paper Hidden causes act by transmitting additional predictive information not summarized by the shallow observed state.
Cite this review
Pith. "Pith review of Conditioning-Depth Diagnostics for Hidden Memory in Temporal Causal Discovery." pith.science (2026). https://pith.science/paper/DAYRVGE3
@misc{pith2026260601214,
author = {Pith},
title = {Pith review of: Conditioning-Depth Diagnostics for Hidden Memory in Temporal Causal Discovery},
year = {2026},
howpublished = {\url{https://pith.science/paper/DAYRVGE3}},
note = {Machine review of arXiv:2606.01214}
}
read the original abstract
Reliable causal discovery in timeseries requires conditioning sets that capture the system state. When predictive history is omitted, residual dependence can appear as direct causal links. We test state adequacy by measuring how inferred graphs change as conditioning depth increases while the reported causal lag stays fixed. Under an adequate finite-order Markov representation, graphs should stabilize once enough observed history is conditioned on; latent common drive, omitted lags, nonstationarity, and measurement dynamics can instead produce depth sensitivity. We formalize this idea with graph instability statistics and evaluate c-GC and c-GC*, the two learners whose depth parameter implements a matched fixed-horizon history intervention. PCMCI+ and JPCMCI+ are excluded from the primary comparison because adaptive parent selection makes nominal depth edge-specific. In paired simulations, a clean order-1 process was stable in every repeat, whereas an AR(1) latent common driver produced positive instability in all c-GC repeats and eight of ten c-GC* repeats. In calcium imaging recordings, connectivity drops at the first transition beyond the one-lag baseline and then levels off, but B=200 bootstrap calibration does not reject the fitted order-1 null. The workflow therefore flags hidden memory or observed-state inadequacy without identifying the generating mechanism or recovering a latent graph.
Figures
Forward citations
Cited by 1 Pith paper
-
Blind Source Separation Can Distort Behavior and Connectivity Analyses of Calcium Transients
Component-removing blind source separation can collapse causal graph recovery to zero in synthetic calcium traces and dramatically densify connectivity estimates from real v2a-RSN traces, so BSS denoising is not a neu...
Reference graph
Works this paper leans on
-
[1]
Adedayo, S. A. (2025). Re-examining granger causality with causal bayesian networks and reichenbach’s principles.arXiv preprint https://arxiv.org/pdf/2501.02672v2. Ahrens, M. B., Orger, M. B., Robson, D. N., Li, J. M., and Keller, P. J. (2013). Whole-brain functional imaging at cellular resolution using light- sheet microscopy.Nature Methods, 10(5):413–42...
Pith/arXiv arXiv 2025
-
[30]
Maeda, T. N. and Shimizu, S. (2020). RCD: Repetitive causal discovery of linear non-gaussian acyclic models with latent confounders. InProceedings of the Twenty Third International Conference on Artificial Intelligence and Statistics, volume 108 of Proceedings of Machine Learning Research, pages 735–745. PMLR. Malinsky, D. and Spirtes, P. (2018). Causal s...
Pith/arXiv arXiv 2020
-
[313]
Geweke, J. F. (1984). Measures of conditional linear dependence and feedback between time series. Journal of the American Statistical Association, 79(388):907–915. Glymour, C., Zhang, K., and Spirtes, P. (2019). Review of causal discovery methods based on graphical models.Frontiers in Genetics, 10:524. 16 Granger, C. W. J. (1969). Investigating causal rel...
Pith/arXiv arXiv 1984
-
[1713]
and Runge, J
Gerhardus, A. and Runge, J. (2020). High-recall causal discovery for autocorrelated time series with latent confounders. InAdvances in Neural Information Processing Systems, volume 33, pages 12615–12625. Geweke, J. (1982). Measurement of linear dependence and feedback between multiple time series.Journal of the American Statistical Association, 77(378):304–
2020
-
[1790]
Chen, L., Li, C., Shen, X., and Pan, W. (2024a). Discovery and inference of a causal network with hidden confounding.Journal of the American Statistical Association, 119(548):2572–2584. Chen, W., Huang, Z., Cai, R., Hao, Z., and Zhang, K. (2024b). Identification of causal structure with latent variables based on higher order cumulants. Proceedings of the ...
Pith/arXiv arXiv 2014
-
[2023]
and Ramsey, J
Kummerfeld, E. and Ramsey, J. (2016). Causal clustering for 1-factor measurement models. Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, pages 1655–1664. Künsch, H. R. (1989). The jackknife and the bootstrap for general stationary observations.Annals of Statistics, 17(3):1217–1241. Kuroki, M. and Pearl, ...
2016
This paper was first reviewed by deepseek-v4-flash on August 3, 2026.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.