Pith. sign in

REVIEW 4 major objections 6 minor 1 cited by

Persistent change in an inferred time-series causal graph as conditioning depth grows signals that the observed state is incomplete — a warning of latent confounding or hidden memory, not a proof of either.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

Depth-dependent instability of inferred causal graphs is proposed as a non-specific warning sign of hidden memory in time-series causal discovery, demonstrated on simulations and four zebrafish calcium-imaging recordings.

T0 review reviewed 2026-08-03 challenge →

load-bearing objection Plausible, cheap model-checking idea with an honest body, but the abstract overclaims a calibration the body denies, and the real-data signature may just be finite-sample power loss. the 4 major comments →

arxiv 2606.01214 v2 pith:DAYRVGE3 submitted 2026-05-31 stat.AP

Conditioning-Depth Diagnostics for Hidden Memory in Temporal Causal Discovery

classification stat.AP
keywords latent confoundersMarkovianityconstraint-based causal discoverytime seriesgraph instabilityconditioning depthhidden memorycalcium imaging
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper turns a known failure mode of time-series causal discovery into a practical diagnostic: instead of reading one learned graph, run the learner across a grid of conditioning depths and watch whether the graph stabilizes. When the observed process is adequately represented by a finite-order Markov state, the graph should settle once enough past history is conditioned on; persistent depth sensitivity indicates the observed state is missing relevant history — a signature consistent with latent common drivers, omitted lags, nonstationarity, or measurement dynamics. The paper formalizes this with graph instability statistics, shows in paired simulations that a clean order-1 process stays stable in every repeat while an AR(1) latent common driver produces positive instability in all c-GC repeats and eight of ten c-GC* repeats, and reports a deletion-dominated drop-and-level pattern in calcium-imaging recordings. The contribution is deliberately modest: a model-checking warning that shallow causal claims deserve caution, not a method that identifies or recovers latent graphs.

Core claim

The central claim is the Heuristic principle of Section 3.2: for an observed process of finite order τ, once the conditioning depth p reaches τ, increasing p to p+1 should not systematically change the population-level graph recovered from the observed process. Persistent changes in the inferred graph as p increases therefore indicate that the chosen observed state representation is inadequate — a warning consistent with latent confounding or hidden memory, though not a proof that a particular latent variable exists. The paper implements this as a depth sweep of a constraint-based learner, with its c-GC and c-GC* variants as primary because their depth parameter is a matched fixed-horizon in

What carries the argument

The workhorse is the conditioning-depth sweep together with graph instability statistics defined on the inferred adjacency matrices. The Heuristic principle (Section 3.2) states that for p ≥ τ, enlarging the conditioning depth from p to p+1 should not induce systematic changes in the population graph; Remark 1 draws the diagnostic converse — persistent changes signal an inadequate observed state representation. The primary statistics are the normalized instability D_p = ||Â(p) − Â(p−1)||_0 / m at each adjacent step, its deletion and addition components D⁻_p and D⁺_p, the largest transition T_obs, and cumulative instability S_obs. c-GC and c-GC* carry the analysis because their depth paramete

Load-bearing premise

The interpretation assumes that the observed depth-sensitivity signature reflects genuine state inadequacy rather than the finite-sample behavior of the conditional-independence tests as the conditioning set grows; the paper lists finite-sample power loss as a possible contributor (Section 4.2) but provides no power-controlled baseline or surrogate-null calibration to rule it out.

What would settle it

Simulate an order-1 Markovian system at the sample sizes of the calcium-imaging recordings, run the same depth sweep (p=1 to 7) with c-GC and c-GC*, and check whether the deletion-dominated drop from p=1 to p=2 appears as often under this adequate-state null as in the non-Markovian simulations. Alternatively, run the paper's own bootstrap calibration (B=200 surrogates from a fitted order-1 null) on the real recordings: if T_obs falls inside the null distribution, the flagging signature is not evidence of hidden memory.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

If this is right

  • Stable graph trajectories across a depth sweep support the adequacy of the observed Markov state used by the learner; depth-sensitive trajectories warn that causal claims from shallow conditioning are not well supported.
  • The diagnostic does not recover a latent graph or prove a latent confounder exists; instability alone cannot separate latent confounding from lag-order misspecification, nonstationarity, measurement error, or finite-sample test-power loss.
  • In the calcium-imaging recordings, the sharp, deletion-dominated drop between p=1 and p=2 followed by leveling-off is compatible with unobserved neural drivers among the many unrecorded v2a-RSNs, and the saturation gives an operational cue for picking a conditioning depth.
  • The bootstrap calibration protocol — surrogate-null bands, a global p-value for T_obs, an estimated onset depth, and edgewise FDR-controlled localization — is specified in the paper but deliberately not applied, so the reported real-data evidence is descriptive sensitivity analysis rather than a calibrated rejection of a Markovian null.
  • Practically, the method implies that a depth sweep should accompany constraint-based time-series causal discovery: if conditioning on one more lag substantially changes the graph, the shallow graph should not be interpreted as causal.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • Testable extension: apply the paper's own surrogate-null calibration (B bootstrapped series from a fitted order-1 null) to the calcium-imaging recordings. If the observed T_obs falls inside the null band, the p=1-to-p=2 drop is as consistent with power loss from larger conditioning sets as with hidden memory; the abstract reports a B=200 calibration that does not reject the order-1 null while the
  • The diagnostic is portable: any learner with a non-adaptive, explicit depth parameter could host the same sweep. A power-controlled baseline — an adequately specified order-1 system subjected to increasing conditioning sets — would map the expected instability under the null and make the signature interpretable.
  • Because instability is nonspecific, the sweep would gain discriminative power in combination with other signals, such as the direction of change (deletion-dominated transitions suggest spurious-edge pruning) or the consistency of the onset depth across multiple recordings.
  • The clean mapping 'stable means adequate state, unstable means inadequate state' assumes the learner's conditional-independence tests are calibrated at every depth; if tests lose power as conditioning sets grow, even a correctly represented system could show a drop-and-level pattern, which is the premise most in need of a controlled check.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. The paper proposes a conditioning-depth diagnostic for temporal causal discovery: run a constraint-based learner across a grid of lag depths and measure instability of the inferred graph. Under an adequate finite-order Markov representation, the population graph should stabilize once p exceeds the true memory order (Heuristic principle, Section 3.2); persistent depth sensitivity is taken as evidence that the observed state representation is inadequate, possibly due to latent confounding or other hidden memory. The diagnostics are normalized instability D_p, its deletion/addition decomposition D^-_p/D^+_p, and summary statistics T_obs and S_obs (Section 3.6). Synthetic experiments compare c-GC, c-GC*, PCMCI+, and JPCMCI+ on Markovian and non-Markovian systems; real-data analyses on four calcium-imaging datasets show a sharp edge-count drop from p=1 to p=2 with deletions dominating, followed by relative stabilization. The paper explicitly disclaims that the method identifies latent confounders or separates alternative mechanisms, and states that the reported statistics are descriptive rather than calibrated.

Significance. If the diagnostic is valid, it addresses a real need: a low-cost model-checking warning for temporal causal discovery, especially for practitioners who must choose a conditioning depth for Granger-style or constraint-based methods. The manuscript is honest about its limitations, explicitly disclaims latent-graph recovery and mechanism identification, and provides reproducible code/notebooks. The population-level intuition is plausible, and the synthetic results show a separation between the Markovian and non-Markovian regimes for the c-GC variants. However, the load-bearing real-data interpretation is not statistically calibrated, and the abstract asserts calibration results that the body disavows. The contribution is modest but potentially useful; its current support is incomplete.

major comments (4)
  1. [Abstract vs. Sections 3.6 and 4.2] The arXiv abstract states: 'B=200 bootstrap calibration does not reject the fitted order-1 null' and reports repeat-level instability counts ('stable in every repeat' / 'positive instability in all c-GC repeats and eight of ten c-GC* repeats'). Section 3.6 explicitly says 'This calibration step is not an empirical component of the experiments reported below' and Section 4.2 says 'No surrogate null calibration, stationarity test, or finite sample power analysis is applied to these recordings.' The body also does not report per-repeat instability counts for the synthetic scenario. This is a direct factual inconsistency. If the abstract were accurate, the B=200 non-rejection would be exactly what a power-loss artifact predicts for an order-1 system, so the abstract's claim would not rescue the hidden-memory interpretation. The abstract must be corrected to match the body, or the missing res
  2. [Section 4.2, Table 2, Figs. 6-7] The central real-data evidence is the deletion-dominated drop from p=1 to p=2, interpreted as 'evidence of observed state inadequacy.' The paper lists 'finite-sample power loss from larger conditioning sets' as a possible contributor but provides no null distribution for D_p, T_obs, or D^-_2 under an adequate order-1 model at the real-data dimensionality (d=92-165). The Heuristic principle in Section 3.2 is population-level; applying it directly to finite-sample permutation tests requires a power-controlled baseline. The simulation baseline (Fig. 2) does not close this gap because the simulation dimensions and sample sizes are unspecified (see next comment). Without a surrogate-null or power analysis, the real-data signature cannot be distinguished from the finite-sample behavior of an actually adequate order-1 system. This is the load-bearing weakness of the empirical claim.
  3. [Section 4.1] The synthetic setup is too underspecified to support the transfer to real data. 'In each repetition, the number of time series and the length are sampled randomly' — no distributions, ranges, or typical values are given. The non-Markovian 'smooth unobserved driver term generated independently of the observed series with random amplitudes, centers, and widths' is not a formal DGP; it is unclear whether the driver causally affects at least two observed variables (i.e., whether it is a latent confounder or only a shared input). Without explicit equations, sample sizes, dimensions, and signal-to-noise settings, the separation shown in Figs. 2-5 cannot be reproduced, and no power comparison to the real-data regime (d=92-165, T≈3600) is possible. This is load-bearing because the real-data interpretation relies on the simulation showing that a latent driver produces depth sensitivity while a Ma
  4. [Sections 3.5 and 5] The primary real-data analysis uses the author's own c-GC and c-GC* learners, from an unreviewed preprint (Adedayo 2025). The simulation results show that these variants give the 'clearest separation,' while PCMCI+ and JPCMCI+ are weaker. Since the diagnostic's behavior depends on the conditional-independence test's power and calibration, it is unclear whether the claimed diagnostic property belongs to the conditioning-depth idea or to the specific testing procedure of c-GC/c-GC*. The manuscript should either provide independent validation of these learners or systematically analyze how test power and calibration affect the instability statistics, especially for the null case. Otherwise, the generality of the proposed diagnostic is not established.
minor comments (6)
  1. [Title] The full text title is 'Markovianity-Based Conditioning Depth Diagnostics for Hidden Confounding in Observational Datasets,' while the arXiv metadata gives 'Conditioning-Depth Diagnostics for Hidden Memory in Temporal Causal Discovery.' These should be reconciled.
  2. [Abstract] Typos: 'promnient' (should be 'prominent') and 'caustiosly' (should be 'cautiously') in the full-text abstract.
  3. [Section 3.6] The notation ||·||_0 for D_p is nonstandard for a normalized count; the text defines D_p as a fraction, so writing it as (1/m) * sum of indicators would be clearer.
  4. [Section 4.1, captions of Figs. 2-5] The figure captions write τ={0,1} for single-lag and τ={0,1,2} for multi-lag; the text elsewhere says τ=1 or τ∈{1,2}. The zero-lag notation is confusing; please align.
  5. [Section 5] Ungrammatical sentence: 'In the both simulations with cases, both variants separate stable Markovian data...' should be rewritten.
  6. [Section 3.7] The interpretation rule uses 'substantially' without an operational threshold. A predefined cutoff or a calibrated null would make the rule actionable; as written, it can only be applied post hoc.

Circularity Check

0 steps flagged

No significant circularity: the instability statistics are defined directly from adjacency outputs, the heuristic is an explicit assumption tested by simulation, and the self-cited learners are operationally specified and externally compared. The abstract/body bootstrap contradiction is a validity issue, not a circularity.

full rationale

The paper's derivation chain is not circular. The graph-instability statistics D_p, D^-_p, D^+_p, T_obs, and S_obs are defined directly from the sequence of inferred adjacency matrices (Sec. 3.6), not from any fitted parameter that is later called a prediction. The central 'Heuristic principle' (Sec. 3.2) is an explicitly stated assumption about population-level conditional tests, not a derived theorem, and the paper does not claim to derive it from first principles. Synthetic data are generated from explicit VAR and latent-driver DGPs (Sec. 4.1) rather than fitted to the instability diagnostics; the Markovian versus non-Markovian contrast is a simulation experiment, not an identity. Real-data results are explicitly described as descriptive model-checking summaries, and the paper repeatedly disclaims that the signature does not uniquely identify latent confounding (Secs. 3.6, 4.2, 5). The c-GC and c-GC* learners come from a self-citation (Adedayo 2025), but Section 3.5 gives a self-contained operational description of their two-stage tests, and the paper includes PCMCI+ and JPCMCI+ as independent external comparisons. Thus the self-citation is not load-bearing for the definition of the diagnostic or for the instability statistics. The main serious problem is the abstract's statement that 'B=200 bootstrap calibration does not reject the fitted order-1 null,' while Section 3.6 states 'This calibration step is not an empirical component of the experiments reported below' and Section 4.2 says 'No surrogate null calibration, stationarity test, or finite sample power analysis is applied to these recordings.' This is an internal inconsistency and a correctness/validity concern, not a circular reduction: no claimed prediction is equivalent by construction to an input. Likewise, the acknowledged finite-sample power-loss confound (Sec. 4.2) is a threat to the real-data interpretation, but it does not make the derivation circular. Under the stated hard rules, these concerns should not raise the circularity score.

Axiom & Free-Parameter Ledger

4 free parameters · 5 axioms · 0 invented entities

The paper's interpretation rests on explicitly stated assumptions (causal Markov, p ≥ τ_max, hidden causes transmitting predictive information) and on an unproven heuristic principle that stable graphs imply adequate Markov representation. No calibration, null model, or power analysis is used to validate this premise; in fact the abstract claims a bootstrap calibration that the body says was not performed. Experimental settings (thresholds, permutation counts, depth grid, ROI subset) are chosen by hand without sensitivity analysis. No new physical or model entities are invented: the latent process C_t is explicitly a conceptual representation "not as an identifiable component of the model" (Section 3.1).

free parameters (4)
  • c-GC / c-GC* significance thresholds α and β = α=0.01, β=0.001
    Chosen by hand with no sensitivity analysis; the diagnostic's separation between Markovian and non-Markovian curves depends on these thresholds (Section 4.1).
  • Non-Markovian smooth driver parameters = random amplitudes, centers, widths (unspecified)
    The hidden-memory DGP that generates the positive-instability signal is described only verbally (Section 4.1); no equation, distribution, or explicit causal role is given.
  • Conditioning-depth grid = npasts = 1..7 (p = 2..7 for multi-lag)
    The grid extent determines which memory lengths the diagnostic can detect; no justification is given that τ_max ≤ 7 for any experiment.
  • fish-3 ROI subset size = 130 of 420 ROIs
    ROIs selected by strongest correlation with recorded behaviour; a data-selection choice that shapes the real-data summaries (Section 4.2).
axioms (5)
  • domain assumption Causal Markov assumption: each variable is independent of its non-effects given its direct causes.
    Stated in Section 3.3 as required for the diagnostic's interpretation; also the basis of constraint-based structure learning.
  • domain assumption Faithfulness: observed conditional independencies arise from graphical separation, not parameter cancellations.
    Invoked in Section 2.2 as a requirement of constraint-based learners; violated faithfulness would break the link between graph and distribution.
  • ad hoc to paper Heuristic principle: for p ≥ τ, enlarging conditioning depth from p to p+1 should not systematically change the population graph.
    Section 3.2; an unproven premise that the learner's population output is insensitive to extra conditioned lags beyond the Markov order.
  • domain assumption A stabilizing depth exists within the tested grid (hidden memory shorter than p_max).
    Section 3.7 interpretation rule 5 assumes instability "before saturating"; if hidden memory exceeds the grid, no saturation occurs and the rule is ambiguous.
  • ad hoc to paper Hidden causes act by transmitting additional predictive information not summarized by the shallow observed state.
    Section 3.3; this defines the mechanism by which latent confounding is supposed to produce the depth signature.

reviewed 2026-08-03 · how reviews work

0 comments
Cite this review

Pith. "Pith review of Conditioning-Depth Diagnostics for Hidden Memory in Temporal Causal Discovery." pith.science (2026). https://pith.science/paper/DAYRVGE3

@misc{pith2026260601214,
  author       = {Pith},
  title        = {Pith review of: Conditioning-Depth Diagnostics for Hidden Memory in Temporal Causal Discovery},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/DAYRVGE3}},
  note         = {Machine review of arXiv:2606.01214}
}
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Reliable causal discovery in timeseries requires conditioning sets that capture the system state. When predictive history is omitted, residual dependence can appear as direct causal links. We test state adequacy by measuring how inferred graphs change as conditioning depth increases while the reported causal lag stays fixed. Under an adequate finite-order Markov representation, graphs should stabilize once enough observed history is conditioned on; latent common drive, omitted lags, nonstationarity, and measurement dynamics can instead produce depth sensitivity. We formalize this idea with graph instability statistics and evaluate c-GC and c-GC*, the two learners whose depth parameter implements a matched fixed-horizon history intervention. PCMCI+ and JPCMCI+ are excluded from the primary comparison because adaptive parent selection makes nominal depth edge-specific. In paired simulations, a clean order-1 process was stable in every repeat, whereas an AR(1) latent common driver produced positive instability in all c-GC repeats and eight of ten c-GC* repeats. In calcium imaging recordings, connectivity drops at the first transition beyond the one-lag baseline and then levels off, but B=200 bootstrap calibration does not reject the fitted order-1 null. The workflow therefore flags hidden memory or observed-state inadequacy without identifying the generating mechanism or recovering a latent graph.

Figures

Figures reproduced from arXiv: 2606.01214 by S. A. Adedayo.

Figure 1
Figure 1. Figure 1: Schematic intuition for the Markovianity [PITH_FULL_IMAGE:figures/full_fig_p007_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Conditioning depth sensitivity metrics for the single lag Markovian simulation, [PITH_FULL_IMAGE:figures/full_fig_p012_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Conditioning depth sensitivity metrics for the single lag non-Markovian simulation with a smooth [PITH_FULL_IMAGE:figures/full_fig_p012_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Conditioning depth sensitivity metrics for the Markovian simulation with multiple lags, [PITH_FULL_IMAGE:figures/full_fig_p013_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: Conditioning depth sensitivity metrics for the non-Markovian simulation with multiple lags and a [PITH_FULL_IMAGE:figures/full_fig_p013_5.png] view at source ↗
Figure 7
Figure 7. Figure 7: visualizes the decomposition into D− p and D+ p clarifies why this should be read as a depth sensitivity result rather than a simple plot of edge counts. For every recording and both methods, the largest graph change is the first transition from p = 1 to p = 2, and deletions dominate additions at that transition. Under c-GC, deletions account for 61.7% to 81.6% of [PITH_FULL_IMAGE:figures/full_fig_p014_7.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Blind Source Separation Can Distort Behavior and Connectivity Analyses of Calcium Transients

    stat.AP 2026-08 conditional novelty 6.0

    Component-removing blind source separation can collapse causal graph recovery to zero in synthetic calcium traces and dramatically densify connectivity estimates from real v2a-RSN traces, so BSS denoising is not a neu...

Reference graph

Works this paper leans on

6 extracted references · 4 linked inside Pith · cited by 1 Pith paper

  1. [1]

    Adedayo, S. A. (2025). Re-examining granger causality with causal bayesian networks and reichenbach’s principles.arXiv preprint https://arxiv.org/pdf/2501.02672v2. Ahrens, M. B., Orger, M. B., Robson, D. N., Li, J. M., and Keller, P. J. (2013). Whole-brain functional imaging at cellular resolution using light- sheet microscopy.Nature Methods, 10(5):413–42...

  2. [30]

    Maeda, T. N. and Shimizu, S. (2020). RCD: Repetitive causal discovery of linear non-gaussian acyclic models with latent confounders. InProceedings of the Twenty Third International Conference on Artificial Intelligence and Statistics, volume 108 of Proceedings of Machine Learning Research, pages 735–745. PMLR. Malinsky, D. and Spirtes, P. (2018). Causal s...

  3. [313]

    Geweke, J. F. (1984). Measures of conditional linear dependence and feedback between time series. Journal of the American Statistical Association, 79(388):907–915. Glymour, C., Zhang, K., and Spirtes, P. (2019). Review of causal discovery methods based on graphical models.Frontiers in Genetics, 10:524. 16 Granger, C. W. J. (1969). Investigating causal rel...

  4. [1713]

    and Runge, J

    Gerhardus, A. and Runge, J. (2020). High-recall causal discovery for autocorrelated time series with latent confounders. InAdvances in Neural Information Processing Systems, volume 33, pages 12615–12625. Geweke, J. (1982). Measurement of linear dependence and feedback between multiple time series.Journal of the American Statistical Association, 77(378):304–

  5. [1790]

    Chen, L., Li, C., Shen, X., and Pan, W. (2024a). Discovery and inference of a causal network with hidden confounding.Journal of the American Statistical Association, 119(548):2572–2584. Chen, W., Huang, Z., Cai, R., Hao, Z., and Zhang, K. (2024b). Identification of causal structure with latent variables based on higher order cumulants. Proceedings of the ...

  6. [2023]

    and Ramsey, J

    Kummerfeld, E. and Ramsey, J. (2016). Causal clustering for 1-factor measurement models. Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, pages 1655–1664. Künsch, H. R. (1989). The jackknife and the bootstrap for general stationary observations.Annals of Statistics, 17(3):1217–1241. Kuroki, M. and Pearl, ...

This paper was first reviewed by deepseek-v4-flash on August 3, 2026.