Pith. sign in

REVIEW 1 major objections 4 minor 89 references

Bandable Cumulant Tensors: Optimal Estimation and Applications in Non-Gaussian Data Modeling

T0 review · 1 major / 4 minor · reviewed 2026-08-14 · deepseek-v4-flash

Pith's one-line read For ordered data with decaying dependence, a tapered sample cumulant estimates the order-d tensor at the minimax rate √(k*/n), with p entering only through log p.

desk verdict Minimax core for bandable cumulant tensors is new and sound; the abstract's 'direct transfer' to applications outruns the theorems' i.i.d. condition, though the paper flags the gap itself. read the letter →

arxiv 2608.10161 v1 pith:DGICP3EU submitted 2026-08-10 stat.ME

classification stat.ME MSC 62G0562H1262M1062C20
keywords higher-ordercumulantscumulanttensorsbandablestructuretaperedestimationtensorspectralnormminimaxnon-Gaussiantimeseriessourcelocalization
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Higher-order cumulants capture non-Gaussian dependence—asymmetry, phase, nonlinear interaction—that covariance cannot see, but their tensors have p^d entries and the standard plug-in estimator is not even rate-optimal under the tensor spectral norm. The paper's central claim is that for ordered coordinates (time, space, frequency, genome), where dependence is strongest locally, both difficulties yield to one assumption: the cumulant tensor is bandable, with entries decaying as they move away from the main tensor diagonal. Under that assumption, the tapered sample cumulant, which computes only O($pk^{{d-1}}$) entries in a diagonal tube of bandwidth k, splits its error into a tapering bias $βk^{{-α-d/2+1}}$ plus a local stochastic term, and at the oracle bandwidth k*=($β^{2}$ n)^{1/(2α+d-1)} attains the minimax spectral-norm rate √(k*/n) for sub-Gaussian data, with the ambient dimension entering only through log p. This matters because it is, the paper argues, the first minimax theory for estimating higher-order cumulant tensors under structural decay, and because the spectral-norm guarantee transfers directly to plug-in procedures: cumulant Yule–Walker estimation in autoregressions, minimum-distance estimation in moving-average models, and matched-filter localization of non-Gaussian sources in sensor arrays.

What carries the argument

The load-bearing object is the bandable cumulant class K^d_{α,β}: order-d tensors whose entries decay as β(1+diam(i))^{-(α+d-1)}, where diam(i) = max_{a,b}|i_a - i_b| measures distance from the main tensor diagonal. The estimator is the tapered sample cumulant bK_{d,T,k} = W_k * bK_d, with weights w_k(m)=1 for m≤k/2, linear down to zero at m=k. The argument rides on two structural facts: a diagonal tube of width k contains O($pk^{{d-1}}$) entries, so the bias from truncating the tail is $βk^{{-α-d/2+1}}$; and localization replaces ambient-dimension fluctuations with local fluctuations governed by k, so the stochastic error is controlled by Δ_{k,x} = √((k+log p+x)/n) + (k+log p+x)^{d/γ}/n. A counting identity represents the taper as a difference of two sliding-window block sums, which lets the proof control the tapering operator through local tensor norms together with concentration inequalities for polynomials in sub-exponential variables.

What would settle it

Run the tapered estimator over a grid of bandwidths on observations from a process whose true cumulant tensor is a single rank-one component spread uniformly over all p coordinates, so K_d = $λu^{{⊗d}}$ with |u_i| = $p^{{-1/2}}$ for every i; such a tensor is not bandable. If the spectral-norm error follows the unstructured √(p/n) scale rather than min_k {$βk^{{-α-d/2+1}}$ + √((k+log p)/n)}, the bandability premise is confirmed as load-bearing, and the theorem's rate is a property of the class K^d_{α,β} rather than of all cumulant tensors.

Watch

Extended reading notes

Core claim

The paper establishes that the order-d cumulant tensor is estimable at a structured minimax rate whenever it is (α,β)-bandable, meaning |(K_d)_i| ≤ β(1+diam(i))^{-(α+d-1)}. The tapered estimator bK_{d,T,k} = W_k * bK_d multiplies the plug-in sample cumulant by weights that are one inside a diagonal tube of radius k/2 and decay linearly to zero at radius k. Theorem 1 separates the error into two controlled parts: E||bK_{d,T,k} - K_d|| ≲ $βk^{{-α-d/2+1}}$ + Δ_{k,0}(1+Δ_{k,0})^{d-1}, where Δ_{k,0} = √((k+log p)/n) + (k+log p)^{d/γ}/n under exponential-type Φ_γ tails. Theorem 2 matches this with a minimax lower bound over sub-Gaussian bandable distributions, yielding the oracle bandwidth k* and rate √(k*/n) when n ≫ (k*+log p)^{d-1}. The key mechanism is that localization replaces the global fluctuation $p^{{d/γ}}$/n that makes raw plug-in cumulants rate-suboptimal with the local fluctuation (k+log p)^{d/γ}/n, so tapering turns the simple plug-in estimator into a minimax rate-optimal one.

Load-bearing premise

The entire rate is purchased by the assumption that the true order-d cumulant tensor is (α,β)-bandable, meaning entries decay polynomially as β(1+diam(i))^{-(α+d-1)} with unknown α and β; if dependence is more long-range or decays more slowly, the tapering bias term dominates and the minimax claim does not apply.

Editorial extensions

If this is right

  • For ordered non-Gaussian data with k+log p ≪ p, the spectral-norm error scales like √(k*/n) rather than the unstructured √(p/n), so bandability turns a seemingly intractable high-dimensional tensor problem into a tractable one.
  • The spectral-norm error bound transfers to parameters: autoregressive coefficients via cumulant Yule–Walker equations, moving-average coefficients via minimum-distance estimation, and source locations via matched filtering are all controlled by the tensor spectral-norm error up to explicit stability constants.
  • Because cumulants of order d≥3 vanish for Gaussian variables, the tapered cumulant equations and localization scores are insensitive to additive Gaussian measurement noise, unlike covariance-based baselines.
  • The estimator runs in O(npk^{d-1}) time and O(pk^{d-1}) memory without ever forming the full tensor, and a Lepski-type stability rule selects the bandwidth from data.
  • In the Neuropixels spike-localization benchmark, the tapered third-order matched-filter score matches the accuracy of the raw third-order score across cell-condition samples while storing roughly five times fewer range summaries, and it localizes more accurately than the covariance-based score.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Inference: the same bias-variance separation should carry to frequency-domain higher-order spectra, since tapering lag cumulants before the Fourier transform gives a lag-window bispectrum estimator whose error splits into the same two terms; the paper sketches this but does not prove a matching minimax result for the bispectrum.
  • Inference: if the coordinate ordering is unknown, one could try to learn a permutation that maximizes bandability, for example by minimizing diameter-decay residuals; this would connect the framework to seriation and manifold-learning problems, but the paper assumes the ordering is known.
  • Inference: the matching lower bound is proved for sub-Gaussian laws (γ=2); for heavier-tailed data with γ<2, the higher-order term (k+log p)^{d/γ}/n may dominate, and the paper does not determine whether the resulting rate is minimax, so a natural next step is a lower bound for exponential-type tails with γ<2.
  • Inference: the main theorem is stated for i.i.d. replicates of a lag vector, while the real-data time series are single long trajectories; extending the concentration argument to dependent data would justify the same guarantees for the air-quality application, which currently relies on approximately independent blocks.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

1 major / 4 minor

Summary. This paper introduces an (α,β)-bandability model for order-d cumulant tensors of ordered p-dimensional random vectors, in which entries decay as β(1+diam(i))^{-(α+d-1)} (Definition 1), and proposes a tapered sample cumulant estimator K̂_{d,T,k} that computes only O(pk^{d-1}) retained entries. Theorem 1 gives a nonasymptotic spectral-norm upper bound that separates the tapering bias βk^{-α-d/2+1} from a localized stochastic error Δ_{k,x}(1+Δ_{k,x})^{d-1}; Theorem 2 establishes a matching minimax lower bound over a sub-Gaussian bandable class, built from a tilted-Gaussian rank-one perturbation that respects the entrywise bandability constraint. For sub-Gaussian observations (γ=2), oracle bandwidth k*=(β^2 n)^{1/(2α+d-1)}, and log p ≲ k*, the tapered estimator attains the rate √(k*/n), which is minimax optimal when n ≫ (k*+log p)^{d-1}. The paper further derives plug-in perturbation bounds for cumulant Yule–Walker estimation in AR models (Theorem 3), minimum-distance estimation in MA models (Theorem 4), and spatial matched-filter source localization (Theorem 5), each stated for i.i.d. observations and conditional on the tensor spectral-norm error. Simulations study the estimator, a Lepski-type bandwidth selector, and the time-series estimators; real-data analyses of RR intervals, air-quality sensors, and Neuropixels waveforms illustrate the methodology, with full preprocessing details in Appendix D and complete proofs in the Supplement.

Significance. The bandable-cumulant minimax theory appears to be the first of its kind, and it is a genuine extension rather than a matrix analogue: the diagonal-tube cardinality O(pk^{d-1}), the bias exponent α+d/2−1, and the higher-order fluctuation (k+log p)^{d/γ}/n each differ from bandable covariance estimation, and the claim that tapering suppresses the plug-in cumulant's non-optimal higher-order fluctuations is interesting and plausible. I checked the load-bearing algebra of Theorem 1 (dyadic annulus bias bound, block-diagonal decomposition in Eq. (32)), Section 3.3 (oracle bandwidth and rate), and Theorem 2 (perturbation amplitude θ ≤ βk^{-α-d/2+1}, Fano separation, KL calibration); the derivations are internally consistent, and the lower bound is a genuine least-favorable construction rather than a circular restatement of the upper bound. The Supplement proves all supporting lemmas, including the tilted-Gaussian perturbation claim, and documents the real-data pipelines in detail; the paper also correctly credits the companion unstructured result [75] for the rank-one perturbation technique, which is here adapted to the bandable geometry.

major comments (1)
  1. [Theorems 3–5 (Eq. (11)); Remark 4; §7.1–7.2; §8] The transfer theorems are stated only for i.i.d. replicates of the relevant random vector, while the abstract and Introduction claim that the spectral-norm guarantee 'transfers directly to downstream tasks.' Theorem 3 and Theorem 4 explicitly assume 'n i.i.d. copies of the marginal distribution of the lag vector Z_t', and Theorem 5 assumes 'i.i.d. copies of the sensor-array snapshot X'. The RR-interval analysis (Section 7.1) and the air-quality analysis (Section 7.2) instead form empirical cumulants as time averages over single long dependent trajectories (subject-wise in Section 7.1, one processed series in Section 7.2), so neither the high-probability bounds of Theorem 1 nor the downstream bounds of Theorems 3–4 apply to those analyses as written. The manuscript itself concedes the gap: Remark 4 states that 'a single stationary trajectory requires a dependent-data analogue,' and Section 8 lists 'extend the concentration theory from independent observations to dependent time-series data' as an open direction. No mixing condition, physical-dependence bound, or effective-sample-size argument is provided to close this gap. Because the unqualified transfer claim is one of the paper's three stated main contributions and is repeated in the abstract, this is load-bearing for the paper's broader framing rather than a local presentation issue. The fix is within scope: either prove the concentration theory under explicit dependence assumptions (for example, physical dependence in the sense used in [86]) for the AR/MA lag vectors, or consistently reword the abstract, Introduction, and Sections 7.1–7.2 to state that rigorous guarantees hold for i.i.d. replicated observations and that the dependent-data analyses are empirical illustrations.
minor comments (4)
  1. [§4.1 (Theorem 3)] The AR application never states an identifiability condition on the innovations' cumulant: Theorem 3 assumes σ_min(A) > 0, but by Lemma 1 every entry of A is proportional to τ_d, so σ_min(A) = 0 whenever τ_d = 0. Since the MA section explicitly assumes τ_d ≠ 0, the AR section should likewise state this assumption before Theorem 3.
  2. [Abstract vs §3.3] The abstract's condition 'n ≫ (k+log p)^{d-1}' for attaining the minimax rate omits the companion conditions stated in Section 3.3 (log p ≲ k* and k_0 ≤ k* ≤ p); aligning the abstract with the theorem's stated regime would avoid an impression that the rate claim is unconditional.
  3. [§7.1 (Figure 8d)] The supplement (Appendix D.1) reports paired subject-bootstrap 95% intervals for the mean portmanteau-score differences, but Figure 8d displays only raw lines and means; displaying these intervals would make the claim that the tapered score is lowest on average directly assessable.
  4. [Appendix F (Lemma 6); §1.1; reference [65]] In the proof of Lemma 6, 'Holder's inequality' should be 'Hölder's inequality'; in Section 1.1 and reference [65], 'PARAF AC' should be 'PARAFAC' (the current text is a typographical split of 'PARAFAC').

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the minimax and plug-in transfer theorems are derived from stated assumptions rather than equivalent to their conclusions.

full rationale

The derivation chain is not circular. Definition 1 is the input model, and Theorem 1's upper bound is proved by a genuine decomposition into tapering bias and stochastic error: the bias term is obtained by a dyadic annulus summation of (W_k-1)*K_d, and the stochastic term follows from local spectral-norm concentration over block-local empirical cumulants, using net arguments and Orlicz tail bounds (Supplement Lemmas 8-10 and the partition-telescoping argument in E.2). The upper bound is not an algebraic restatement of the bandable class definition. Theorem 2's lower bound is also self-contained: it constructs least-favorable distributions by Hermite-tilted Gaussian perturbations, calibrates the perturbation amplitude so that the resulting order-d cumulant tensor lies in the same bandable class, and applies Fano's inequality. The citation [75] appears only as motivation and comparison ('The construction is related to rank-one perturbation methods for unstructured high-order cumulant estimation [75]'); the actual perturbation claim is proved in the supplement, so that self-citation is not load-bearing. The oracle bandwidth k*=(beta^2 n)^{1/(2alpha+d-1)} is a theoretical balancing point between the bias and stochastic terms, not a fitted parameter, and the simulations use oracle k only as a benchmark while evaluating the Lepski-type selector separately. The downstream results are deterministic perturbation bounds that transfer spectral-norm error into parameter or localization error, and their high-probability versions are explicitly stated for i.i.d. copies; Remark 4 and Section 8 acknowledge that dependent time series require separate concentration theory. This is a scoping limitation of the applications, not circularity, because the real-data dependent trajectories do not feed back into the minimax theorems.

Assumptions & free parameters 3 free parameters · 4 assumptions · 0 invented entities

The central result rests on the bandability class, an exponential-tail condition, and standard concentration/Fano machinery. The bandwidth is the main free parameter; the practical selector is heuristic. No new physical entities are introduced, and no parameter is fitted to data in order to state the minimax rate.

free parameters (3)
  • bandwidth k = oracle k* = (β^2 n)^{1/(2α+d-1)}; Lepski-selected in practice
    The minimax guarantee requires k at the oracle order. In practice k is chosen from a grid by a Lepski-type stability rule with threshold A=1, but no adaptation theorem is proved.
  • Lepski threshold multiplier A = A=1
    Chosen by hand as the default in all numerical work based on simulation; no theoretical criterion for choosing A is provided.
  • bandability parameters α and β = assumed known in theory; unknown in practice
    The class K^d_{α,β} and the oracle bandwidth formula depend on α and β. The theorems treat them as fixed and known; no estimation or adaptation procedure for α and β is developed.
assumptions (4)
  • domain assumption The order-d cumulant tensor is (α,β)-bandable: |(K_d)_i| ≤ β(1+diam(i))^{-(α+d-1)} for all multi-indices (Definition 1, Section 2.2).
    This is the entire bandability premise. The tapering bias term βk^{-α-d/2+1} and the oracle bandwidth formula follow from it, so the central claim fails without it.
  • domain assumption Exponential-type tails: sup_{u∈S^{p-1}} ||u^T X||_{Φγ} ≤ K for some 0<γ≤d.
    Used in the local empirical moment and cumulant concentration bounds (Lemma 10 and Lemma 8). It restricts the class to sub-Gaussian or sub-exponential observations and excludes heavy tails.
  • standard math Standard concentration and Fano tools: Orlicz concentration inequalities from [31], net arguments, and generalized Fano's inequality.
    Imported background results. They are not proved in the paper and are used in the upper and lower bound arguments.
  • domain assumption For the applications, AR/MA/sensor models satisfy the bandability lemmas (Lemmas 1-3), and Gaussian contamination has zero cumulants of order d≥3.
    These assumptions justify the cumulant Yule-Walker, minimum-distance, and matched-filter plug-in analyses. They are standard for the models but are restrictions on the data-generating process.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Bandable Cumulant Tensors: Optimal Estimation and Applications in Non-Gaussian Data Modeling." pith.science (2026). https://pith.science/paper/DGICP3EU

@misc{pith2026260810161,
  author       = {Pith},
  title        = {Pith review of: Bandable Cumulant Tensors: Optimal Estimation and Applications in Non-Gaussian Data Modeling},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/DGICP3EU}},
  note         = {Machine review of arXiv:2608.10161}
}
abstract

Higher-order cumulants capture the non-Gaussian dependence that covariance misses, but they are hard to use in high dimensions. An order-$d$ cumulant tensor has $p^d$ entries, and the plug-in sample cumulant is generally not even rate-optimal under the tensor spectral norm. For ordered data, both difficulties admit one remedy: assuming that higher-order interactions decay away from the main tensor diagonal, we introduce a bandable cumulant class and a tapered sample cumulant estimator that computes only $O(pk^{d-1})$ local entries at bandwidth $k$ and never forms the full tensor. Under exponential-type tail conditions, we prove nonasymptotic spectral-norm bounds that separate tapering bias from stochastic error, and a minimax lower bound over the same class that matches the leading bias and stochastic terms of the upper bound; for sub-Gaussian observations, the tapered estimator attains the minimax rate whenever $n\gg(k+\log p)^{d-1}$ at the oracle bandwidth $k$, with the ambient dimension entering only through $\log p$. Localization also suppresses the higher-order fluctuations behind this suboptimality, so tapering plays a stronger role here than in bandable covariance estimation. The spectral-norm guarantee transfers directly to downstream tasks, yielding plug-in error bounds for cumulant Yule--Walker estimation in autoregressive models, minimum-distance estimation in moving-average models, and matched-filter source localization in sensor arrays. Simulations corroborate the theory, and real-data analyses of RR-interval, air-quality, and Neuropixels recordings illustrate the resulting stability gains.

Figures

Figures reproduced from arXiv: 2608.10161 by the authors.

Figure 1
Figure 1. Illustration of banding and tapering for an order-3 cumulant tensor. [PITH_FULL_IMAGE:figures/full_fig_p009_1.png] view at source ↗
Figure 2
Figure 2. Mean tensor estimation error as a function of the sample size [PITH_FULL_IMAGE:figures/full_fig_p028_2.png] view at source ↗
Figure 3
Figure 3. Mean tensor estimation error as a function of the dimension [PITH_FULL_IMAGE:figures/full_fig_p028_3.png] view at source ↗
Figures from the paper (9 more)
Figure 4
Figure 4. Figure 4: Risk ratio for data-driven bandwidth selection. [PITH_FULL_IMAGE:figures/full_fig_p030_4.png]
Figure 5
Figure 5. Figure 5: AR coefficient estimation. 0.04 0.06 0.08 0.10 500 1000 2000 Sample size Mean L2 coefficient error MA minimum−distance estimation from third−order cumulants [PITH_FULL_IMAGE:figures/full_fig_p031_5.png]
Figure 6
Figure 6. Figure 6: MA coefficient estimation using the third-order cumulant minimum-distance estima [PITH_FULL_IMAGE:figures/full_fig_p031_6.png]
Figure 7
Figure 7. Figure 7: Bandwidth selection [PITH_FULL_IMAGE:figures/full_fig_p032_7.png]
Figure 8
Figure 8. Figure 8: RR interval analysis. 33 [PITH_FULL_IMAGE:figures/full_fig_p033_8.png]
Figure 9
Figure 9. Figure 9: Preprocessing and order selection [PITH_FULL_IMAGE:figures/full_fig_p035_9.png]
Figure 10
Figure 10. Figure 10: Air Quality analysis. 7.3 Spatial Localization on Paired Patch-Clamp and Neuropixels Waveforms We evaluate the proposed tapered cumulant localization procedure in Section 4.3 on the SPE-1 paired patch-clamp and Neuropixels benchmark archive [87, 88]. Neuropixels probe…
Figure 11
Figure 11. Figure 11: SPE-1 localization experiment. 8 Discussion This paper develops a bandability framework for estimating higher-order cumulant tensors in ordered high-dimensional problems. The main idea is that, in many non-Gaussian settings, higher-order dependence is strongest among …
Figure 12
Figure 12. Figure 12: (b) shows why the tapered score can also be preferable for individual samples. In the c28, D11 example, the raw cumulant score has a large off-target peak at 2660 µm. The ground-truth depth is 2373.4 µm, so the raw estimate has error 286.6 µm. The tapered score suppre…

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

89 extracted references · 65 canonical work pages

  1. [75]

    Detection is harder than estimation in cer- tain regimes: Inference for moment and cumulant tensors.arXiv preprint arXiv:2603.26029, 2026

    Runshi Tang, Yuefeng Han, and Anru R Zhang. Detection is harder than estimation in cer- tain regimes: Inference for moment and cumulant tensors.arXiv preprint arXiv:2603.26029, 2026

  2. [86]

    Asymptotic theory for estimators of high-order statistics of stationary processes.IEEE Transactions on Information Theory, 64(7):4907–4922, 2018

    Danna Zhang and Wei Biao Wu. Asymptotic theory for estimators of high-order statistics of stationary processes.IEEE Transactions on Information Theory, 64(7):4907–4922, 2018

  3. [1]

    Allen and Robert Tibshirani

    Genevera I. Allen and Robert Tibshirani. Transposable regularized covariance models with an application to missing data imputation.The Annals of Applied Statistics, 4(2):764–790, 2010

  4. [2]

    Tensor decompositions for learning latent variable models.J

    Animashree Anandkumar, Rong Ge, Daniel J Hsu, Sham M Kakade, Matus Telgarsky, et al. Tensor decompositions for learning latent variable models.J. Mach. Learn. Res., 15(1):2773–2832, 2014

  5. [3]

    A method of moments for mixture models and hidden markov models

    Animashree Anandkumar, Daniel Hsu, and Sham M Kakade. A method of moments for mixture models and hidden markov models. InConference on learning theory, pages 33–1. JMLR Workshop and Conference Proceedings, 2012

  6. [4]

    Anderson.An Introduction to Multivariate Statistical Analysis

    Theodore W. Anderson.An Introduction to Multivariate Statistical Analysis. Wiley, Hobo- ken, NJ, 3 edition, 2003

  7. [5]

    Ashley, Douglas M

    Richard A. Ashley, Douglas M. Patterson, and Melvin J. Hinich. A diagnostic test for nonlinear serial dependence in time series fitting errors.Journal of Time Series Analysis, 7(3):165–178, 1986

  8. [6]

    Large-dimensional independent component analysis: Sta- tistical optimality and computational tractability.The Annals of Statistics, 53(2):477–505, 2025

    Arnab Auddy and Ming Yuan. Large-dimensional independent component analysis: Sta- tistical optimality and computational tractability.The Annals of Statistics, 53(2):477–505, 2025. 38

Show all 89 references
  1. [7]

    Ilmoniemi, Gian Luca Romani, Vittorio Pizzella, and Laura Marzetti

    Alessio Basti, Guido Nolte, Roberto Guidotti, Risto J. Ilmoniemi, Gian Luca Romani, Vittorio Pizzella, and Laura Marzetti. A bicoherence approach to analyze multi-dimensional cross-frequency coupling in eeg/meg data.Scientific Reports, 14:8461, 2024

  2. [8]

    Bickel and Elizaveta Levina

    Peter J. Bickel and Elizaveta Levina. Regularized estimation of large covariance matrices. The Annals of Statistics, 36(1):199–227, 2008

  3. [9]

    Brillinger

    David R. Brillinger. The identification of a particular nonlinear time series system. Biometrika, 64(3):509–515, 1977

  4. [10]

    Brillinger.Time Series: Data Analysis and Theory

    David R. Brillinger.Time Series: Data Analysis and Theory. Holden-Day, San Francisco, expanded edition, 1981

  5. [11]

    Brockwell and Richard A

    Peter J. Brockwell and Richard A. Davis.Time Series: Theory and Methods. Springer, New York, 2 edition, 1991

  6. [12]

    Buccino, Cole L

    Alessio P. Buccino, Cole L. Hurwitz, Samuel Garcia, Jeremy Magland, Joshua H. Siegle, Roger Hurwitz, and Matthias H. Hennig. SpikeInterface, a unified framework for spike sorting.eLife, 9:e61834, 2020

  7. [13]

    Byrd, Peihuang Lu, Jorge Nocedal, and Ciyou Zhu

    Richard H. Byrd, Peihuang Lu, Jorge Nocedal, and Ciyou Zhu. A limited memory algorithm for bound constrained optimization.SIAM Journal on Scientific Computing, 16(5):1190– 1208, 1995

  8. [14]

    Tony Cai, Zhao Ren, and Harrison H

    T. Tony Cai, Zhao Ren, and Harrison H. Zhou. Estimating structured high-dimensional co- variance and precision matrices: Optimal rates and adaptive estimation.Electronic Journal of Statistics, 10(1):1–59, 2016

  9. [15]

    Tony Cai, Cun-Hui Zhang, and Harrison H

    T. Tony Cai, Cun-Hui Zhang, and Harrison H. Zhou. Optimal rates of convergence for covariance matrix estimation.The Annals of Statistics, 38(4):2118–2144, 2010

  10. [16]

    Tony Cai and Harrison H

    T. Tony Cai and Harrison H. Zhou. Minimax estimation of large covariance matrices under ℓ1 norm.Statistica Sinica, 22(4):1319–1349, 2012

  11. [17]

    High-order contrasts for independent component analysis.Neural Computation, 11(1):157–192, 1999

    Jean-Fran¸ cois Cardoso. High-order contrasts for independent component analysis.Neural Computation, 11(1):157–192, 1999

  12. [18]

    Blind beamforming for non-gaussian sig- nals.IEE Proceedings F: Radar and Signal Processing, 140(6):362–370, 1993

    Jean-Fran¸ cois Cardoso and Antoine Souloumiac. Blind beamforming for non-gaussian sig- nals.IEE Proceedings F: Radar and Signal Processing, 140(6):362–370, 1993

  13. [19]

    Stability and asymptotics for autoregressive processes

    Likai Chen and Wei Biao Wu. Stability and asymptotics for autoregressive processes. Electronic Journal of Statistics, 10(2):3723–3751, 2016

  14. [20]

    Independent component analysis, a new concept?Signal processing, 36(3):287–314, 1994

    Pierre Comon. Independent component analysis, a new concept?Signal processing, 36(3):287–314, 1994. 39

  15. [21]

    A multilinear singular value decomposition.SIAM Journal on Matrix Analysis and Applications, 21(4):1253–1278, 2000

    Lieven De Lathauwer, Bart De Moor, and Joos Vandewalle. A multilinear singular value decomposition.SIAM Journal on Matrix Analysis and Applications, 21(4):1253–1278, 2000

  16. [22]

    The MLE algorithm for the matrix normal distribution.Journal of Sta- tistical Computation and Simulation, 64(2):105–123, 1999

    Pierre Dutilleul. The MLE algorithm for the matrix normal distribution.Journal of Sta- tistical Computation and Simulation, 64(2):105–123, 1999

  17. [23]

    Asymptotically optimal estimation of ma and arma parameters of non-gaussian processes from high-order moments.IEEE Transactions on Automatic Control, 35(1):27–35, 1990

    Benjamin Friedlander and Boaz Porat. Asymptotically optimal estimation of ma and arma parameters of non-gaussian processes from high-order moments.IEEE Transactions on Automatic Control, 35(1):27–35, 1990

  18. [24]

    Giannakis

    Georgios B. Giannakis. On the identifiability of non-gaussian ARMA models using cumu- lants.IEEE Transactions on Automatic Control, 35(1):18–26, 1990

  19. [25]

    Giannakis, Yoshio Inouye, and Jerry M

    Georgios B. Giannakis, Yoshio Inouye, and Jerry M. Mendel. Cumulant based identifica- tion of multichannel moving-average models.IEEE Transactions on Automatic Control, 34(7):783–787, 1989

  20. [26]

    Bandwidth selection in kernel density esti- mation: Oracle inequalities and adaptive minimax optimality.The Annals of Statistics, 39(3):1608–1632, 2011

    Alexander Goldenshluger and Oleg Lepski. Bandwidth selection in kernel density esti- mation: Oracle inequalities and adaptive minimax optimality.The Annals of Statistics, 39(3):1608–1632, 2011

  21. [27]

    Minimum distance estimation of possibly noninvertible moving average models.Journal of Business & Economic Statistics, 33(3):403–417, 2015

    Nikolay Gospodinov and Serena Ng. Minimum distance estimation of possibly noninvertible moving average models.Journal of Business & Economic Statistics, 33(3):403–417, 2015

  22. [28]

    Kristjan Greenewald and Alfred O. Hero. Robust Kronecker product PCA for spatio- temporal covariance estimation.IEEE Transactions on Signal Processing, 63(23):6368– 6378, 2015

  23. [29]

    Kristjan Greenewald, Theodoros Tsiligkaridis, and Alfred O. Hero. Kronecker sum decom- positions of space-time data. In2013 IEEE 5th International Workshop on Computational Advances in Multi-Sensor Adaptive Processing, pages 65–68, 2013

  24. [30]

    Kristjan Greenewald, Shuheng Zhou, and Alfred O. Hero. Tensor graphical lasso (Ter- aLasso).Journal of the Royal Statistical Society: Series B (Statistical Methodology), 81(5):901–931, 2019

  25. [31]

    Concentration inequalities for polyno- mials inα-sub-exponential random variables.Electronic Journal of Probability, 26(none):1– 22, January 2021

    Friedrich G¨ otze, Holger Sambale, and Arthur Sinulis. Concentration inequalities for polyno- mials inα-sub-exponential random variables.Electronic Journal of Probability, 26(none):1– 22, January 2021. Publisher: Institute of Mathematical Statistics and Bernoulli Society

  26. [32]

    Hamilton.Time Series Analysis

    James D. Hamilton.Time Series Analysis. Princeton University Press, Princeton, NJ, 1994. 40

  27. [33]

    Platten, Jan R

    David Hickey, Keith Worden, Michael F. Platten, Jan R. Wright, and Jonathan E. Cooper. Higher-order spectra for identification of nonlinear modal coupling.Mechanical Systems and Signal Processing, 23(4):1037–1061, 2009

  28. [34]

    Hillar and Lek-Heng Lim

    Christopher J. Hillar and Lek-Heng Lim. Most tensor problems are NP-hard.Journal of the ACM, 60(6):45:1–45:39, 2013

  29. [35]

    Melvin J. Hinich. Testing for dependence in the input to a linear time series model.Journal of Nonparametric Statistics, 6(2–3):205–221, 1996

  30. [36]

    Hinich and G

    Melvin J. Hinich and G. R. Wilson. Time delay estimation using the cross bispectrum. IEEE Transactions on Signal Processing, 40(1):106–113, 1992

  31. [37]

    Peter Hoff, Andrew McCormack, and Anru R. Zhang. Core shrinkage covariance estima- tion for matrix-variate data.Journal of the Royal Statistical Society Series B: Statistical Methodology, 85(5):1659–1679, 2023

  32. [38]

    M. Huzii. Estimation of coefficients of an autoregressive process by using a higher order moment.Journal of Time Series Analysis, 2(2):87–93, 1981

  33. [39]

    Wiley, New York, 2001

    Aapo Hyv¨ arinen, Juha Karhunen, and Erkki Oja.Independent Component Analysis. Wiley, New York, 2001

  34. [40]

    RR interval time series from healthy subjects, 2021

    Isabel Mar ´ ıa Irurzun, Leopoldo Garavaglia, Maria Magdalena Defeo, et al. RR interval time series from healthy subjects, 2021. Version 1.0.0

  35. [41]

    Jolliffe.Principal Component Analysis

    Ian T. Jolliffe.Principal Component Analysis. Springer, New York, 2 edition, 2002

  36. [42]

    Jun, Nicholas A

    James J. Jun, Nicholas A. Steinmetz, Joshua H. Siegle, Daniel J. Denman, Marius Bauza, Brian Barbarits, Albert K. Lee, Costas A. Anastassiou, Alexandru Andrei, C ¸ a˘ gatay Ay- din, Mladen Barbic, Timothy J. Blanche, Vincent Bonin, Jo˜ ao Couto, Barundeb Dutta, Sergey L. Grati...

  37. [43]

    Kolda and Brett W

    Tamara G. Kolda and Brett W. Bader. Tensor decompositions and applications.SIAM Review, 51(3):455–500, 2009

  38. [44]

    Concentration inequalities and moment bounds for sample covariance operators.Bernoulli, pages 110–133, 2017

    Vladimir Koltchinskii and Karim Lounici. Concentration inequalities and moment bounds for sample covariance operators.Bernoulli, pages 110–133, 2017. 41

  39. [45]

    Two decades of array signal processing research: The parametric approach.IEEE Signal Processing Magazine, 13(4):67–94, 1996

    Hamid Krim and Mats Viberg. Two decades of array signal processing research: The parametric approach.IEEE Signal Processing Magazine, 13(4):67–94, 1996

  40. [46]

    O. V. Lepskii. Asymptotically minimax adaptive estimation. i. upper bounds. optimally adaptive estimates.Theory of Probability and its Applications, 36(4):682–697, 1991

  41. [47]

    Deconvolution and estimation of transfer function phase and coefficients for nongaussian linear processes.The Annals of Statistics, 10(4):1195– 1208, 1982

    Keh-Shin Lii and Murray Rosenblatt. Deconvolution and estimation of transfer function phase and coefficients for nongaussian linear processes.The Annals of Statistics, 10(4):1195– 1208, 1982

  42. [48]

    Tensor-on-tensor regression: Riemannian optimization, over-parameterization, statistical-computational gap and their interplay.The Annals of Statistics, 52(6):2583–2612, 2024

    Yuetian Luo and Anru R Zhang. Tensor-on-tensor regression: Riemannian optimization, over-parameterization, statistical-computational gap and their interplay.The Annals of Statistics, 52(6):2583–2612, 2024

  43. [49]

    Toshifumi Matsuoka and T. J. Ulrych. Phase estimation using the bispectrum.Proceedings of the IEEE, 72(10):1403–1411, 1984

  44. [50]

    Chapman and Hall, London, 1987

    Peter McCullagh.Tensor Methods in Statistics. Chapman and Hall, London, 1987

  45. [51]

    Jerry M. Mendel. Tutorial on higher-order statistics (spectra) in signal processing and sys- tem theory: Theoretical results and some applications.Proceedings of the IEEE, 79(3):278– 305, 1991

  46. [52]

    Muirhead.Aspects of Multivariate Statistical Theory

    Robb J. Muirhead.Aspects of Multivariate Statistical Theory. Wiley, New York, 1982

  47. [53]

    Nikias and Jerry M

    Chrysostomos L. Nikias and Jerry M. Mendel. Signal processing with higher-order spectra. IEEE Signal Processing Magazine, 10(3):10–37, 1993

  48. [54]

    Nikias and Athina P

    Chrysostomos L. Nikias and Athina P. Petropulu.Higher-Order Spectra Analysis: A Non- linear Signal Processing Framework. Prentice-Hall, Englewood Cliffs, NJ, 1993

  49. [55]

    Nikias and M

    Chrysostomos L. Nikias and M. R. Raghuveer. Bispectrum estimation: A digital signal processing framework.Proceedings of the IEEE, 75(7):869–891, 1987

  50. [56]

    Wright.Numerical Optimization

    Jorge Nocedal and Stephen J. Wright.Numerical Optimization. Springer, New York, 2 edition, 2006

  51. [57]

    Independent component analysis: A statistical per- spective.Wiley Interdisciplinary Reviews: Computational Statistics, 10(5):e1440, 2018

    Klaus Nordhausen and Hannu Oja. Independent component analysis: A statistical per- spective.Wiley Interdisciplinary Reviews: Computational Statistics, 10(5):e1440, 2018

  52. [58]

    Rong Pan and Chrysostomos L. Nikias. The complex cepstrum of higher order cumulants and nonminimum phase system identification.IEEE Transactions on Acoustics, Speech, and Signal Processing, 36(2):186–205, 1988. 42

  53. [59]

    Parker, Scott H

    Paul A. Parker, Scott H. Holan, and Nalini Ravishanker. Nonlinear time series classification using bispectrum-based deep convolutional neural networks.Applied Stochastic Models in Business and Industry, 36(5):877–890, 2020

  54. [60]

    Structured co- variance estimation via tensor-train decomposition.arXiv preprint arXiv:2510.08174, 2025

    Artsiom Patarasau, Nikita Puchkin, Maxim Rakhuba, and Fedor Noskov. Structured co- variance estimation via tensor-train decomposition.arXiv preprint arXiv:2510.08174, 2025

  55. [61]

    Contributions to the mathematical theory of evolution.Philosophical Trans- actions of the Royal Society of London

    Karl Pearson. Contributions to the mathematical theory of evolution.Philosophical Trans- actions of the Royal Society of London. A, 185:71–110, 1894

  56. [62]

    Petropulu

    Athina P. Petropulu. Noncausal nonminimum phase arma modeling of non-gaussian pro- cesses.IEEE Transactions on Signal Processing, 43(8):1946–1954, 1995

  57. [63]

    Petropulu and Chrysostomos L

    Athina P. Petropulu and Chrysostomos L. Nikias. The complex cepstrum and bicepstrum: Analytic performance evaluation in the presence of gaussian noise.IEEE Transactions on Acoustics, Speech, and Signal Processing, 38(7):1246–1256, 1990

  58. [64]

    Dimension-free structured covariance estimation

    Nikita Puchkin and Maxim Rakhuba. Dimension-free structured covariance estimation. In Proceedings of Thirty Seventh Conference on Learning Theory, volume 247 ofProceedings of Machine Learning Research, pages 4276–4306. PMLR, 2024

  59. [65]

    Provable online CP/PARAF AC decom- position of a structured tensor via dictionary learning

    Sirisha Rambhatla, Xingguo Li, and Jarvis Haupt. Provable online CP/PARAF AC decom- position of a structured tensor via dictionary learning. InAdvances in Neural Information Processing Systems, volume 33, pages 8610–8621, 2020

  60. [66]

    ESPRIT: Estimation of signal parameters via rotational invariance techniques.IEEE Transactions on Acoustics, Speech, and Signal Processing, 37(7):984–995, 1989

    Richard Roy and Thomas Kailath. ESPRIT: Estimation of signal parameters via rotational invariance techniques.IEEE Transactions on Acoustics, Speech, and Signal Processing, 37(7):984–995, 1989

  61. [67]

    Krieger Pub- lishing, Malabar, FL, 1989

    Martin Schetzen.The Volterra and Wiener Theories of Nonlinear Systems. Krieger Pub- lishing, Malabar, FL, 1989

  62. [68]

    Ralph O. Schmidt. Multiple emitter location and signal parameter estimation.IEEE Transactions on Antennas and Propagation, 34(3):276–280, 1986

  63. [69]

    Springer, 2017

    Konrad Schm¨ udgen et al.The moment problem, volume 9. Springer, 2017

  64. [70]

    Shamsunder and Georgios B

    S. Shamsunder and Georgios B. Giannakis. Detection and parameter estimation of multiple nongaussian sources via higher order statistics.IEEE Transactions on Signal Processing, 42(5):1145–1155, 1994

  65. [71]

    Correct es- timation of higher-order spectra: From theoretical challenges to practical multi-channel implementation in signalsnap.arXiv preprint arXiv:2505.01231, 2025

    Markus Sifft, Armin Ghorbanietemad, Fabian Wagner, and Daniel H¨ agele. Correct es- timation of higher-order spectra: From theoretical challenges to practical multi-channel implementation in signalsnap.arXiv preprint arXiv:2505.01231, 2025. 43

  66. [72]

    Cumulants and partition lattices 1.Australian Journal of Statistics, 25(2):378–388, 1983

    Terry P Speed. Cumulants and partition lattices 1.Australian Journal of Statistics, 25(2):378–388, 1983

  67. [73]

    Giannakis, and Jerry M

    Ananthram Swami, Georgios B. Giannakis, and Jerry M. Mendel. Linear modeling of multidimensional non-Gaussian processes using cumulants.Multidimensional Systems and Signal Processing, 1:11–37, 1990

  68. [74]

    Revisit cp tensor decom- position: Statistical optimality and fast convergence.arXiv preprint arXiv:2505.23046, 2025

    Runshi Tang, Julien Chhor, Olga Klopp, and Anru R Zhang. Revisit cp tensor decom- position: Statistical optimality and fast convergence.arXiv preprint arXiv:2505.23046, 2025

  69. [76]

    Runshi Tang, Ming Yuan, and Anru R. Zhang. Mode-wise principal subspace pursuit and matrix spiked covariance model.Journal of the Royal Statistical Society Series B: Statistical Methodology, 87(1):232–255, 2025

  70. [77]

    Springer, New York, 1991

    Masanobu Taniguchi.Higher Order Asymptotic Theory for Time Series Analysis, volume 68 ofLecture Notes in Statistics. Springer, New York, 1991

  71. [78]

    Theodoros Tsiligkaridis and Alfred O. Hero. Covariance estimation in high dimensions via Kronecker product expansions.IEEE Transactions on Signal Processing, 61(21):5347–5360, 2013

  72. [79]

    Jitendra K. Tugnait. Time delay estimation in unknown spatially correlated gaussian noise using higher-order statistics. InProceedings of the 23rd Asilomar Conference on Signals, Systems and Computers, volume 1, pages 211–215, 1989

  73. [80]

    Air Quality

    Saverio Vito. Air Quality. UCI Machine Learning Repository, 2008. DOI: https://doi.org/10.24432/C59K5F

  74. [81]

    Third-order statistics reconstruction from compressive mea- surements.IEEE Transactions on Signal Processing, 69:2888–2901, 2021

    Yanbo Wang and Zhi Tian. Third-order statistics reconstruction from compressive mea- surements.IEEE Transactions on Signal Processing, 69:2888–2901, 2021

  75. [82]

    Yu Wang, Zeyu Sun, Dogyoon Song, and Alfred O. Hero. Kronecker-structured covariance models for multiway data.Statistics Surveys, 16:238–270, 2022

  76. [83]

    On estimation of covariance matrices with Kronecker product structure.IEEE Transactions on Signal Processing, 56(2):478– 491, 2008

    Karl Werner, Magnus Jansson, and Petre Stoica. On estimation of covariance matrices with Kronecker product structure.IEEE Transactions on Signal Processing, 56(2):478– 491, 2008. 44

  77. [84]

    Banding sample autocovariance matrices of sta- tionary processes.Statistica Sinica, pages 1755–1768, 2009

    Wei Biao Wu and Mohsen Pourahmadi. Banding sample autocovariance matrices of sta- tionary processes.Statistica Sinica, pages 1755–1768, 2009

  78. [85]

    Tensor svd: Statistical and computational limits.IEEE Transactions on Information Theory, 64(11):7311–7338, 2018

    Anru Zhang and Dong Xia. Tensor svd: Statistical and computational limits.IEEE Transactions on Information Theory, 64(11):7311–7338, 2018

  79. [87]

    Spike localization algorithms, 2026

    Hao Zhao. Spike localization algorithms, 2026. Files used:spe1.zipandseed 42.zip; accessed 2026-07-03

  80. [88]

    Benchmarking spike source localization algorithms in high density probes.PLOS Computational Biology, 22(3):e1014059, 2026

    Hao Zhao, Xinyu Zhang, Alba Marin-Llobet, Xiaoyun Lin, and Jun Liu. Benchmarking spike source localization algorithms in high density probes.PLOS Computational Biology, 22(3):e1014059, 2026

  81. [89]

    Gemini: Graph estimation with matrix variate normal instances.The Annals of Statistics, 42(2):532–562, 2014

    Shuheng Zhou. Gemini: Graph estimation with matrix variate normal instances.The Annals of Statistics, 42(2):532–562, 2014. 45 A Upper Bounds for Bandable Moment Tensors Definition 2(Bandable moment tensor).Forα >0andβ >0, we say that an order-dmoment tensorM d is(α, β)-bandabl...

Pith tools

Reviewed August 14, 2026 · model on record in the stance chip above.