REVIEW 2 major objections 5 minor 1 cited by
Strong Gaussian approximations with random multipliers
T0 review · 2 major / 5 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read The paper proves that data-dependent random multipliers can be incorporated into strong Gaussian approximations for non-stationary, possibly high-dimensional time series, with explicit error rates and a functional central limit theorem as…
desk verdict Useful random-multiplier extension with a real proof gap in the de-randomization step of Theorem 3; the central result is likely correct but needs a genuine rework. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The proof rests on a block-martingale decomposition of the multiplier error. Write $\hat g_{t,n}=g_{t,n}+e_t$; divide the indices into blocks of length $L=L_n$, and in block $j$ replace $X_{t,n}$ by a copy $\tilde X_{t,n}$ built from fresh innovations. Assumption (A.M) makes $e_t$ measurable with respect to innovations at or before the block start, while $\tilde X_{t,n}$ uses later innovations, so $e_t$ and $\tilde X_{t,n}$ are independent inside the block; the block sums are then martingale differences whose $L^2$ norm is bounded by $\Theta_n\Lambda_n$. A second ingredient is the improved covariance estimate of Proposition 1, $\|\operatorname{Cov}(X_{s,n},X_{t,n})\|_{\mathrm{tr}}\le C_\beta(|s-t|+1)^{-\beta}$, which lets the Gaussian approximation run with only $\beta>1$ instead of $\beta>2$. The Gaussian vectors are then attached to the averaged process $g_{t,n}X_{t,n}$ via the strong-approximation theorem.
What would settle it
For iid standard normal $X_t$, let $\hat g_t=1+n^{-1/2}X_{t-1}$; the theorem predicts the normalized partial sum converges to a standard Brownian motion, and the correction term is $n^{-1}\sum X_{t-1}X_t\to0$. Replacing the multiplier by $\hat g_t=1+n^{-1/2}X_t$ (zero lag) makes the normalized partial sum process converge to $W_u+u$ rather than $W_u$, because $\frac1{\sqrt n}\sum_{t=1}^{\lfloor un\rfloor}\hat g_tX_t = \frac1{\sqrt n}\sum_{t=1}^{\lfloor un\rfloor}X_t + \frac1n\sum_{t=1}^{\lfloor un\rfloor}X_t^2 \Rightarrow W_u+u$; this concrete difference shows the positive lag in Assumption (A.M) is what prevents a bias term.
Extended reading notes
Core claim
On a possibly enlarged probability space, for a non-stationary nonlinear time series $X_{t,n}=G_{t,n}(\epsilon_t)$ with zero mean and physical dependence coefficients decaying as $\Theta_n(h+1)^{-\beta}$ with $\beta>1$, and for random multiplier matrices $\hat g_{t,n}\in\mathbb{R}^{m\times d}$ that are measurable with respect to the innovations up to time $t-L_n$, one can construct independent Gaussian vectors $Y_{t,n}\sim N(0,\Sigma_{t,n})$ such that the maximum over $k\le n$ of $\|\frac{1}{\sqrt n}\sum_{t=1}^k(\hat g_{t,n}X_{t,n}-Y_{t,n})\|$ is stochastically bounded by the displayed rate. The rate contains a de-randomization error controlled by the $L^2$ distance $\Lambda_n$ between $\hat g_{t,n}$ and a deterministic approximation $g_{t,n}$, a dependence-tail term $\Xi_n=\sum_{h\ge L_n}\delta_n(h)$, and the Gaussian-approximation rate for $g_{t,n}X_{t,n}$. When the dimension is fixed and the sequence is locally stationary, Theorem 5 turns this into the functional convergence $\frac{1}{\sqrt n}\sum_{t=1}^{\lfloor un\rfloor}\hat g_{t,n}(X_{t,n}-E X_{t,n})\Rightarrow\int_0^u \Sigma_v^{1/2}\,dW_v$.
Load-bearing premise
The results stand or fall on Assumption (A.M): each multiplier at time $t$ may use only information up to time $t-L_n$ for some positive lag $L_n\ge1$, so that within each block the multiplier error is independent of a freshly randomized copy of the data; if $L_n=0$ the block argument collapses and the stated error bounds no longer follow.
Editorial extensions
If this is right
- In a finite-dimensional locally stationary model, the studentized partial sum process $\frac1{\sqrt n}\sum_{t=1}^{\lfloor un\rfloor}\tilde\Sigma(t/n)^{-1/2}(X_{t,n}-E X_{t,n})$ converges to standard Brownian motion, so the critical values of the resulting sequential test do not depend on the unknown covariance function $\Sigma(u)$.
- The error bound quantifies the price of data-dependent weights: the approximation degrades with the estimation error $\Lambda_n$, the dependence tail $\Xi_n$ beyond the lag, and the block length $L_n$.
- Because $d$ and $m$ may grow with $n$, the result covers high-dimensional partial sums with high-dimensional multipliers, where no classical limit object exists.
- The underlying Gaussian approximation is improved to any $\beta>1$, extending the earlier $\beta>2$ regime.
Reading between the lines
- The paper leaves implicit that the same block-martingale argument should work for online estimators updated with a rolling window of length $L_n$ growing slowly, and the bound's $L_n\Phi_n n^{1/q-1/2}$ term gives the exact trade-off between adaptivity and accuracy.
- Because the critical value in the studentized test is distribution-free, a practical extension is to turn the FCLT into a stopping rule for online monitoring: reject at the first $u$ where the studentized path crosses the Brownian supremum quantile, without knowing $\Sigma$ in advance.
- For independent data the paper notes $\Xi_n=0$ for any $L_n\ge1$; this suggests the lag requirement is mainly a tool for dependence, and a version with contemporaneous multipliers may be possible for independent or conditionally independent arrays via different martingale arguments.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper extends strong Gaussian approximation results for partial sums of non-stationary, potentially high-dimensional time series to sums whose summands are multiplied by data-dependent random matrices. The main result (Theorem 3) gives an explicit coupling between the multiplier-weighted partial sum process and a sequence of independent Gaussian vectors, with an error rate expressed in terms of the multiplier estimation error (Λ_n), the dependence tail (Ξ_n), the multiplier regularity (Φ_n, Ψ_n), and a block-length term. In the finite-dimensional locally stationary setting (Theorem 5), the coupling yields a functional central limit theorem. An application to studentized partial sums and sequential testing is given in Theorem 6, with a small simulation study. The paper also improves the underlying strong approximation (Theorem 2) by weakening the dependence condition from β > 2 to β > 1. The central claims are plausible and the rates are clearly stated, but the proof of the de-randomization step in Theorem 3 currently contains a load-bearing gap, in addition to a false equality in a maximal inequality display.
Significance. If the results are correct, they constitute a useful extension of strong Gaussian approximations to settings with estimated or data-dependent multipliers, which is relevant for bootstrap and sequential inference in non-stationary environments. The explicit treatment of the lagged measurability assumption (A.M) and the resulting Ξ_n terms is a genuine contribution, and the improved β > 1 condition in Theorem 2 over the prior work of Mies and Steland is valuable. The paper ships a concrete statistical application (studentized partial sums) with simulation evidence, and the main theorem is stated with explicit rates. These are real strengths. However, the proof of Theorem 3 is not complete as written: one equality is invalid, and the residual bound for the de-randomization step relies on a coupling that does not produce the stated rate. Both issues appear repairable without changing the stated rates, but they are central to the main theorem.
major comments (2)
- [Section 3, Proof of Theorem 3, display after Eq. (4)] The identity E max_{r} ||∑_{j=0}^{r} W_j||^2 = ∑_{j} E||W_j||^2 is not valid for the block martingale differences W_j; only an inequality, e.g. Doob's L2 maximal inequality, gives E max_r ||S_r||^2 ≤ C ∑_j E||W_j||^2. Since the subsequent computation bounds exactly the sum of block second moments, the final rate is unchanged up to a universal constant, but the displayed equality must be replaced by an inequality.
- [Section 3, Proof of Theorem 3, second term in Eq. (4)] The coupling \tilde X_{t,n} = G_{t,n}(\tilde ε^j_{t,t-jL}) replaces innovations at indices ≤ jL, so for t = jL + d with 1 ≤ d ≤ L, the residual X_{t,n} - \tilde X_{t,n} depends on innovations at lags at least d. By (A.2), ||X_{t,n} - \tilde X_{t,n}||_{L2} is of order Θ_n d^{1-β}. Consequently, ∑_t ||X_{t,n} - \tilde X_{t,n}||_{L2}^2 is of order (n/L) ∑_{d=1}^L d^{2-2β}, which for β > 3/2 is of order n/L, whereas (√n Ξ_n)^2 is of order n L^{2-2β} and is strictly smaller when β > 3/2. Hence the displayed bound OP(Λ_n) OP(√n Ξ_n) for the second term in (4) does not follow from the construction used. This gap is load-bearing because the rate Λ_n Ξ_n enters Theorem 3 and is one of the conditions for Theorem 5. The proof should be repaired, for example by resampling at a fixed lag L in a way that still preserves the block martingale-difference structure, so that the residual tail starts at L.
minor comments (5)
- [Section 4, Theorem 5] The statement should explicitly require L_n ≥ 1 (or L_n ∈ N) in addition to L_n = o(n^{1/2 - 1/q}), so that Theorem 3 applies; the current condition alone does not guarantee the lagged measurability needed for the de-randomization argument.
- [Section 5, Proof of Theorem 6] The sentence 'with L_n = 1 and Θ_n = 0' should read 'with L_n = 1 and Ξ_n = 0': the dependence constant Θ_n is positive even for independent data, while the tail sum Ξ_n vanishes for any L_n ≥ 1.
- [Section 5, Proof of Theorem 6] The statement 'for k_n ≍ n^{-1/3}' is a typo; it should be 'k_n ≍ n^{2/3}' (or a positive power consistent with k_n = \lfloor n^{2/3} \rfloor) to match the subsequent Λ_n = O(n^{1/6}).
- [Section 3, after Theorem 3] The phrase 'Assumption (A.4) implies' should refer to Assumption (A.M), not (A.4), which is defined later in Section 4.
- [Section 2, notation for ε_{t,h} and \tilde ε_{t,h}] The notation for the coupled innovation sequences is confusing: \tilde ε_{t,h} and ε_{t,h} are defined with different replacement schemes, and the proof of Theorem 3 uses \tilde ε^j_{t,t-jL} but the explanatory formula in the text appears to describe ε_{t,h}. The authors should clarify which sequence is used in the block coupling.
Circularity Check
No significant circularity: the random-multiplier approximation is derived from an externally published base theorem and does not assume its own conclusion.
full rationale
The central derivation chain is not circular. Theorem 3 reduces the target random-multiplier statement to (i) a de-randomization step that bounds the effect of replacing the random multiplier \hat g_{t,n} by the non-random g_{t,n} and (ii) the deterministic-multiplier strong approximation Theorem 2, applied to Z_{t,n}=g_{t,n}G_{t,n}(\epsilon_t). Theorem 2 is not assumed as the target result; it is an improved version of a previously published theorem (Mies and Steland 2023), and the improvement is justified in the paper via Proposition 1, which is proved from assumption (A.2). Thus the random-multiplier result has independent content beyond the prior theorem. No parameters are fitted and renamed as predictions; the rates are derived from the stated constants Λ_n, Ξ_n, Φ_n, Ψ_n, and the physical-dependence decay. The paper also explicitly notes that the FCLT without multipliers is standard in the literature, so Theorem 5 is not a renaming of a known result. The proof of Theorem 3 contains a potentially serious gap: the bound on the second term in decomposition (4) is asserted as OP(Λ_n)OP(√n Ξ_n), whereas the blockwise coupling \tilde X_t=G_{t,n}(\tilde\epsilon^j_{t,t-jL}) gives an L2 sum of order n/L, which for β>3/2 is larger than n Ξ_n^2; this is a correctness concern, not a circularity, and there is no reduction of the theorem to its own assumptions.
Assumptions & free parameters
assumptions (7)
- standard math Physical dependence representation: X_{t,n}=G_{t,n}(epsilon_t) with iid seeds and martingale-difference projection (Wu, 2005).
- domain assumption (A.1) Time smoothness: ||G_{1,n}(epsilon_0)||_{L2} + sum_{t=2}^n ||G_{t,n}(epsilon_0)-G_{t-1,n}(epsilon_0)||_{L2} <= Theta_n Gamma_n.
- domain assumption (A.2) Polynomial decay of physical dependence: delta_n(h) <= Theta_n (h+1)^{-beta} with beta > 1.
- domain assumption (A.M) Lagged measurability: ghat_t is epsilon_{t-L_n}-measurable.
- domain assumption (A.3) and (A.4) Local stationarity: maps G_u and g_u approximate G_{floor(vn),n} and g_{floor(vn),n} in L2 as n tends to infinity.
- domain assumption The base strong approximation theorem of Mies and Steland (2023, Thm 3.1) and its proof skeleton.
- standard math L2 martingale maximal inequality for vector-valued martingale difference sequences.
Cite this review
Pith. "Pith review of Strong Gaussian approximations with random multipliers." pith.science (2026). https://pith.science/paper/2WRKTBQY
@misc{pith2026241214346,
author = {Pith},
title = {Pith review of: Strong Gaussian approximations with random multipliers},
year = {2026},
howpublished = {\url{https://pith.science/paper/2WRKTBQY}},
note = {Machine review of arXiv:2412.14346}
}
read the original abstract
One reason why standard formulations of the central limit theorems are not applicable in high-dimensional and non-stationary regimes is the lack of a suitable limit object. Instead, suitable distributional approximations can be used, where the approximating object is not constant, but a sequence as well. We extend Gaussian approximation results for the partial sum process by allowing each summand to be multiplied by a data-dependent matrix. The results allow for serial dependence of the data, and for high-dimensionality of both the data and the multipliers. In the finite-dimensional and locally-stationary setting, we obtain a functional central limit theorem as a direct consequence. An application to sequential testing in non-stationary environments is described.
Forward citations
Cited by 1 Pith paper
-
At the edge of Donsker's Theorem: Asymptotics of multiscale scan statistics
Replacing the additive penalty in multiscale scan tests by a multiplicative weighting yields critical values that are asymptotically valid for sub-Gaussian noise, based on a new thresholded weak convergence result.
Reference graph
Works this paper leans on
-
[1]
Berkes, I., Liu, W., and Wu, W. B. (2014). Koml´ os-Major-Tusn´ ady approximation under dependence. Annals of Probability, 42(2):794–817
work page 2014
-
[2]
Billingsley, P. (1999). Convergence of Probability Measures. John Wiley & Sons
work page 1999
-
[3]
Bonnerjee, S., Karmakar, S., and Wu, W. B. (2024). Gaussian Appr oximation For Non-stationary Time Series with Optimal Rate and Explicit Construct ion. Annals of Statistics, 52(5):2293–2317
work page 2024
-
[4]
Dahlhaus, R. (1997). Fitting time series models to nonstationary pr ocesses. Annals of Statistics, 25(1):1–37
work page 1997
-
[5]
Dahlhaus, R. and Polonik, W. (2009). Empirical spectral processe s for locally stationary time series. Bernoulli, 15(1):1–39
work page 2009
-
[6]
Dahlhaus, R., Richter, S., and Wu, W. B. (2019). Towards a general theory for nonlinear locally stationary processes. Bernoulli, 25(2):1013–1044
work page 2019
-
[7]
Giraud, C., Roueff, F., and Sanchez-Perez, A. (2015). Aggregatio n of predictors for nonstationary sub-linear processes and online adaptive forecast ing of time varying autoregressive processes. The Annals of Statistics, 43(6):2412–2450. 13
work page 2015
-
[8]
Grenier, Y. (1983). Time-dependent ARMA modeling of nonstationa ry signals. IEEE Transactions on Acoustics, Speech, and Signal Processing, 31(4):899–911
work page 1983
Show all 17 references
-
[9]
and Wu, W
Karmakar, S. and Wu, W. B. (2020). Optimal Gaussian Approximatio n For Multiple Time Series. Statistica Sinica, 30:1399–1417. Koml´ os, J., Major, P., and Tusn´ ady, G. (1975). An approximation of partial sums of independent R V’-s, and the sample DF. I. Zeitschrift f¨ urWahrs...
2020
-
[10]
and Lin, Z
Liu, W. and Lin, Z. (2009). Strong approximation for a class of stat ionary processes. Stochastic Processes and their Applications, 119:249–280
2009
-
[11]
Mies, F. (2023). Functional estimation and change detection for n onstationary time series. Journal of the American Statistical Association, 118(542):1011–1022
2023
-
[12]
and Steland, A
Mies, F. and Steland, A. (2023). Sequential Gaussian approximatio n for nonstationary time series in high dimensions. Bernoulli, 29(4):3114–3140
2023
-
[13]
and Steland, A
Mies, F. and Steland, A. (2024). Projection inference for high-dim ensional covari- ance matrices with structured shrinkage targets. Electronic Journal of Statistics, 18(1):1643–1676
2024
-
[14]
Moulines, E., Priouret, P., and Roueff, F. (2005). On recursive estim ation for time varying autoregressive processes. The Annals of Statistics, 33(6). Subba Rao, T. (1970). The Fitting of Non-stationary Time-series M odels with Time- dependent Parameters. Journal of the Royal ...
2005
-
[15]
Wu, W. B. (2005). Nonlinear system theory: Another look at depen dence. Proceedings of the National Academy of Sciences, 102(40):14150–14154
2005
-
[16]
Wu, W. B. and Zhou, Z. (2011). Gaussian approximations for non-s tationary multiple time series. Statistica Sinica, 21(3):1397–1413
2011
-
[17]
Zaitsev, A. Y. (2013). The accuracy of strong Gaussian approxim ation for sums of independent random vectors. Russian Mathematical Surveys, 68(4):721–761. 14
2013
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.