{"id":"3e5d803d-cae2-48ce-a14c-adcd65fe4c23","arxiv_id":"2412.14346","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Partial sums with data-dependent multiplier matrices admit a strong Gaussian approximation whose error is controlled by a lagged measurability condition, with a functional central limit theorem in the locally stationary case.","lead":"This paper proves Gaussian approximation bounds for partial sums in which every observation is multiplied by a random, data-dependent matrix. The results cover high-dimensional, serially dependent and non-stationary data, and imply a functional central limit theorem for locally stationary sequences with a sequential testing rule.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The de-randomization step in Theorem 3 couples with the wrong lag: the residual bound OP(sqrt(n) Xi_n) fails for beta > 3/2, so the stated error rate is not established as written.","rationale":"The reader's weakest-assumption focus on Assumption (A.M) is a legitimate limitation, but it is an explicit hypothesis rather than an internal inconsistency: the theorem only claims the result under the stated lag, and the application uses L_n = 1. The more load-bearing issue is an internal proof gap in the de-randomization argument. The blockwise coupling \\tilde X_t = G_{t,n}(\\tilde\\epsilon^j_{t,t-jL}) leaves residual tails that start at distance d = t - jL, which can be as small as 1, not at the lag L. The crude Cauchy-Schwarz bound in the proof then requires \\sum_t \\|X_t - \\tilde X_t\\|^2 to be O(n \\Xi_n^2), but for beta > 3/2 the actual size is O(n/L), which is larger by a factor of order L^{beta - 3/2}. This affects the central claim directly because Theorem 3's de-randomization error feeds into the FCLT and into the conditions of Theorem 5. The flaw is repairable by replacing the block-boundary coupling with a per-time lag-L coupling, and the theorem is probably correct after that repair; thus I would keep the reader's CONDITIONAL verdict rather than reject. I also agree with the reader's minor points about the false Doob-type equality and the notation/typo issues, but those are less central than the residual-tail problem identified here.","tokens_in":11744,"tokens_out":24945,"duration_ms":223909,"concrete_test":"Recompute the residual sum in (4) with the blockwise \\tilde X_t = G_{t,n}(\\tilde\\epsilon^j_{t,t-jL}) for the MA(1) process X_t = \\epsilon_t + \\epsilon_{t-1} with L = L_n \\to \\infty and beta > 3/2: evaluate \\sum_{t=1}^n \\|X_t - \\tilde X_t\\|^2 and verify that it is of order n/L, not n L^{2-2\\beta}. If confirmed, the proof's OP(\\sqrt n \\Xi_n) bound is invalid; then repeat the de-randomization with the lag-L coupling \\tilde X_t = G_{t,n}(\\epsilon_{t,t-L}) and check that the residual term is indeed \\Lambda_n \\Xi_n, which would confirm that the theorem statement is likely correct but the proof must be revised.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"In the proof of Theorem 3, after equation (4), the blockwise coupling is \\tilde X_t = G_{t,n}(\\tilde\\epsilon^j_{t,t-jL}) for t = jL+1, ..., (j+1)L. Thus \\tilde X_t resamples the innovation at index jL, whose lag from t is d = t - jL in {1, ..., L}. By (A.2), \\|X_t - \\tilde X_t\\|_{L2} is bounded by \\Theta_n \\sum_{h=d}^\\infty (h+1)^{-\\beta}, which is of order \\Theta_n d^{1-\\beta}. Squaring and summing over blocks gives \\sum_t \\|X_t - \\tilde X_t\\|^2 of order (n/L) \\sum_{d=1}^L d^{2-2\\beta}. For beta > 3/2 this is of order n/L, whereas (\\sqrt n \\Xi_n)^2 is of order n L^{2-2\\beta}, which is strictly smaller because 2-2\\beta < -1. Hence the displayed bound OP(\\Lambda_n) * OP(\\sqrt n \\Xi_n) for the second term in (4) does not follow from the construction used in the paper. This gap is load-bearing because the de-randomization rate \\Lambda_n \\Xi_n appears in Theorem 3 and is one of the hypotheses needed for Theorem 5. The natural fix is to couple with G_{t,n}(\\epsilon_{t,t-L}) for each t, resampling at lag L rather than at the block boundary jL; then the residual tail starts at L, the claimed \\sqrt n \\Xi_n bound is restored, and the block-length term n^{1/q-1/2} L \\Phi_n becomes unnecessary. This is a repairable proof gap, but Theorem 3 is not fully proved as written.","agreement_with_reader":"disagree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper extends strong Gaussian approximation results for partial sums of non-stationary, potentially high-dimensional time series to sums whose summands are multiplied by data-dependent random matrices. The main result (Theorem 3) gives an explicit coupling between the multiplier-weighted partial sum process and a sequence of independent Gaussian vectors, with an error rate expressed in terms of the multiplier estimation error (Λ_n), the dependence tail (Ξ_n), the multiplier regularity (Φ_n, Ψ_n), and a block-length term. In the finite-dimensional locally stationary setting (Theorem 5), the coupling yields a functional central limit theorem. An application to studentized partial sums and sequential testing is given in Theorem 6, with a small simulation study. The paper also improves the underlying strong approximation (Theorem 2) by weakening the dependence condition from β > 2 to β > 1. The central claims are plausible and the rates are clearly stated, but the proof of the de-randomization step in Theorem 3 currently contains a load-bearing gap, in addition to a false equality in a maximal inequality display.","tokens_in":12083,"tokens_out":11653,"duration_ms":95015,"significance":"If the results are correct, they constitute a useful extension of strong Gaussian approximations to settings with estimated or data-dependent multipliers, which is relevant for bootstrap and sequential inference in non-stationary environments. The explicit treatment of the lagged measurability assumption (A.M) and the resulting Ξ_n terms is a genuine contribution, and the improved β > 1 condition in Theorem 2 over the prior work of Mies and Steland is valuable. The paper ships a concrete statistical application (studentized partial sums) with simulation evidence, and the main theorem is stated with explicit rates. These are real strengths. However, the proof of Theorem 3 is not complete as written: one equality is invalid, and the residual bound for the de-randomization step relies on a coupling that does not produce the stated rate. Both issues appear repairable without changing the stated rates, but they are central to the main theorem.","major_comments":[{"comment":"The identity E max_{r} ||∑_{j=0}^{r} W_j||^2 = ∑_{j} E||W_j||^2 is not valid for the block martingale differences W_j; only an inequality, e.g. Doob's L2 maximal inequality, gives E max_r ||S_r||^2 ≤ C ∑_j E||W_j||^2. Since the subsequent computation bounds exactly the sum of block second moments, the final rate is unchanged up to a universal constant, but the displayed equality must be replaced by an inequality.","section":"Section 3, Proof of Theorem 3, display after Eq. (4)"},{"comment":"The coupling \\tilde X_{t,n} = G_{t,n}(\\tilde ε^j_{t,t-jL}) replaces innovations at indices ≤ jL, so for t = jL + d with 1 ≤ d ≤ L, the residual X_{t,n} - \\tilde X_{t,n} depends on innovations at lags at least d. By (A.2), ||X_{t,n} - \\tilde X_{t,n}||_{L2} is of order Θ_n d^{1-β}. Consequently, ∑_t ||X_{t,n} - \\tilde X_{t,n}||_{L2}^2 is of order (n/L) ∑_{d=1}^L d^{2-2β}, which for β > 3/2 is of order n/L, whereas (√n Ξ_n)^2 is of order n L^{2-2β} and is strictly smaller when β > 3/2. Hence the displayed bound OP(Λ_n) OP(√n Ξ_n) for the second term in (4) does not follow from the construction used. This gap is load-bearing because the rate Λ_n Ξ_n enters Theorem 3 and is one of the conditions for Theorem 5. The proof should be repaired, for example by resampling at a fixed lag L in a way that still preserves the block martingale-difference structure, so that the residual tail starts at L.","section":"Section 3, Proof of Theorem 3, second term in Eq. (4)"}],"minor_comments":[{"comment":"The statement should explicitly require L_n ≥ 1 (or L_n ∈ N) in addition to L_n = o(n^{1/2 - 1/q}), so that Theorem 3 applies; the current condition alone does not guarantee the lagged measurability needed for the de-randomization argument.","section":"Section 4, Theorem 5"},{"comment":"The sentence 'with L_n = 1 and Θ_n = 0' should read 'with L_n = 1 and Ξ_n = 0': the dependence constant Θ_n is positive even for independent data, while the tail sum Ξ_n vanishes for any L_n ≥ 1.","section":"Section 5, Proof of Theorem 6"},{"comment":"The statement 'for k_n ≍ n^{-1/3}' is a typo; it should be 'k_n ≍ n^{2/3}' (or a positive power consistent with k_n = \\lfloor n^{2/3} \\rfloor) to match the subsequent Λ_n = O(n^{1/6}).","section":"Section 5, Proof of Theorem 6"},{"comment":"The phrase 'Assumption (A.4) implies' should refer to Assumption (A.M), not (A.4), which is defined later in Section 4.","section":"Section 3, after Theorem 3"},{"comment":"The notation for the coupled innovation sequences is confusing: \\tilde ε_{t,h} and ε_{t,h} are defined with different replacement schemes, and the proof of Theorem 3 uses \\tilde ε^j_{t,t-jL} but the explanatory formula in the text appears to describe ε_{t,h}. The authors should clarify which sequence is used in the block coupling.","section":"Section 2, notation for ε_{t,h} and \\tilde ε_{t,h}"}],"recommendation":"major_revision","confidential_remarks":"The paper presents a natural extension of strong approximations to data-dependent multipliers, and the stated rates are plausible. The main novelty is the de-randomization argument in Theorem 3, but that argument currently has two technical flaws: an invalid equality for the maximum of martingale difference blocks and a residual bound that does not follow from the chosen coupling. Both appear repairable without changing the rates, so I am not recommending rejection. The revision should make the proof of Theorem 3 fully rigorous, in particular by selecting a coupling that yields the claimed Λ_n Ξ_n term while preserving the martingale-difference structure. The simulation section is brief but adequate for illustration. The paper relies heavily on the author's earlier work, which is acceptable given that the new ingredients are clearly identified."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a genuinely useful extension of the strong approximation framework to random multipliers, and the studentization FCLT is a natural payoff. But the proof of Theorem 3 has a gap in the de-randomization step that is load-bearing, and the fix is not a one-line constant change.\n\nWhat is actually new: Theorem 3 allows data-dependent multipliers with a lag, and Theorem 5 gives an FCLT for locally stationary sequences. Theorem 2 improves the beta range from beta>2 to beta>1 by exploiting Proposition 1's covariance bound. That is a real improvement. The application to sequential testing is coherent, and the simulations, while small, illustrate the power gain.\n\nWhat is done well: The block-martingale idea is sensible, and the structure of the proof is clear. The paper is honest about relying on Mies-Steland (2023) for the base approximation, which is acceptable given that the target is an extension. The rate calculations in Theorem 2 are careful.\n\nThe soft spots: First, the proof of Theorem 3 asserts an equality for the maximal expectation of the martingale block sums, but only an inequality with a constant (Doob) is valid. That is minor and repairable. Second, and more serious, in equation (4) tilde X_t is coupled at the block boundary jL, so the lag from t to the resampled innovation is d = t - jL, ranging from 1 to L. The L2 error for each term is of order Theta_n d^{1-beta}, not Theta_n L^{1-beta}. Summing squares over blocks gives roughly (n/L) sum_{d=1}^L d^{2-2beta} ~ n/L for beta>3/2, which is larger than the claimed n L^{2-2beta} by a factor L^{2beta-3}. So the displayed OP(Lambda_n) OP(sqrt(n) Xi_n) bound for the residual does not follow from the construction as written. That matters because the de-randomization term Lambda_n Xi_n feeds into Theorem 3 and into the hypotheses of Theorem 5. The natural fix—resampling at lag L for each t—would break the martingale difference structure of the first term, so the repair is not trivial. The theorem may well be true with a moderately worse rate, but it needs a real argument.\n\nThere are also typos: a B/L mismatch in the residual term, an incorrect exponent k_n ~ n^{-1/3} in the proof of Theorem 6 (should be n^{2/3}), and a line saying Theta_n = 0 where the context indicates Xi_n = 0.\n\nBottom line: the paper deserves a serious referee, but the referee should require a corrected proof of Theorem 3. The central contribution is real and likely correct; it just isn't fully proved as written.","headline":"Useful random-multiplier extension with a real proof gap in the de-randomization step of Theorem 3; the central result is likely correct but needs a genuine rework.","tokens_in":12643,"tokens_out":11643,"would_cite":true,"duration_ms":91948,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["60F17","62M10","60F05","62L10"],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper proves that data-dependent random multipliers can be incorporated into strong Gaussian approximations for non-stationary, possibly high-dimensional time series, with explicit error rates and a functional central limit theorem as…","keywords":["Gaussian approximation","strong approximation","functional central limit theorem","locally stationary time series","random multipliers","high-dimensional time series","physical dependence","sequential testing"],"falsifier":"For iid standard normal $X_t$, let $\\hat g_t=1+n^{-1/2}X_{t-1}$; the theorem predicts the normalized partial sum converges to a standard Brownian motion, and the correction term is $n^{-1}\\sum X_{t-1}X_t\\to0$. Replacing the multiplier by $\\hat g_t=1+n^{-1/2}X_t$ (zero lag) makes the normalized partial sum process converge to $W_u+u$ rather than $W_u$, because $\\frac1{\\sqrt n}\\sum_{t=1}^{\\lfloor un\\rfloor}\\hat g_tX_t = \\frac1{\\sqrt n}\\sum_{t=1}^{\\lfloor un\\rfloor}X_t + \\frac1n\\sum_{t=1}^{\\lfloor un\\rfloor}X_t^2 \\Rightarrow W_u+u$; this concrete difference shows the positive lag in Assumption (A.M) is what prevents a bias term.","tokens_in":11493,"feed_emoji":"📊","tokens_out":14031,"duration_ms":117368,"temperature":0.7,"pith_summary":"Classical functional central limit theorems fail for non-stationary or high-dimensional data because there is no single limit distribution to converge to. This paper proves that the partial sum process can still be uniformly approximated by a sequence of independent Gaussian vectors whose covariance matrices come from the time-varying local law, and that this approximation survives when every summand is multiplied by a data-dependent random matrix. The main theorem gives an explicit error bound in terms of the multiplier estimation error, the serial dependence tail, a delay parameter, and the growing dimensions. In the finite-dimensional locally stationary case the approximation upgrades to a functional central limit theorem, with an application to studentized sequential tests whose critical values do not depend on the unknown covariance.","feed_headline":"Nonstationary sums with data-driven weights still go Gaussian","feed_subtitle":"A positive lag in the multipliers yields explicit error rates and a studentized functional central limit theorem.","key_machinery":"The proof rests on a block-martingale decomposition of the multiplier error. Write $\\hat g_{t,n}=g_{t,n}+e_t$; divide the indices into blocks of length $L=L_n$, and in block $j$ replace $X_{t,n}$ by a copy $\\tilde X_{t,n}$ built from fresh innovations. Assumption (A.M) makes $e_t$ measurable with respect to innovations at or before the block start, while $\\tilde X_{t,n}$ uses later innovations, so $e_t$ and $\\tilde X_{t,n}$ are independent inside the block; the block sums are then martingale differences whose $L^2$ norm is bounded by $\\Theta_n\\Lambda_n$. A second ingredient is the improved covariance estimate of Proposition 1, $\\|\\operatorname{Cov}(X_{s,n},X_{t,n})\\|_{\\mathrm{tr}}\\le C_\\beta(|s-t|+1)^{-\\beta}$, which lets the Gaussian approximation run with only $\\beta>1$ instead of $\\beta>2$. The Gaussian vectors are then attached to the averaged process $g_{t,n}X_{t,n}$ via the strong-approximation theorem.","core_discovery":"On a possibly enlarged probability space, for a non-stationary nonlinear time series $X_{t,n}=G_{t,n}(\\epsilon_t)$ with zero mean and physical dependence coefficients decaying as $\\Theta_n(h+1)^{-\\beta}$ with $\\beta>1$, and for random multiplier matrices $\\hat g_{t,n}\\in\\mathbb{R}^{m\\times d}$ that are measurable with respect to the innovations up to time $t-L_n$, one can construct independent Gaussian vectors $Y_{t,n}\\sim N(0,\\Sigma_{t,n})$ such that the maximum over $k\\le n$ of $\\|\\frac{1}{\\sqrt n}\\sum_{t=1}^k(\\hat g_{t,n}X_{t,n}-Y_{t,n})\\|$ is stochastically bounded by the displayed rate. The rate contains a de-randomization error controlled by the $L^2$ distance $\\Lambda_n$ between $\\hat g_{t,n}$ and a deterministic approximation $g_{t,n}$, a dependence-tail term $\\Xi_n=\\sum_{h\\ge L_n}\\delta_n(h)$, and the Gaussian-approximation rate for $g_{t,n}X_{t,n}$. When the dimension is fixed and the sequence is locally stationary, Theorem 5 turns this into the functional convergence $\\frac{1}{\\sqrt n}\\sum_{t=1}^{\\lfloor un\\rfloor}\\hat g_{t,n}(X_{t,n}-E X_{t,n})\\Rightarrow\\int_0^u \\Sigma_v^{1/2}\\,dW_v$.","pith_inferences":["The paper leaves implicit that the same block-martingale argument should work for online estimators updated with a rolling window of length $L_n$ growing slowly, and the bound's $L_n\\Phi_n n^{1/q-1/2}$ term gives the exact trade-off between adaptivity and accuracy.","Because the critical value in the studentized test is distribution-free, a practical extension is to turn the FCLT into a stopping rule for online monitoring: reject at the first $u$ where the studentized path crosses the Brownian supremum quantile, without knowing $\\Sigma$ in advance.","For independent data the paper notes $\\Xi_n=0$ for any $L_n\\ge1$; this suggests the lag requirement is mainly a tool for dependence, and a version with contemporaneous multipliers may be possible for independent or conditionally independent arrays via different martingale arguments."],"forward_implications":["In a finite-dimensional locally stationary model, the studentized partial sum process $\\frac1{\\sqrt n}\\sum_{t=1}^{\\lfloor un\\rfloor}\\tilde\\Sigma(t/n)^{-1/2}(X_{t,n}-E X_{t,n})$ converges to standard Brownian motion, so the critical values of the resulting sequential test do not depend on the unknown covariance function $\\Sigma(u)$.","The error bound quantifies the price of data-dependent weights: the approximation degrades with the estimation error $\\Lambda_n$, the dependence tail $\\Xi_n$ beyond the lag, and the block length $L_n$.","Because $d$ and $m$ may grow with $n$, the result covers high-dimensional partial sums with high-dimensional multipliers, where no classical limit object exists.","The underlying Gaussian approximation is improved to any $\\beta>1$, extending the earlier $\\beta>2$ regime."],"supporting_citations":[{"why":"Supplies the sequential Gaussian approximation for $g_{t,n}X_{t,n}$ that Theorem 3 couples with the de-randomization step; its Theorem 3.1 is the result being improved and extended.","marker":"Mies and Steland (2023)"},{"why":"Provides the physical dependence measure $\\delta_n(h)$ used in Assumption (A.2) and throughout the error bounds.","marker":"Wu (2005)"},{"why":"Supplies the locally stationary framework that the finite-dimensional functional central limit theorem uses to define the local covariance $\\Sigma_u$.","marker":"Dahlhaus (1997)"},{"why":"Frames the multiplier central limit theorem context for bootstrap methods that the paper generalizes to dependent multipliers.","marker":"van der Vaart and Wellner (1996)"},{"why":"Introduces the lagged-measurability condition for multipliers that Assumption (A.M) adopts, and provides the change-detection setting.","marker":"Mies (2023)"}],"fun_headline_variants":["Random multipliers preserve Gaussian approximation in non-stationary sums","Data-driven weights don't break Gaussian limits for non-stationary sums","Gaussian approximation with random multipliers: beyond stationary settings","Explicit error rates for Gaussian approximation with random multipliers","Random multipliers and non-stationarity: Gaussian bounds remain"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The results stand or fall on Assumption (A.M): each multiplier at time $t$ may use only information up to time $t-L_n$ for some positive lag $L_n\\ge1$, so that within each block the multiplier error is independent of a freshly randomized copy of the data; if $L_n=0$ the block argument collapses and the stated error bounds no longer follow.","fun_headline_variants_meta":{"raw":{"variants":["Random multipliers preserve Gaussian approximation in non-stationary sums","Data-driven weights don't break Gaussian limits for non-stationary sums","Gaussian approximation with random multipliers: beyond stationary settings","Explicit error rates for Gaussian approximation with random multipliers","Random multipliers and non-stationarity: Gaussian bounds remain"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00156,"raw_usage":{"total_tokens":6240,"prompt_tokens":962,"completion_tokens":5278,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":578,"completion_tokens_details":{"reasoning_tokens":5196}},"tokens_in":578,"tokens_out":5278,"duration_ms":30202,"temperature":1.0,"reasoning_tokens":5196,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T12:19:58.656367+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"For iid standard normal $X_t$, let $\\hat g_t=1+n^{-1/2}X_{t-1}$; the theorem predicts the normalized partial sum converges to a standard Brownian motion, and the correction term is $n^{-1}\\sum X_{t-1}X_t\\to0$. Replacing the multiplier by $\\hat g_t=1+n^{-1/2}X_t$ (zero lag) makes the normalized partial sum process converge to $W_u+u$ rather than $W_u$, because $\\frac1{\\sqrt n}\\sum_{t=1}^{\\lfloor un\\rfloor}\\hat g_tX_t = \\frac1{\\sqrt n}\\sum_{t=1}^{\\lfloor un\\rfloor}X_t + \\frac1n\\sum_{t=1}^{\\lfloor un\\rfloor}X_t^2 \\Rightarrow W_u+u$; this concrete difference shows the positive lag in Assumption (A.M) is what prevents a bias term.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the locally stationary framework that the finite-dimensional functional central limit theorem uses to define the local covariance $\\Sigma_u$."}],"review_version":1}