{"id":"ea404284-3483-4fb8-b094-1a59409ae93b","arxiv_id":"2506.08723","paper_version":1,"verdict":"ACCEPT","confidence":"MODERATE","novelty_score":8.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"For non-stationary high-dimensional time series with short memory and light tails, normalized sums can be approximated by Gaussian vectors in Wasserstein distance and on all convex sets at nearly optimal rates.","lead":"This paper proves new Gaussian approximation bounds for high-dimensional non-stationary time series, in Wasserstein distance and over all convex sets, and justifies a multiplier bootstrap for inference. The bounds improve previously known rates for dependent data and approach the best known rates for independent data under short memory and light tails.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The claimed near-optimal dimension dependence of the W2 rate is not supported: the bound scales as d, while independent coordinate examples achieve √(d/n).","rationale":"The reader's weakest_assumption is the eigenvalue lower bound λ* > 0. That is an explicit, standard high-dimensional assumption, and the proofs use it consistently; I do not see a hidden gap there. The more load-bearing concern is the paper's headline optimality claim for the W2 rate. Theorem 3.1 and Lemma A.2 produce a factor d, but the independent-coordinate Rademacher example shows the true W2 scales as √(d/n), so the dimension factor is not optimal. This does not invalidate the theorem as an upper bound, but it does invalidate the abstract's statement that the rates are 'nearly optimal with respect to dimensionality' and 'nearly identical to the corresponding best-known GA rates established for independent data' for the Wasserstein part. The proofs are substantial and appear internally consistent; the issue is the interpretation and the claimed optimality. Therefore the paper should be accepted only after the optimality claims in Sections 1.2 and 3.1 are revised to describe the bounds as conservative in d rather than near-optimal in d. The convex-set results and bootstrap consistency may still stand as claimed, since their dimension rates are matched to independent-data benchmarks.","tokens_in":52620,"tokens_out":39430,"duration_ms":488611,"concrete_test":"Compute or bound W2 for iid d-dimensional Rademacher vectors with d = n^{0.8}, n = 10^4. The product structure gives W2^2 = d W2_1D^2 ≈ c^2 d/n, so W2 ≈ c n^{-0.1} → 0. Compare this with the Theorem 3.1 bound O(d n^{-1/2}(log n)^5), which diverges for d = n^{0.8}. If the true W2 vanishes while the stated bound diverges, the claimed 'factor of d is optimal' and the near-optimality of the dimension range d = o(n^{1/2}) are refuted.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The abstract and Section 1.2 advertise the W2 Gaussian approximation rates as nearly optimal in both dimension n and dimensionality d, and Section 3.1 explicitly says 'we expect a factor of d to be optimal when multiplied into the bound.' The proof of Lemma A.2 and the final bound in Theorem 3.1 indeed carry a factor d, giving W2 = O(d n^{-1/2} (log n)^5) in the light-tailed, short-memory case. This dimension factor is not near-optimal. Consider iid X_i with independent Rademacher coordinates, which satisfies Assumption 1(b) with λ*=1 and Assumption 2(b). By tensorization of the 2-Wasserstein distance for product measures, W2(S_n/√n, N(0,I_d))^2 = d W2_1D(Binomial(n,1/2), N(0,1))^2 ≈ c^2 d/n, so W2 ≈ c√(d/n), which vanishes for d=o(n). For d=n^{0.8}, the true W2 tends to 0, while the paper's bound d n^{-1/2}(log n)^5 diverges. The theorem's inequality remains a valid upper bound, but the central claim that the results are nearly optimal in d, or nearly identical to the best-known independent-data rates, is not correct. This overstatement affects the paper's main advertised contribution, though not the validity of the bounds themselves.","agreement_with_reader":"disagree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper develops Gaussian approximation (GA) bounds for sums of triangular arrays of d-dimensional non-stationary Bernoulli-shift time series. Theorem 3.1 bounds the 2-Wasserstein distance between the normalized sum and a matched-covariance Gaussian by O(d n^{1/r-1/2} log n) under finite p-th moments and polynomial memory, and by O(d n^{-1/2}(log n)^5) under exponential moments and exponential memory. Theorem 3.2 gives convex-set variation bounds that decay at rates permitting d up to O(n^{2/5-δ}) under short memory and light tails. Section 4 proves consistency of a multiplier bootstrap for both distances via Gaussian coupling lemmas. Sections 5 and 6 apply the theory to combined L2/L∞ regression inference and thresholded inference, with simulations.","tokens_in":52912,"tokens_out":9707,"duration_ms":119435,"significance":"If the bounds and bootstrap results are correct, the paper extends high-dimensional Gaussian approximation beyond hyper-rectangles to all convex sets and to the 2-Wasserstein distance for non-stationary time series, with explicit dependence on d and n and with constructive Gaussian couplings. The appendices are unusually detailed: they include an L2 truncation lemma, martingale embedding, blocking and M-dependence arguments, Stein's method, and bootstrap consistency proofs, and they contain no fitted parameters or circular steps. These are genuine strengths. However, the advertised near-optimality of the W2 dimension dependence and the sufficiency of the conditions in the application theorems both need to be reassessed before the paper is accepted.","major_comments":[{"comment":"The advertised 'nearly optimal ... with respect to dimensionality' for the W2 rate is not supported by the stated bounds. For iid Rademacher coordinates, Assumption 1(b) and Assumption 2(b) hold, and tensorization gives W2(S_n/√n, N(0,I_d))^2 = d W2_1D(Binomial(n,1/2), N(0,1))^2 ≈ c^2 d/n, so W2 ≈ c√(d/n). The theorem's bound O(d n^{-1/2}(log n)^5) is larger by a factor √d(log n)^5 and diverges for d = n^{0.8}, where the true distance tends to 0. The bound itself remains a valid upper bound, but the statements in the Abstract, Section 1.2, and Section 3.1 that a 'factor of d' is optimal and that the rates are nearly optimal in d are incorrect as written.","section":"Theorem 3.1 and Section 3.1"},{"comment":"Under the stated condition d^2/n → 0 with fixed p > 4, the displayed upper bounds do not tend to zero, so the claimed asymptotic validity of the bootstrap tests does not follow from those theorems. For example, with d = n^{1/2-ε} and ε sufficiently small, the term d^{7/8} n^{1/p-1/4} in Theorem 5.1 behaves like n^{3/16-7ε/8+1/p}, which diverges for, say, p = 5 and ε = 0.01; the same is true for the d^{3/2} n^{2/p-1/2} term in Theorem 5.2. The proofs in Appendices D and E bound these quantities only in probability and do not make them o(1) under the stated assumptions. Either the dimension condition must be strengthened to a rate such as d = o(n^{2/7-δ}) (with the power depending on p), or the rates in the theorems must be corrected.","section":"Theorems 5.1 and 5.2"}],"minor_comments":[{"comment":"The notation |·| is used for the Euclidean norm of vectors, the operator norm of matrices, and the absolute value of scalars. This is acceptable after the definitions, but it occasionally makes expressions such as those in Eq. (37) harder to parse; a separate symbol for the operator norm would improve readability.","section":"Section 2"},{"comment":"In the proof of Lemma A.3, the text refers to 'Assumption (i)' and later to '(I)' in the exponential-moment case; these labels should be unified.","section":"Lemma A.3 proof"},{"comment":"The simulation study uses n = 500 and d = 25, giving d^2/n = 1.25, which does not satisfy d^2/n → 0 required by Theorems 5.1 and 5.2. This does not invalidate the simulation results, but the mismatch between the asymptotic condition and the simulation design should be acknowledged.","section":"Section 6"},{"comment":"There are typographical errors in Appendix E ('Lipstchiz' appears twice), and the long display following Eq. (58) is written without additional parentheses, making it harder to follow than necessary.","section":"Appendix E"}],"recommendation":"major_revision","confidential_remarks":"The technical core of the paper appears sound and the appendices are detailed, but the near-optimality claim for the W2 dimension rate and the sufficiency of the application-theorem conditions both need substantive correction. These are fixable within the manuscript's scope, so I recommend major revision rather than rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two results here are worth knowing about. The first is a convex-set Gaussian approximation for non-stationary time series of diverging dimension, with a dimension range that matches the best independent-data result (d up to o(n^{2/5})). The second is a 2-Wasserstein GA with the right n-rate (n^{-1/2} up to logs) under short memory and light tails. The bootstrap theory for both is also new, and the two applications—combined L2/L∞ testing and thresholded inference—are sensible. The appendices do real work; the proofs are long but explicit, and the conditions are stated clearly. No circular fitting shows up; the self-citations are to earlier tools, not to the conclusion. The soft spot is the dimension-dependence claim for the Wasserstein bound. The theorem gives W2 = O(d n^{-1/2} (log n)^5) in the exponential-moment case. That's a correct upper bound, but it is not near-optimal in d. The stress-test example is right: for iid Rademacher vectors with independent coordinates—which satisfy the assumptions—the actual W2 is O(√(d/n)). So the d factor is loose by a factor of √d, and the paper's statement that 'we expect a factor of d to be optimal' is not supported. This is an overstatement in the abstract and Section 3.1, not a flaw in the theorem itself. The convex GA dimension range is fine; it matches known independent-data bounds. The eigenvalue assumption on Cov(X_n/√n) is strong but standard; it shows up in the covariance-correction step and bootstrap lemmas, so if λ* decays with n the rates become worse. Who should read this: anyone working on high-dimensional time series inference, especially for statistics that are convex or Lipschitz functionals. It deserves a serious referee. I'd send it to a good theory journal and ask the authors to temper the near-optimality language and add a comment about the product-measure gap.","headline":"Real new GA bounds for non-stationary high-dimensional time series, but the Wasserstein near-optimal dimension claim is not backed by the proof.","tokens_in":53472,"tokens_out":9722,"would_cite":true,"duration_ms":110634,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62M10","60B12","62F40"],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper proves that normalized sums of a broad class of high-dimensional non-stationary time series are close to a matching Gaussian in 2-Wasserstein distance and uniformly over all convex sets, at rates that are nearly optimal in the…","keywords":["non-stationary time series","high-dimensional inference","Gaussian approximation","2-Wasserstein distance","convex sets","physical dependence measures","multiplier bootstrap","central limit theorem"],"falsifier":"Build a triangular array where all d coordinates share a common scalar non-stationary process plus independent noise whose variance shrinks like $n^{{-1}}$, so the smallest eigenvalue of Cov(X_n/√n) tends to zero. For growing d, the paper's covariance-correction step would require a blowup of order λ_*^{-1/2}; computing the actual W2 distance, for example with Gaussian innovations where W2 has a closed form, would show whether the claimed O(d $n^{{-1/2}}$ (log n)^5) rate still holds or fails.","tokens_in":52427,"feed_emoji":"📈","tokens_out":5446,"duration_ms":66514,"temperature":0.7,"pith_summary":"This paper tries to establish a general Gaussian approximation for triangular arrays of high-dimensional non-stationary time series: the normalized sum X_n/√n is close, in 2-Wasserstein distance and uniformly over all convex sets, to a Gaussian vector with the same covariance. The rates depend on sample size, dimension, moment order, and the decay of physical dependence; for short memory and light tails they are nearly optimal in n, and the approximation holds while d diverges at rates up to o($n^{{1/2}}$) for Wasserstein distance and o($n^{{2/5}}$) for convex sets. The argument includes a multiplier bootstrap whose validity is proved for both distances, so the Gaussian approximation can be implemented without knowing the dependence structure. If these claims are right, many inference procedures that use Lipschitz or convex functionals of time-series sums—quadratic forms, thresholded estimators, combined L2/L∞ tests—inherit Gaussian limit behavior at near-parametric rates.","feed_headline":"Time-series sums approximate Gaussians even as dimension grows","feed_subtitle":"Explicit rates for Wasserstein distance and all convex sets extend bootstrap inference to non-stationary, high-dimensional data.","key_machinery":"The paper represents a non-stationary d-dimensional series as a triangular array of time-varying Bernoulli shifts x_i = G_i(F_i), and measures dependence by physical dependence coefficients θ_{k,j,q} and their cumulative tail Θ_{m,p}. The proof approximates X_n in stages: truncate coordinates, replace the series by an M-dependent approximation, divide into blocks whose border innovations are conditioned on, apply a martingale-embedding Gaussian approximation to conditionally independent block sums, and then correct the covariance mismatch using the uniform lower eigenvalue bound λ_*. Truncation is controlled by a new L2-norm truncation lemma with error O(√n $n^{{1/p}}$, which the paper claims is the first optimal-rate L2 truncation bound for time series. The multiplier bootstrap statistic τ_{n,L} = Σ_i B_i ψ_{i,L} uses overlapping block sums ψ_{i,L} and i.i.d. standard normals B_i; its consistency is proved through Gaussian coupling lemmas that compare covariances by Frobenius or max norm.","core_discovery":"The central discovery is a pair of explicit bounds. Under finite p-th moments and polynomially decaying dependence, the 2-Wasserstein distance between X_n/√n and a Gaussian vector Y_n/√n with the same covariance is O(d $n^{{1/r-1/2}}$ log n), where r is determined by p and the dependence exponent; under exponential moments and exponentially decaying memory the rate becomes O(d $n^{{-1/2}}$ (log n)^5). For the collection of all convex sets, the distance is bounded by a minimum of two terms and vanishes for d as large as O($n^{{2/5-δ}}$) under sufficiently short memory and light tails. The proof constructs the approximating Gaussian explicitly, then justifies a multiplier bootstrap based on overlapping block sums, so the Gaussian law is not only shown to exist but is accessible computationally.","pith_inferences":["Because W2 control is stronger than control in probability, the same bounds should transfer to strong-approximation statements in the Lévy-Prohorov metric for diverging dimension; the paper notes the connection but does not develop it as a separate theorem.","The covariance-correction step depends on the uniform eigenvalue bound, which suggests the method may be adapted to ill-conditioned settings by adding regularization or by replacing λ_* with a data-dependent threshold; a testable extension is to verify whether bootstrap coverage remains valid when λ_* is allowed to decay slowly with n.","The explicit Gaussian couplings in the paper could yield finite-sample confidence sets for functions of the covariance matrix, since Lemma 4.2 gives a constructive coupling rather than an existence statement."],"forward_implications":["For any Lipschitz functional of X_n/√n, the Wasserstein bound controls the difference in expectation, so threshold-type statistics are asymptotically Gaussian with the same covariance.","For combined L2 and L∞ regression tests, the convex Gaussian approximation justifies the multiplier bootstrap critical value under H0 whenever d^2/n → 0 and the moment and dependence assumptions hold.","For thresholded inference, the bootstrap error is bounded explicitly and becomes negligible when d grows slower than n^{1/3} with finite high moments.","The bootstrap is consistent uniformly over all convex sets, not just hyper-rectangles, so inference for quadratic forms, eigenvalue-type statistics, and other non-rectangular regions becomes accessible through one unified procedure."],"supporting_citations":[{"why":"Supplies the Rosenblatt transform that represents every triangular array of non-stationary time series as time-varying Bernoulli shifts.","marker":"Rosenblatt (1952)"},{"why":"Introduces the physical dependence measures used throughout the paper to quantify temporal dependence.","marker":"Wu (2005)"},{"why":"Supplies the martingale-embedding Gaussian approximation with near-optimal dimension and sample-size rates that the paper extends to blocks.","marker":"Eldan, Mikulincer and Zhai (2020)"},{"why":"Provides the prior sequential Gaussian approximation for high-dimensional non-stationary time series and technical steps used in Lemma A.2.","marker":"Mies and Steland (2023)"},{"why":"Provides boundary-conditioning methods and optimal univariate strong-approximation rates that motivate the fixed-dimension benchmark.","marker":"Berkes, Liu and Wu (2014)"},{"why":"Supplies the M-dependence, blocking, and covariance-correction steps that the proof generalizes to diverging dimension.","marker":"Karmakar and Wu (2020)"},{"why":"Gives the dimension-dependent convex-set Gaussian approximation bound used in Lemma B.1.","marker":"Bentkus (2003)"},{"why":"Establishes the best-known independent-data convex Gaussian approximation rate that the paper matches under dependence.","marker":"Fang and Koike (2024)"},{"why":"Supplies the multiplier bootstrap construction whose high-dimensional extension is validated in the paper.","marker":"Zhou (2013)"}],"fun_headline_variants":["Gaussian approximation extended to non-stationary high-dim series","Nearly optimal Wasserstein rates for dependent high-dim sums","Convex-set Gaussian bounds for diverging-dimension time series","Bootstrap Gaussian inference justified for non-stationary data","Gaussian limits for high-dim dependent series with light tails"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing condition is Assumption 1's uniform lower bound on the smallest eigenvalue of Cov(X_n/√n). The covariance-correction step divides by this eigenvalue, and if it decays with n or approaches zero, the stated Wasserstein and bootstrap error rates collapse through the factor λ_*^{-1/2}.","fun_headline_variants_meta":{"raw":{"variants":["Gaussian approximation extended to non-stationary high-dim series","Nearly optimal Wasserstein rates for dependent high-dim sums","Convex-set Gaussian bounds for diverging-dimension time series","Bootstrap Gaussian inference justified for non-stationary data","Gaussian limits for high-dim dependent series with light tails"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000257,"raw_usage":{"total_tokens":1568,"prompt_tokens":925,"completion_tokens":643,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":541,"completion_tokens_details":{"reasoning_tokens":561}},"tokens_in":541,"tokens_out":643,"duration_ms":7611,"temperature":1.0,"reasoning_tokens":561,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T05:03:42.676677+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Build a triangular array where all d coordinates share a common scalar non-stationary process plus independent noise whose variance shrinks like $n^{{-1}}$, so the smallest eigenvalue of Cov(X_n/√n) tends to zero. For growing d, the paper's covariance-correction step would require a blowup of order λ_*^{-1/2}; computing the actual W2 distance, for example with Gaussian innovations where W2 has a closed form, would show whether the claimed O(d $n^{{-1/2}}$ (log n)^5) rate still holds or fails.","supporting_citations":[{"cited_title":"( 1952 )","cited_arxiv_id":null,"evidence_quote":"Supplies the Rosenblatt transform that represents every triangular array of non-stationary time series as time-varying Bernoulli shifts."},{"cited_title":", Mikulincer , Dan D","cited_arxiv_id":null,"evidence_quote":"Supplies the martingale-embedding Gaussian approximation with near-optimal dimension and sample-size rates that the paper extends to blocks."},{"cited_title":"Steland , Ansgar A","cited_arxiv_id":null,"evidence_quote":"Provides the prior sequential Gaussian approximation for high-dimensional non-stationary time series and technical steps used in Lemma A.2."},{"cited_title":", Liu , Weidong W","cited_arxiv_id":null,"evidence_quote":"Provides boundary-conditioning methods and optimal univariate strong-approximation rates that motivate the fixed-dimension benchmark."},{"cited_title":"Wu , Wei Biao W","cited_arxiv_id":null,"evidence_quote":"Supplies the M-dependence, blocking, and covariance-correction steps that the proof generalizes to diverging dimension."},{"cited_title":"( 2003 )","cited_arxiv_id":null,"evidence_quote":"Gives the dimension-dependent convex-set Gaussian approximation bound used in Lemma B.1."},{"cited_title":"Koike , Yuta Y","cited_arxiv_id":null,"evidence_quote":"Establishes the best-known independent-data convex Gaussian approximation rate that the paper matches under dependence."},{"cited_title":"( 2013 )","cited_arxiv_id":null,"evidence_quote":"Supplies the multiplier bootstrap construction whose high-dimensional extension is validated in the paper."}],"review_version":1}