{"id":"6150b709-58de-4c8c-ad56-f7616aa83795","arxiv_id":"2510.22368","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Sequential CUSUM, Page, and recycling detectors based on degenerate U-statistics have limiting null distributions and detection-delay laws under square-summable kernels.","lead":"This paper develops online detectors for changes in the full distribution of a data stream, based on degenerate U-statistics, and proves their asymptotic behavior under weaker conditions than previous work. It also proposes a detector that recycles old monitored observations into the training baseline, with simulations and an application to infant heart-rate data.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Theorem 4.4's Monte Carlo critical-value proof relies on an unverified spectral-consistency theorem; if that theorem needs more than Assumption 2.3, the 'square summability only' claim does not cover the practical method.","rationale":"The reader's stated weakest assumption was independence, which is indeed a limitation: the Wiener limits and martingale tail bounds in the supplement all require Assumption 2.2, and serial dependence would invalidate the critical values. However, independence is an explicit model assumption and is checked via BDS in the empirical section; it is a scope limitation rather than an internal gap in the central claim. The more load-bearing concern for the paper's distinctive claim—'only square summability'—is the unverified eigenvalue-consistency step in Theorem 4.4. The proof imports a strong ℓ^2 spectral-consistency result without stating or checking its hypotheses, and the paper's own kernels are unbounded. If the imported theorem is not applicable, the Monte Carlo critical-value approximation fails, and the practical implementation advertised in the abstract and Section 6 is unsupported. This does not necessarily refute Theorems 3.1–3.5, so the appropriate disposition remains conditional on resolving this point; hence the reader's overall CONDITIONAL verdict is unchanged. I mark agreement as partial because the reader did mention the imported, unverified eigenvalue-consistency result in the rationale, but did not identify it as the weakest load-bearing assumption.","tokens_in":66303,"tokens_out":9475,"duration_ms":107069,"concrete_test":"Write out the exact hypotheses of Koltchinskii and Gin\\'e (2000, Theorem 3.1). Then check them for kernels satisfying only Assumption 2.3, in particular for h(x,y)=\\|x-y\\|^2 on Gaussian data and h(x,y)=[1-\\exp(-\\|x-y\\|^2/(2a^2))]^{1/2}. If the theorem requires boundedness of h or a stronger moment condition, Theorem 4.4 is not proved as stated. If the theorem does apply, verify explicitly that its conditions are met by these kernels, or give a direct proof that inf_\\pi \\sum_{\\ell}(\\lambda_\\ell-\\hat\\lambda_{\\pi(\\ell),m})^2 \\to 0 a.s. under only Eh^2(X,Y)<\\infty.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The proof of Theorem 4.4 is the only place where the estimated eigenvalues \\hat\\lambda_{\\ell,m} of the empirical matrix A_m in (4.7) are connected to the true eigenvalues \\lambda_\\ell. The paper asserts that, by Koltchinskii and Gin\\'e (2000, Theorem 3.1), inf_\\pi \\sum_{\\ell\\ge 1}(\\lambda_\\ell - \\hat\\lambda_{\\pi(\\ell),m})^2 \\to 0 almost surely. This is a strong ℓ^2-consistency statement, but the cited theorem is not stated and its hypotheses are not checked. Assumption 2.3 only requires Eh^2(X,Y)<\\infty; for the paper's leading examples, h is unbounded (e.g., h(x,y)=\\|x-y\\|^2). General spectral-consistency results for empirical Gram matrices typically require bounded kernels or additional moment/regularity assumptions, and it is not immediate that the cited theorem applies to unbounded kernels satisfying only Assumption 2.3. Without this eigenvalue consistency in ℓ^2, the conditional finite-dimensional convergence in (4.11) does not follow: the random weights \\hat\\lambda_{\\ell,m} must converge to the true \\lambda_\\ell in exactly the metric that controls the tail of the infinite Gaussian sum. The Monte Carlo critical-value procedure used in Section 5 and advertised in the abstract and Section 6 therefore rests on an unverified link. Theorems 3.1–3.5 may still be sound, but the headline claim that 'all the asymptotic theory is derived by requiring only the square summability of the eigenvalues' is not established for Theorem 4.4.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper develops online changepoint-detection procedures based on degenerate U-statistics. Three detectors are treated: a CUSUM-type detector, a Page-type detector, and a novel 'repurposing' detector that expands the training sample with old monitoring observations. For each, null weak limits are proved under open- and closed-ended monitoring, and detection-delay limit theorems are given for early and late changes. The paper also proposes a retrospective test for the stability of the training sample, a Monte Carlo method for critical values, kernel examples, simulations in R^5, and an application to infant ECG data. The headline claim is that all the asymptotic theory is derived under only square summability of the kernel eigenvalues, rather than the absolute summability used in earlier work.","tokens_in":66704,"tokens_out":30573,"duration_ms":294601,"significance":"If the results hold, this is a substantial contribution to the U-statistic-based monitoring literature. The paper unifies CUSUM, Page, and recycling schemes under one framework; the delay distribution results for both early and late changes are more complete than in most prior work; and replacing absolute summability by square summability is a genuine technical improvement that is also practically checkable. The proofs are detailed and largely self-contained, and the empirical section is careful, including checks of the independence and moment assumptions. The main caveat is that the Monte Carlo critical-value justification depends on an unstated spectral-consistency theorem whose applicability to the paper's kernel class is not verified; this is the only load-bearing gap I found.","major_comments":[{"comment":"The Monte Carlo critical-value approximation, used in Section 5 and advertised in the abstract, is justified by the assertion that inf_pi sum_l (lambda_l - hat_lambda_{pi(l),m})^2 -> 0 a.s., stated in the proof of Theorem 4.4 to follow from Koltchinskii and Giné (2000, Theorem 3.1). The cited theorem is not stated and its hypotheses are not verified. The paper's leading kernels are unbounded (e.g., h(x,y)=||x-y||^2), and Assumption 2.3 only gives E h^2(X,Y)<infinity. It is not immediate that the theorem applies to the doubly centered empirical matrix A_m in (4.7), nor that it gives the required l^2 convergence of eigenvalues. Without this l^2 spectral consistency, the conditional finite-dimensional convergence in (4.11) does not follow. Since the square-summability claim is repeated as the paper's main advance, this gap is load-bearing. Please state the theorem and verify its hypotheses,","section":"Section 4.4 (Theorem 4.4), proof in Appendix D"},{"comment":"Even if the cited l^2 spectral-consistency result is granted, the proof as written establishes convergence for the process with the optimally permuted coefficients pi_m, while the statistics in (4.8)-(4.10) use the identity-ordered eigenvalues hat_lambda_{ell,m}. The displayed bounds in the tightness part switch to the ordered eigenvalues without comment. The missing step is exchangeability of the independent Wiener processes: permuting the coefficients does not change the conditional law. This must be stated explicitly; as written, the transition from the permuted process to hat_Gamma_m is a logical gap.","section":"Proof of Theorem 4.4 (ordered vs permuted eigenvalues)"}],"minor_comments":[{"comment":"The domain in the definition of Gamma(u) is written as '0 <= v <= 1'; this should be '0 <= u <= 1'.","section":"Equation (3.8)"},{"comment":"The theorem statement omits the positive semidefiniteness and continuity conditions on K that the proof uses via the Moore-Aronszajn theorem. Please add them to the statement.","section":"Theorem 4.3"},{"comment":"The assertion that using only a fraction (e.g., m/2) of the eigenvalues 'still yields the same result' is unproved and not immediate. It should be justified or removed.","section":"Section 4.4"},{"comment":"The infinite sums in the proof use hat_lambda_{ell,m} for ell > m, but hat_lambda_{ell,m} is only defined for ell <= m. Please define these coefficients as zero for ell > m.","section":"Proof of Theorem 4.4"},{"comment":"The text says the D(3) scheme improves on D(1) and D(2) across all alternatives and kernels for both strong and weak changes. Table 5.3 shows D(3) has power 0.251 for weak HA,2 with kernel h(1), while D(1) and D(2) have 0.577 and 0.592. Please qualify this claim.","section":"Section 5.1, after Tables 5.2-5.3"},{"comment":"There is a typo: '{Xi, i >= i}' should be '{Xi, i >= 1}'.","section":"Assumption 2.2"}],"recommendation":"major_revision","confidential_remarks":"The substantive risk is concentrated in Theorem 4.4. If the authors can verify that the cited Koltchinskii-Giné theorem applies under Assumption 2.3 to the centered empirical matrix (4.7), I would be inclined to accept after a careful revision. If not, the square-summability claim should be restricted to Theorems 3.1-3.5 and 4.1-4.3, and the Monte Carlo procedure in Section 4.4 presented as heuristic. I found no circularity in the main proofs; the core null and delay asymptotics do not depend on the authors' own earlier papers."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The short version: the null-limit and delay asymptotics are real, new, and mostly hold up; the Monte Carlo critical-value theorem (4.4) has a load-bearing gap the authors paper over with an unverified citation. Also, the abstract and the body describe different experiments.\n\nWhat's genuinely new: square summability of the kernel eigenvalues replacing absolute summability (a real improvement over Biau et al. 2016, and checkable in practice), late-change detection-delay limits, Theorem 4.3 on strong negative type for δ^{1/2}, and the recycling detector D^(3). The proofs are detailed and, as far as I can see, internally coherent—Lemmas C.1–C.5 supply the martingale bounds and Gaussian approximations, and the delay proofs follow the announced decomposition. The simulations are extensive, and the authors are honest about D^(3)'s poor behavior under weak, late changes.\n\nThe soft spots, in proportion:\n\nTheorem 4.4 is the serious one. The proof rests on an ℓ²-consistency statement for estimated eigenvalues attributed to Koltchinskii and Giné (2000, Thm 3.1), but the hypotheses are never checked. Assumption 2.3 (Eh² < ∞) is all you get, and the leading kernels—||x−y||², for instance—are unbounded. It is not immediate that the cited theorem applies there, and the random weights in (4.8)–(4.10) must converge in exactly the metric that controls the tail of the infinite Gaussian sum. Without that, the conditional convergence in (4.11) doesn't follow. The main results (Theorems 3.1–3.5) don't depend on this, but the practical critical-value procedure used in the simulations does. The headline claim that 'all the asymptotic theory' needs only square summability is not established for Theorem 4.4.\n\nMinor but real: the abstract mentions compressor-sensor data from a metro train; the application is infant ECG heart-rate data. The abstract also claims comparisons with mean-, covariance-, and empirical-CDF-based monitors; the simulations compare only with CUSUM and CUSUM-cov detectors. Probably leftover from an earlier draft, but it's careless and will confuse readers. No code or data archive either.\n\nIndependence (Assumption 2.2) is load-bearing but honestly stated, and the empirical section at least tries to check it. Fine as a limitation, not a flaw.\n\nWho this is for: researchers working on sequential changepoint detection, degenerate U-statistics, or energy-distance methodology. They should rely on the null limits but treat the Monte Carlo critical values as heuristic until Theorem 4.4 is fixed.\n\nSend it to peer review. The core results deserve referee time, but Theorem 4.4 needs either a verified condition or a softened claim, and the abstract needs to match the body.","headline":"Genuinely useful advance in sequential distributional changepoint theory under square summability, but the Monte Carlo critical-value theorem (4.4) rides on an unverified eigenvalue-consistency citation and the abstract doesn't match the text.","tokens_in":67204,"tokens_out":4709,"would_cite":true,"duration_ms":48486,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["60F17"],"pacs":[],"model":"deepseek-v4-flash","headline":"Online detection of distributional changepoints can be built from degenerate U-statistic processes whose kernel eigenvalues need only be square summable, not absolutely summable.","keywords":["degenerate U-statistics","sequential changepoint detection","distributional changepoint","CUSUM","restarting CUSUM","energy distance","square-summable eigenvalues","simulation-based critical values"],"falsifier":"Run a long i.i.d. sequence with no change and a kernel whose eigenvalues are square-summable but not absolutely summable, e.g. λℓ = (−1)ℓ/ℓ, and compare the empirical null rejection frequency of the CUSUM detector with the paper's theoretical critical-value approximation. If the frequency does not approach the nominal level as the training length grows, square summability alone is insufficient.","tokens_in":66190,"feed_emoji":"📈","tokens_out":7069,"duration_ms":71863,"temperature":0.7,"pith_summary":"This paper tries to establish that online monitoring for changes in a distribution can be built on degenerate U-statistic processes under a mild condition: the kernel's eigenvalues need only be square summable, not absolutely summable. It proposes three detectors—a plain CUSUM-type statistic, a restarting CUSUM variant, and a new recycling scheme that absorbs old monitored observations into the training baseline—and derives their null limiting distributions as suprema of weighted infinite sums of squared standard Brownian motions. It also gives limiting laws for detection delays for both early and late changes, enabling a user to anticipate how quickly an alarm will sound. If the central claims are right, applied monitoring of multivariate data can use a wide family of distance-based kernels, with critical values obtained by simulating the limit from estimated eigenvalues, and the theory covers both open-ended and closed-ended monitoring.","feed_headline":"Square-summable kernels suffice for online changepoint detection","feed_subtitle":"CUSUM, restarting, and recycling detectors get null limits and delay laws under one mild eigenvalue condition.","key_machinery":"The carrier of the argument is the degenerate U-statistic built from a symmetric kernel h(x,y); after centring, the kernel has spectral expansion Σ λℓ φℓ(x)φℓ(y), with φℓ orthonormal eigenfunctions of the integral operator Ag(x) = E[h(x,Y)g(Y)]. The detector is m^{−1}k²|U_m(h;k)|, comparing training and monitored observations, and the boundary function g_m(k) = (k/m)/(1+k/m)^β (1+k/m)² controls false alarms. The proofs approximate the U-statistic by squared CUSUM processes of the eigenfunctions; the only eigenvalue condition needed for the remainder bounds is square summability Σ λℓ² < ∞, which already follows from E h²(X,Y) < ∞. This is what lets the limit be an infinite weighted squared-Br","core_discovery":"The central discovery is that the asymptotic behaviour of online distributional-change detectors based on degenerate U-statistics is governed by the weighted process Γ(u) = Σ λℓ (Wℓ²(u) − u), with weights λℓ the eigenvalues of the kernel's integral operator, and that this limit is valid under only Σ λℓ² < ∞. Under the null, the ordinary CUSUM detector, the restarting version, and the recycling detector all converge to functionals of Γ; in particular, the Type I error probability tends to P{sup u^{−β}|Γ(u)| > c}. Under the alternative, the detection delay has a normal limit for early changes and a directly characterised non-Gaussian limit for late changes; the paper also shows that simulation","pith_inferences":["Because only square summability is needed, any kernel with finite second moment is admissible; in particular, unbounded distance kernels such as |x−y|^η can be used for heavy-tailed multivariate data, and the paper's construction results suggest genuinely omnibus monitors can be built from positive-definite kernels.","The theory rests on independence of the observations; extending the results to weak dependence would require additional arguments, and a natural robustness check is to run the proposed detector on an autoregressive null and compare false-alarm rates.","The recycling scheme's window parameters (minimum window size and retention proportion) are not covered by a delay theory; simulation evidence suggests low β and short windows for small late changes, so a data-driven rule for choosing these parameters would be a next step.","If square summability is indeed sufficient, the same proof strategy may transfer to functional or network-valued data, since the kernel would be the only ingredient that needs to change."],"forward_implications":["If correct, asymptotic critical values for open-ended, long-horizon, and short-horizon monitoring are available from the same weighted squared-Brownian functional, with only the supremum interval changing.","The delay laws quantify detection lag: roughly w m^ρ observations after an early break, and a delay proportional to √m for fixed-size late breaks, allowing users to choose the weight β and detector in advance.","The recycling detector is valid under the null and, in simulations, detects strong breaks faster than the other two; for small late breaks, the simulations warn that recycling post-change data can contaminate the baseline.","Simulation-based critical values using empirical eigenvalues of the training sample converge, so in practice one does not need to know the true eigenvalues of the kernel.","The retrospective training-sample test makes the usual no-break-in-training assumption testable, again under square summability."],"fun_headline_variants":["Weak kernel condition powers online distributional change detection","Streaming changepoint detectors drop strong assumptions","Degenerate U-statistics ease online monitoring theory","Online change detection: square-summable kernels are enough","CUSUM and Page detectors now work under lighter kernel assumptions"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The load-bearing premise is that the observations are an independent sequence; without independence, the Brownian-motion limits, martingale tail bounds, and simulation-based critical values are not justified.","fun_headline_variants_meta":{"raw":{"variants":["Weak kernel condition powers online distributional change detection","Streaming changepoint detectors drop strong assumptions","Degenerate U-statistics ease online monitoring theory","Online change detection: square-summable kernels are enough","CUSUM and Page detectors now work under lighter kernel assumptions"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000349,"raw_usage":{"total_tokens":1725,"prompt_tokens":709,"completion_tokens":1016,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":453,"completion_tokens_details":{"reasoning_tokens":940}},"tokens_in":453,"tokens_out":1016,"duration_ms":9674,"temperature":1.0,"reasoning_tokens":940,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-04T08:06:16.507810+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run a long i.i.d. sequence with no change and a kernel whose eigenvalues are square-summable but not absolutely summable, e.g. λℓ = (−1)ℓ/ℓ, and compare the empirical null rejection frequency of the CUSUM detector with the paper's theoretical critical-value approximation. If the frequency does not approach the nominal level as the training length grows, square summability alone is insufficient.","supporting_citations":[],"review_version":1}