{"id":"030dde3c-416f-470e-942e-b12065fbee0f","arxiv_id":"2607.03367","paper_version":1,"verdict":"ACCEPT","confidence":"HIGH","novelty_score":6.5,"correctness_risk":"low","formal_verification":"none","parameter_count":2,"one_line_summary":"Minimax risk for estimating KSD is governed by the Hilbert–Schmidt norm of the Stein covariance, attained by a square-root U-statistic but not by the standard V-statistic.","lead":"The paper shows that the sharpest possible accuracy for estimating Kernel Stein Discrepancy is set by the Hilbert–Schmidt size of the Stein covariance operator, not its trace. A simple debiased square-root U-statistic attains this rate, while the usual plug-in estimator is worse by a factor that can grow exponentially with dimension.","discovery_kind":"extension","skeptic_critique":{"model":"grok-4.5","headline":"No significant objection identified","rationale":"The paper’s strongest claim is a clean fixed-target minimax result that refines the earlier global-rate lower bound of Cribeiro-Ramallo et al. The upper bound for the debiased U-statistic follows from standard U-statistic variance formulas plus the 1/2-Hölder property of the positive square root; the matching lower bound is obtained by a two-fuzzy-hypothesis construction along the leading eigen-directions of C⋆, with a careful truncation argument that removes the artificial boundedness assumption on the coordinates gj. Both sides are controlled by the same spectral functional tr(C⋆²)^{1/4}/√n, so the identification of the Hilbert–Schmidt scale is tight. The V-statistic’s pinning to the trace scale is elementary (reverse triangle inequality plus a fourth-moment interpolation) and is therefore not an artifact. The known-score assumption is a genuine practical limitation, but it is already stated as such in the conclusion and does not affect the correctness of the pure estimation theory under Assumption 1. Consequently the reader’s ACCEPT verdict stands; no adjustment is warranted.","tokens_in":16600,"tokens_out":516,"duration_ms":4889,"concrete_test":"Independently recompute the Gaussian spectral quantities of Proposition 10 (tr(C⋆) and tr(C⋆²)) by Monte-Carlo integration of KP0(X,Y)² under X,Y~N(0,Id) for γ=1/2 and d=2,5,10; if the numerical values deviate from the closed form by more than a few percent after 10^6 samples, the explicit V/U gap formulas would need re-examination.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim (Corollary 6) is that the minimax risk of estimating KSD(P0,P) over the target-normalized class P(A) is of order √(∥C⋆∥HS/n), attained by the positive-part square-root U-statistic while the V-statistic is pinned to the larger trace scale. Matching upper and lower bounds are supplied (Theorems 2 and 5, Proposition 7), the spectral gap is made explicit for the Gaussian target (Section 3.4), and the known-score premise is already flagged by the authors as outside the present scope. No internal inconsistency or gap that would overturn the fixed-target minimax identification was found.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper studies minimax estimation of the scalar Kernel Stein Discrepancy KSD(P0,P) for a fixed target P0 known only through its score, from n i.i.d. samples of P. It identifies the sharp spectral constant as the Hilbert–Schmidt norm of the Stein covariance operator C⋆, so that the minimax scale is √(∥C⋆∥HS/n). This scale is attained by the positive-part square-root U-statistic that removes diagonal terms; the usual plug-in V-statistic remains at the larger trace scale √(tr(C⋆)/n) and is therefore suboptimal by reff(C⋆)^{1/4}. Matching upper bounds (Theorems 1–2), a fuzzy-hypothesis lower bound over the target-normalized class P(A) (Theorem 5), and an explicit Gaussian calculation showing an exponential-in-dimension gap for fixed bandwidth (Section 3.4) are provided.","tokens_in":16777,"tokens_out":764,"duration_ms":5749,"significance":"The contribution is a clean refinement of existing rate-optimal results for KSD estimation: once the target is fixed, the relevant constant is the HS norm of C⋆ rather than its trace, and the debiased U-statistic is minimax optimal while the V-statistic is not. The argument is standard and transparent (reverse triangle inequality, U-statistic variance with degeneracy at the target, Tsybakov two fuzzy hypotheses along Stein eigen-coordinates), the class P(A) is well-motivated by the off-diagonal representation of KSD², and the Gaussian spectral formulas make the V/U gap concrete and dimension-dependent. The known-score premise is already flagged by the authors as outside the present scope. The result is of clear interest for KSD-based diagnostics and goodness-of-fit, especially in moderate-to-high dimension.","major_comments":[],"minor_comments":[{"comment":"In the abstract and introduction the minimax scale is written √(∥C⋆∥HS/n); later (Theorem 2, Corollary 6) it is equivalently tr(C⋆²)^{1/4}/√n. A single consistent notation for the HS norm would reduce minor ambiguity for readers who do not immediately equate the two.","section":null},{"comment":"Figure 1 caption and the surrounding text in §3.4 refer to Monte Carlo risks of [KSD_V and [KSD_U; the notation for the estimators is slightly inconsistent with the definitions (4)–(5) (hat vs. bracket). Aligning the symbols would help.","section":null},{"comment":"Appendix D: the truncation argument that removes the boundedness assumption on the Stein coordinates g_j is correct but dense; a short remark that the same construction works under a uniform L^{2+δ} moment on the first m coordinates would make the scope of the lower bound clearer without lengthening the proof.","section":null},{"comment":"Typographical: several displayed equations have line-break artifacts (e.g., the definition of reff(C⋆) and the Gaussian tr(C⋆²) formula). These are purely presentational.","section":null}],"recommendation":"accept","confidential_remarks":"The manuscript is short, self-contained, and technically solid. It is a natural fit for a statistics journal that publishes sharp minimax results for kernel methods. I see no novelty or citation concerns; the comparison with Cribeiro-Ramallo et al. is fair and correctly positions the fixed-target spectral refinement."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The real news is the spectral constant, not the n^{-1/2} rate. With the target fixed, the minimax risk of estimating KSD(P0,P) over the natural off-diagonal class P(A) is of order √(∥C⋆∥_HS / n). The positive-part square-root U-statistic hits it; the plug-in V-statistic is stuck at the larger trace scale and loses a factor reff(C⋆)^{1/4}. For a Gaussian target and fixed-bandwidth Gaussian kernel that factor is exponential in dimension. That is new relative to the global-rate result of Cribeiro-Ramallo et al., which treated both target and sampling law as free and therefore could not see the constant.\n\nThe math is careful and complete. Upper bounds are standard reverse-triangle and U-statistic variance, with the degeneracy at the target made explicit. The lower bound uses Tsybakov fuzzy hypotheses along the leading eigen-directions of C⋆, with a clean truncation argument to remove the boundedness assumption on the coordinates. The Gaussian calculations are closed-form and match the Monte Carlo figure. Citation pattern is appropriate; no circularity.\n\nSoft spots are minor and already flagged by the author. The score is assumed known exactly, so the result does not cover score-matching or learned energy-based models; that is outside the stated scope. The class P(A) is target-normalized by tr(C⋆^{2}), which is the right choice for exposing the spectral geometry but means the constant A is not universal. Neither issue undercuts the fixed-target identification.\n\nThis is for people who use KSD for diagnostics or GOF and for anyone working on minimax rates of kernel discrepancies. It is a solid math.ST paper that deserves a serious referee. I would accept it for peer review and would cite the V/U gap and the Gaussian dimension calculation.","headline":"Clean fixed-target minimax theory that pins KSD estimation to the HS scale of C⋆ and shows the usual V-statistic is strictly worse by reff^{1/4}.","tokens_in":17392,"tokens_out":507,"would_cite":true,"duration_ms":5001,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62G05","62G20","46E22"],"pacs":[],"model":"grok-4.5","headline":"The sharp minimax scale for estimating kernel Stein discrepancy is set by the Hilbert–Schmidt norm of the Stein covariance, and the debiased U-statistic attains it while the plug-in V-statistic stays at the larger trace scale.","keywords":["kernel Stein discrepancy","minimax estimation","Hilbert–Schmidt norm","U-statistic","V-statistic","Stein covariance operator","effective rank","goodness-of-fit"],"falsifier":"At a high-dimensional standard Gaussian target with fixed-bandwidth Gaussian kernel, compute Monte-Carlo risks of both the V-statistic and the square-root U-statistic for growing d at fixed n; if the observed gap fails to track (1+8γ)^{d/8} while the U-statistic risk tracks the explicit Hilbert–Schmidt formula, the spectral claim is false.","tokens_in":17506,"feed_emoji":"📐","tokens_out":730,"duration_ms":5799,"temperature":0.7,"pith_summary":"Kernel Stein discrepancy (KSD) measures how far a sample sits from a fixed target that is known only through its score function. The paper asks for the precise statistical difficulty of estimating that scalar discrepancy from n independent draws. It shows that the governing constant is not the familiar trace of the Stein covariance operator, but its Hilbert–Schmidt norm, so the minimax risk is of order the square root of that norm over n. A simple debiased estimator—the positive square root of the off-diagonal U-statistic—achieves this rate. The ordinary plug-in V-statistic, which keeps the diagonal terms, is stuck at the larger trace scale and therefore loses a factor equal to the fourth root of the effective rank of the same operator. For a Gaussian target and a fixed-bandwidth Gaussian kernel that factor grows exponentially with dimension, so the choice between the two estimators is not cosmetic.","feed_headline":"KSD estimation has a sharp Hilbert–Schmidt scale","feed_subtitle":"Debiased U-statistic attains it; the plug-in V-statistic loses an exponential factor in dimension","key_machinery":"The Stein covariance operator C⋆ = E_{P0}[ξ_{P0}(X) ⊗ ξ_{P0}(X)], whose Hilbert–Schmidt norm (equivalently tr(C⋆²)^{1/2}) sets the minimax constant, together with the degeneracy of the associated U-statistic at the target forced by Stein’s identity.","core_discovery":"When the target P0 is fixed, the minimax risk of estimating KSD(P0,P) over a natural off-diagonal moment class is of order the square root of the Hilbert–Schmidt norm of the Stein covariance C⋆ divided by n. The positive-part square-root U-statistic attains this scale; the plug-in V-statistic cannot, remaining pinned to the larger trace scale √(tr(C⋆)/n) and therefore suboptimal by the fourth root of the effective rank of C⋆.","pith_inferences":[],"forward_implications":[],"fun_headline_variants":["Minimax KSD risk set by Hilbert–Schmidt norm of Stein covariance","Positive-part U-statistic attains sharp HS scale for KSD estimation","V-statistic stuck at trace scale; U-stat hits HS minimax for KSD","KSD minimax rate is HS of Stein operator, not its trace","Debiased U-stat matches HS scale; plug-in loses rank factor"],"cache_read_input_tokens":128,"weakest_assumption_plain":"The target score function is treated as known exactly; any estimation error in the score is outside the present bounds and could erase the claimed Hilbert–Schmidt advantage.","fun_headline_variants_meta":{"raw":{"variants":["Minimax KSD risk set by Hilbert–Schmidt norm of Stein covariance","Positive-part U-statistic attains sharp HS scale for KSD estimation","V-statistic stuck at trace scale; U-stat hits HS minimax for KSD","KSD minimax rate is HS of Stein operator, not its trace","Debiased U-stat matches HS scale; plug-in loses rank factor"]},"model":"grok-4.5","effort":"low","cost_usd":0.004642,"raw_usage":{"total_tokens":1311,"prompt_tokens":758,"num_sources_used":0,"completion_tokens":105,"cost_in_usd_ticks":46420000,"prompt_tokens_details":{"text_tokens":758,"audio_tokens":0,"image_tokens":0,"cached_tokens":128},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":448,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":758,"tokens_out":105,"duration_ms":3719,"temperature":1.0,"reasoning_tokens":448,"cache_read_input_tokens":128,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-12T02:58:33.490466+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"At a high-dimensional standard Gaussian target with fixed-bandwidth Gaussian kernel, compute Monte-Carlo risks of both the V-statistic and the square-root U-statistic for growing d at fixed n; if the observed gap fails to track (1+8γ)^{d/8} while the U-statistic risk tracks the explicit Hilbert–Schmidt formula, the spectral claim is false.","supporting_citations":[],"review_version":1}