{"id":"df38c233-6870-4e10-93ce-355104d5be4f","arxiv_id":"2507.09496","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"high","formal_verification":"none","parameter_count":0,"one_line_summary":"Gaussian maxima to Gumbel convergence rates are computed exactly for the Kolmogorov, W1, total variation, KL and Fisher metrics, with explicit constants depending on powers of log log n and log n.","lead":"This paper computes exact asymptotic convergence rates for the maximum of n independent standard Gaussian variables to the Gumbel distribution under five probability metrics. It also shows how the choice of centering and scaling constants changes these rates.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The Fisher-information rate in Theorem 1 is off by a factor of 2: §3.5 computes c_n^4/(64(log n)^2) times ∫ e^{-3x-e^{-x}} dx, but that integral equals 2, so the constant should be (log log n)^4/(512(log n)^2), not /1024.","rationale":"The reader's weakest_assumption concerned uniform validity of the refined Mills-ratio expansions; I do not find a concrete failure there, since the t_n^4/(log n)^2 errors are uniformly small on [-ℓ1, ℓ2] and the tail contributions are exponentially small. The reader did flag the Fisher-number inconsistency in their rationale, and I agree that this is the correct load-bearing issue. I separately verified the central expansion behind Lemma 2.4: the (log n)^{-1} coefficient cancels and the (log n)^{-2} coefficient gives c_n^4/32, so the KL rate in Theorem 1 is supported. The Fisher rate, however, is internally wrong by a factor of 2 because the final integral is evaluated as 1 instead of 2. This is localized but load-bearing: Theorem 1 asserts five exact asymptotic constants, and one false constant invalidates the theorem as stated. The appropriate disposition is conditional acceptance pending correction of the Fisher line and the corresponding proof display; no other asserted rate appears to require change.","tokens_in":15810,"tokens_out":27957,"duration_ms":301519,"concrete_test":"Recompute the decisive integral of §3.5: I_0 = ∫_{-∞}^{∞} e^{-3x-e^{-x}} dx. Substitute u = e^{-x}; then dx = -du/u and I_0 = ∫_0^∞ u^2 e^{-u} du = 2. Re-evaluate (3.8) with h_n(x) = -t_n^2(x)/(4 log n)(1+o(1)), obtaining I = (c_n^4/(64 log^2 n)) I_0 (1+o(1)) = (log log n)^4/(512 log^2 n)(1+o(1)). If the corrected constant appears, the Fisher assertion in Theorem 1 must be changed from denominator 1024 to 512.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"Section 3.5 derives I(L(Yn)|Λ) = 1/4 E[(log(fn/g))'^2] and reduces the leading term to c_n^4/(64(log n)^2) ∫ e^{-3x-e^{-x}} dx. The displayed conversion then treats this integral as 1 and concludes c_n^4/(64(log n)^2) = (log log n)^4/(1024(log n)^2). With u = e^{-x}, e^{-3x-e^{-x}} dx = -u^2 e^{-u} du, so the integral is Γ(3) = 2. The correct value is therefore c_n^4/(32(log n)^2)(1+o(1)) = (log log n)^4/(512(log n)^2)(1+o(1)), matching the KL constant at leading order. The same conclusion follows directly from h_n(x) ≈ -t_n^2(x)/(4 log n) in (3.8): (1/4)E[h_n^2 e^{-2Y_n}] ≈ (1/(64 log^2 n)) ∫ t_n^4 e^{-3x-e^{-x}} dx, whose c_n^4 coefficient is 2/64 = 1/32. Theorem 1 asserts exact constants, so the printed Fisher line is false as stated; the other four rates are not affected by this arithmetic slip. A smaller, non-load-bearing inconsistency also appears in Lemma 2.4, where the statement of E[Y_n] has coefficient γ while the proof obtains (γ+1); that quantity is not used in Theorem 1.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper studies the rate of convergence of the linearly normalized maximum of i.i.d. standard normals, Y_n = a_n(X_(n) - b_n), to the standard Gumbel distribution, with a_n = sqrt(2 log n) and b_n = sqrt(2 log n) - (log log n + log(4 pi))/(2 sqrt(2 log n)). Theorem 1 claims exact asymptotic constants for the Kolmogorov distance, the W1 Wasserstein distance, total variation, Kullback-Leibler divergence, and Fisher information, all of order (log log n)^2/log n or (log log n)^4/log^2 n. Sections 2 and 3 derive these results from Mills-ratio expansions for the tail, distribution, and density of Y_n, together with Gamma-integral evaluations. Section 4 quantifies how the choice of norming constants changes the rates (Theorems 2 and 3). The paper is largely self-contained; the self-citations [9]-[12] supply methodology rather than the target constants.","tokens_in":16154,"tokens_out":15722,"duration_ms":155267,"significance":"The constants in Theorem 1, if correct, would sharpen Hall's 1/log n bounds to exact asymptotics and extend them to strong metrics; the TV, KL, and Fisher information results appear to be new. The derivations are detailed and do not rely on numerical fitting: all constants come from explicit Mills-ratio expansions and Gamma identities, and the self-citations are methodological only, so there is no circularity. I also checked the uniformity concern about Lemmas 2.1 and 2.3 on [-(1/4)log log n, (log n)^{1/4}]: at the endpoints the quoted error terms are o(1), so I do not see a gap there. The central defect is a factor-of-two error in the Fisher information constant, which is localized and repairable.","major_comments":[{"comment":"The Fisher information constant in Theorem 1 is off by a factor of 2. The chain after Eq. (3.8) evaluates the integral J = integral_{-infty}^{infty} e^{-3x-e^{-x}} dx as if it were 1, but with u = e^{-x} one has J = integral_0^infty u^2 e^{-u} du = Gamma(3) = 2. Therefore the displayed conclusion should be I(L(Y_n)|Lambda) = c_n^4/(32 log^2 n) (1+o(1)) = (log log n)^4/(512 log^2 n) (1+o(1)), not (log log n)^4/(1024 log^2 n). Since Theorem 1 asserts exact constants, the fifth line of Theorem 1 is false as written; the other four assertions in Theorem 1 are not affected by this arithmetic slip.","section":"Section 3.5, Eq. (3.8); Theorem 1"}],"minor_comments":[{"comment":"The notation 'c_n = log(sqrt(2 pi a_n))' is ambiguous and appears to mean 'log(sqrt(2 pi) a_n)' = log sqrt(4 pi log n), which is the value used later in Section 3. Please correct the parenthesis placement so that the two definitions agree.","section":"Section 2, after Eq. (2.1)"},{"comment":"The statement of Lemma 2.4 gives E[Y_n] = gamma - c_n^2/(4 log n) + gamma c_n/(2 log n), but the proof concludes with (gamma+1) c_n/(2 log n). This quantity is not used in the proof of Theorem 1, so it does not affect the main results, but the statement and proof should be made consistent.","section":"Lemma 2.4"},{"comment":"In the bound for sup_{x <= -ell_1(n)} |F_n(x) - Lambda(x)|, the second term should be e^{-e^{ell_1(n)}}, not e^{e^{ell_1(n)}}; the following line correctly uses the exponentially decaying expression.","section":"Section 3.1, around Eq. (3.1)"},{"comment":"There are several typographical issues: 'Premililaries' in the Section 2 heading, 'the the scaling' in Section 4, and 'forth section' in the introduction. These should be corrected.","section":"Throughout"}],"recommendation":"major_revision","confidential_remarks":"The paper is within the scope of the journal and the main results are plausible after fixing the arithmetic slip in the Fisher information constant. The self-citation pattern is appropriate and there is no circularity. I recommend major revision rather than rejection because the error is localized and does not undermine the method or the other four rates in Theorem 1."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear colleague,\n\nThe real news here: Ma and Tian compute exact asymptotic constants for five distances between the normalized Gaussian maximum and Gumbel, and four of the five are right. The Fisher-information constant in Theorem 1 is off by a factor of two. In Section 3.5 they reduce the leading term to c_n^4/(64 log^2 n) times the integral of e^{-3x-e^{-x}} dx, and that integral is 2, not 1. So the printed constant (log log n)^4/(1024 log^2 n) should be (log log n)^4/(512 log^2 n), which is exactly their KL constant. The paper's own display immediately before that last step shows the c_n^4/64 coefficient; the missing factor comes from treating the integral as 1. This is load-bearing because Theorem 1 asserts exact constants, but it is localized: the other four rates and the W1/TV/BE arguments do not depend on that integral.\n\nWhat is genuinely new: exact coefficients for W1, TV, KL, and a refined Berry-Esseen constant under the natural normalization a_n = sqrt(2 log n), b_n = sqrt(2 log n) - (log log n + log(4π))/(2 sqrt(2 log n)). The paper also compares other normalizations, reproducing Hall's order and giving explicit constants in Theorems 2–3. That is useful. The derivations are detailed, the Mills-ratio expansions are consistent with Leadbetter's pointwise expansion, and the self-citations are methodological, not circular.\n\nSoft spots beyond the factor-two slip: Lemma 2.4 states E[Y_n] with coefficient γ c_n/(2 log n), but the proof ends with (γ+1)c_n/(2 log n). Minor, since that moment is not used in Theorem 1. The harder question is whether the sixth-order density expansion (2.6) really holds uniformly on [-(1/4)log log n, (log n)^{1/4}] with the claimed O(t_n^6/(log n)^{9/4}) error. The KL and Fisher computations lean on that precision. The tail estimates outside the interval look exponential, so a referee should be able to verify quickly, but it needs checking.\n\nWho this is for: extreme value theorists, and random-matrix people who want clean constants for Gumbel approximations. If the Fisher line is corrected, this is a solid contribution. Even with the error, it deserves a serious referee; this is exactly the kind of fixable arithmetic slip peer review is for.\n\nMy recommendation: send it out. The referee should verify the Fisher integral and the Lemma 2.4 coefficient, and then it should be publishable.\n\nBest,\n[You]","headline":"Exact rates for four of five distances are right and new; the Fisher-information constant in Theorem 1 is off by a factor of two, so the paper needs a small but load-bearing correction before the theorem is quoted.","tokens_in":16670,"tokens_out":3533,"would_cite":true,"duration_ms":36004,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["60G70","60B20","60B10"],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper determines the exact convergence rate of the normalized Gaussian maximum to the Gumbel distribution, giving sharp constants for five distances and showing how the centering choice changes the speed.","keywords":["extreme value theory","Gumbel distribution","normal extremes","Berry-Esseen bound","Wasserstein distance","total variation distance","Kullback-Leibler divergence","Fisher information"],"falsifier":"Recompute the Fisher-information constant from the paper's own Section 3.5: with $h_n(x) \\sim -\\frac{t_n(x)^2}{4\\log n}$ and $\\int_{-\\infty}^{\\infty} e^{-3x - e^{-x}}dx = \\int_0^\\infty t^2 e^{-t}dt = 2$, the displayed integral in (3.8) evaluates to $\\frac{(\\log\\log n)^4}{512\\log^2 n}(1+o(1))$, which disagrees with the $1024$ stated in Theorem 1; a numerical evaluation of all five distances for $n$ up to $10^{15}$ — feasible because $F_n(x) = \\Phi^n(a_n + (x - c_n)/a_n)$ is explicit — would settle which coefficient is right.","tokens_in":15602,"feed_emoji":"📈","tokens_out":36150,"duration_ms":341914,"temperature":0.7,"pith_summary":"This paper establishes the exact speed at which the largest of $n$ independent standard normal draws approaches the Gumbel distribution after classical linear rescaling. With $a_n = \\sqrt{2\\log n}$ and $b_n = \\sqrt{2\\log n} - \\frac{\\log\\log n + \\log(4\\pi)}{2\\sqrt{2\\log n}}$, it derives precise asymptotic constants for five measures of discrepancy: the uniform (Berry–Esseen) bound, the $W_1$ Wasserstein distance, total variation, Kullback–Leibler divergence, and Fisher information. The rates are powers of $\\log\\log n$ over powers of $\\log n$ with explicit coefficients — for instance $(\\log\\log n)^2/(16e\\log n)$ for the uniform bound and $(\\log\\log n)^4/(512\\log^2 n)$ for the Kullback–Leibler divergence. The paper also shows that alternative centering constants produce rates proportional to $1/\\log n$ or $\\log\\log n/\\log n$ with explicit constants, refining the classical order-of-magnitude picture. This matters because the Gaussian maximum is the canonical example in extreme value theory, and the exact distance-by-distance picture with explicit constants is new.","feed_headline":"Gaussian maxima reach Gumbel at (log log n)²/log n","feed_subtitle":"Uniform, Wasserstein, TV, KL, and Fisher distances each get sharp constants for this classical limit.","key_machinery":"The workhorse is a pair of refined Mills-ratio expansions of the Gaussian tail — Lemma 2.1 for the survival function and Lemma 2.3 (with the sixth-order version (2.6)) for the induced density — valid uniformly on the central interval $[-\\ell_1(n), \\ell_2(n)] = [-\\tfrac14\\log\\log n, (\\log n)^{1/4}]$. The interval is chosen so that the leading correction $\\frac{e^{-x}(t_n(x)^2 + 2t_n(x) + 2)}{4\\log n}$ is $O((\\log n)^{-1/4})$, small enough for a single Taylor expansion to capture the rate, while contributions outside the interval are shown to be only $O(e^{-(\\log n)^{1/4}})$. Here $t_n(x) = x - c_n$ with $c_n = \\log\\sqrt{4\\pi\\log n}$, so the dominant coefficient is $c_n^2 \\approx \\tfrac14(\\log\\log n)^2$, which produces the $(\\log\\log n)^2/\\log n$ order. For the distribution-function distances the sharp constants come from three classical integrals: $\\sup_x e^{-e^{-x}-x} = 1/e$, $\\int_{-\\infty}^{\\infty} e^{-e^{-x}-x}dx = 1$, and $\\int_{-\\infty}^{\\infty} e^{-x-e^{-x}}|e^{-x}-1|dx = 2/e$. The Kullback–Leibler and Fisher computations need the sixth-order density expansion because the $1/\\log n$ coefficients integrate to zero; the surviving $(\\log n)^{-2}$ term is proportional to $c_n^4$, giving the $(\\log\\log n)^4/\\log^2 n$ order. A final device is the exact identity $\\mathbb{E}\\log\\Phi(X_{(n)}) = -1/n$, which collapses the Kullback–Leibler computation to the moment calculation of Lemma 2.4.","core_discovery":"The central claim, Theorem 1, is that for i.i.d. standard normal variables and the normalized maximum $Y_n = a_n(X_{(n)} - b_n)$ with $a_n = \\sqrt{2\\log n}$ and $b_n = \\sqrt{2\\log n} - \\frac{\\log\\log n + \\log(4\\pi)}{2\\sqrt{2\\log n}}$, the distribution of $Y_n$ approaches the standard Gumbel law $\\Lambda$ at the exact rates $\\sup_x |P(Y_n \\le x) - \\Lambda(x)| = \\frac{(\\log\\log n)^2}{16e\\log n}(1+o(1))$, $W_1 = \\frac{(\\log\\log n)^2}{16\\log n}(1+o(1))$, total variation $= \\frac{(\\log\\log n)^2}{8e\\log n}(1+o(1))$, Kullback–Leibler $= \\frac{(\\log\\log n)^4}{512\\log^2 n}(1+o(1))$, and Fisher information $= \\frac{(\\log\\log n)^4}{1024\\log^2 n}(1+o(1))$. This upgrades the pointwise distribution expansion known previously at each fixed $x$ (the expression (1.3) from [8]) to a uniform statement with the sharp constant, and extends the refinement to the density, which is what the integral and derivative distances require. Theorems 2 and 3 then compare centering schemes: the centering defined in [5] by $2\\pi b^2 e^{b^2} = n^2$ gives pure $1/\\log n$ rates with explicit constants such as $d_1/(4\\log n)$ with $d_1 \\approx 1.305$, while an intermediate second-order centering gives rates of order $\\log\\log n/\\log n$, showing that the order of convergence is governed by how accurately the centering absorbs the $\\log\\log n$ term.","pith_inferences":["If the sixth-order tail expansion holds for other light-tailed parents in the Gumbel domain of attraction (Weibull-type tails with shape parameter greater than 1, say), the same truncation-and-cancel scheme should give exact constants with the same structure; the paper itself treats only Gaussian tails.","The mechanism behind the $(\\log\\log n)^4/\\log^2 n$ rate for KL and Fisher is that the leading $1/\\log n$ corrections cancel in the relevant expectations; this suggests the general rule that information-style divergences scale as the square of distribution-function distances in smooth parametric limits, a rule the paper does not state.","Reading Theorem 1 against the $d_1/4$ constant of Theorem 2, the classical centering gives a smaller uniform error than the centering of [5] for every sample size below roughly $10^{19}$; only beyond that astronomical size does the asymptotically optimal $1/\\log n$ rate win, a concrete comparison the paper does not draw.","The identity $\\mathbb{E}\\log F(X_{(n)}) = -1/n$ used to collapse the KL computation actually holds for any continuous parent distribution $F$ by the change of variable $u = F(x)$, so the trick transfers verbatim to other extreme-value problems; the paper states it only in the Gaussian setting."],"forward_implications":["With the classical norming constants, the Berry–Esseen constant for Gaussian maxima is exactly $(\\log\\log n)^2/(16e\\log n)$, upgrading the order-of-magnitude bound of [5] to a sharp asymptotic for this normalization.","The five distances separate cleanly: the uniform and total variation constants are $1/e$ and $2/e$ multiples of the $W_1$ scale, while Kullback–Leibler and Fisher information converge faster, at order $(\\log\\log n)^4/\\log^2 n$, because they respond to the squared density correction.","Changing the centering constant changes the order of convergence: the centering of [5] gives $1/\\log n$, the classical centering gives $(\\log\\log n)^2/\\log n$, and an intermediate choice gives $\\log\\log n/\\log n$; no choice beats $1/\\log n$ in the limit, consistent with the sharpness result of [5].","The classical upper-bound constant $c_2 \\le 3/\\log n$ from [5] is replaced by the explicit constant $d_1/4 \\approx 0.326$ for the uniform distance under that centering, with analogous explicit constants $d_3 \\approx 2.6$, $d_4 \\approx 30.8$, and $d_5 \\approx 15.4$ for total variation, Kullback–Leibler, and Fisher information."],"supporting_citations":[{"why":"It supplies the Gamma-function integral identities used in Lemma 2.4 to evaluate the moments that fix the Kullback–Leibler constant.","marker":"[1]"},{"why":"It establishes the sharp 1/log n uniform rate and its optimality; Theorem 1 refines this rate and Theorem 2 benchmarks against its centering choice.","marker":"[5]"},{"why":"It provides the pointwise expansion (1.3) that Lemma 2.2 reproduces and upgrades to a uniform asymptotic.","marker":"[8]"},{"why":"It supplies the truncation scheme (the central interval) that Section 3 follows for the W1 computation.","marker":"[9]"},{"why":"It supplies the Berry–Esseen treatment that Section 3.1 follows for the uniform bound.","marker":"[10]"},{"why":"It is the source of the definitions of the five distances and divergences used in Theorem 1.","marker":"[14]"}],"fun_headline_variants":["Exact rates for Gaussian max to Gumbel law","Sharp constants for five distances to Gumbel","Centering shifts the convergence speed of extremes","Gaussian extremes: precise Gumbel approximation"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the refined Gaussian-tail and density expansions (Lemmas 2.1–2.3, including the sixth-order version (2.6)) are uniformly accurate to the claimed order on the central interval $[-\\tfrac14\\log\\log n, (\\log n)^{1/4}]$ with negligible contributions outside it; if those error terms are actually larger than asserted, the exact constants in Theorem 1 change.","fun_headline_variants_meta":{"raw":{"variants":["Exact rates for Gaussian max to Gumbel law","Sharp constants for five distances to Gumbel","Centering shifts the convergence speed of extremes","Gaussian extremes: precise Gumbel approximation"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000239,"raw_usage":{"total_tokens":1636,"prompt_tokens":1189,"completion_tokens":447,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":805,"completion_tokens_details":{"reasoning_tokens":387}},"tokens_in":805,"tokens_out":447,"duration_ms":5511,"temperature":1.0,"reasoning_tokens":387,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T17:57:30.132191+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Recompute the Fisher-information constant from the paper's own Section 3.5: with $h_n(x) \\sim -\\frac{t_n(x)^2}{4\\log n}$ and $\\int_{-\\infty}^{\\infty} e^{-3x - e^{-x}}dx = \\int_0^\\infty t^2 e^{-t}dt = 2$, the displayed integral in (3.8) evaluates to $\\frac{(\\log\\log n)^4}{512\\log^2 n}(1+o(1))$, which disagrees with the $1024$ stated in Theorem 1; a numerical evaluation of all five distances for $n$ up to $10^{15}$ — feasible because $F_n(x) = \\Phi^n(a_n + (x - c_n)/a_n)$ is explicit — would settle which coefficient is right.","supporting_citations":[{"cited_title":"and Stegun, I","cited_arxiv_id":null,"evidence_quote":"It supplies the Gamma-function integral identities used in Lemma 2.4 to evaluate the moments that fix the Kullback–Leibler constant."},{"cited_title":"(1979) On the rate of convergence of normal extremes","cited_arxiv_id":null,"evidence_quote":"It establishes the sharp 1/log n uniform rate and its optimality; Theorem 1 refines this rate and Theorem 2 benchmarks against its centering choice."},{"cited_title":"R., Lindgren, G., and Rootz´ en, H","cited_arxiv_id":null,"evidence_quote":"It provides the pointwise expansion (1.3) that Lemma 2.2 reproduces and upgrades to a uniform asymptotic."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"It is the source of the definitions of the five distances and divergences used in Theorem 1."}],"review_version":1}