REVIEW 1 major objections 4 minor 14 references
Revisit on the convergence rate of normal extremes
T0 review · 1 major / 4 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read The paper determines the exact convergence rate of the normalized Gaussian maximum to the Gumbel distribution, giving sharp constants for five distances and showing how the centering choice changes the speed.
desk verdict Exact rates for four of five distances are right and new; the Fisher-information constant in Theorem 1 is off by a factor of two, so the paper needs a small but load-bearing correction before the theorem is quoted. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The workhorse is a pair of refined Mills-ratio expansions of the Gaussian tail — Lemma 2.1 for the survival function and Lemma 2.3 (with the sixth-order version (2.6)) for the induced density — valid uniformly on the central interval $[-\ell_1(n), \ell_2(n)] = [-\tfrac14\log\log n, (\log n)^{1/4}]$. The interval is chosen so that the leading correction $\frac{e^{-x}(t_n(x)^2 + 2t_n(x) + 2)}{4\log n}$ is $O((\log n)^{-1/4})$, small enough for a single Taylor expansion to capture the rate, while contributions outside the interval are shown to be only $O(e^{-(\log n)^{1/4}})$. Here $t_n(x) = x - c_n$ with $c_n = \log\sqrt{4\pi\log n}$, so the dominant coefficient is $c_n^2 \approx \tfrac14(\log\log n)^2$, which produces the $(\log\log n)^2/\log n$ order. For the distribution-function distances the sharp constants come from three classical integrals: $\sup_x e^{-e^{-x}-x} = 1/e$, $\int_{-\infty}^{\infty} e^{-e^{-x}-x}dx = 1$, and $\int_{-\infty}^{\infty} e^{-x-e^{-x}}|e^{-x}-1|dx = 2/e$. The Kullback–Leibler and Fisher computations need the sixth-order density expansion because the $1/\log n$ coefficients integrate to zero; the surviving $(\log n)^{-2}$ term is proportional to $c_n^4$, giving the $(\log\log n)^4/\log^2 n$ order. A final device is the exact identity $\mathbb{E}\log\Phi(X_{(n)}) = -1/n$, which collapses the Kullback–Leibler computation to the moment calculation of Lemma 2.4.
What would settle it
Recompute the Fisher-information constant from the paper's own Section 3.5: with $h_n(x) \sim -\frac{t_n(x)^2}{4\log n}$ and $\int_{-\infty}^{\infty} e^{-3x - e^{-x}}dx = \int_0^\infty t^2 e^{-t}dt = 2$, the displayed integral in (3.8) evaluates to $\frac{(\log\log n)^4}{512\log^2 n}(1+o(1))$, which disagrees with the $1024$ stated in Theorem 1; a numerical evaluation of all five distances for $n$ up to $10^{15}$ — feasible because $F_n(x) = \Phi^n(a_n + (x - c_n)/a_n)$ is explicit — would settle which coefficient is right.
Extended reading notes
Core claim
The central claim, Theorem 1, is that for i.i.d. standard normal variables and the normalized maximum $Y_n = a_n(X_{(n)} - b_n)$ with $a_n = \sqrt{2\log n}$ and $b_n = \sqrt{2\log n} - \frac{\log\log n + \log(4\pi)}{2\sqrt{2\log n}}$, the distribution of $Y_n$ approaches the standard Gumbel law $\Lambda$ at the exact rates $\sup_x |P(Y_n \le x) - \Lambda(x)| = \frac{(\log\log n)^2}{16e\log n}(1+o(1))$, $W_1 = \frac{(\log\log n)^2}{16\log n}(1+o(1))$, total variation $= \frac{(\log\log n)^2}{8e\log n}(1+o(1))$, Kullback–Leibler $= \frac{(\log\log n)^4}{512\log^2 n}(1+o(1))$, and Fisher information $= \frac{(\log\log n)^4}{1024\log^2 n}(1+o(1))$. This upgrades the pointwise distribution expansion known previously at each fixed $x$ (the expression (1.3) from [8]) to a uniform statement with the sharp constant, and extends the refinement to the density, which is what the integral and derivative distances require. Theorems 2 and 3 then compare centering schemes: the centering defined in [5] by $2\pi b^2 e^{b^2} = n^2$ gives pure $1/\log n$ rates with explicit constants such as $d_1/(4\log n)$ with $d_1 \approx 1.305$, while an intermediate second-order centering gives rates of order $\log\log n/\log n$, showing that the order of convergence is governed by how accurately the centering absorbs the $\log\log n$ term.
Load-bearing premise
The load-bearing premise is that the refined Gaussian-tail and density expansions (Lemmas 2.1–2.3, including the sixth-order version (2.6)) are uniformly accurate to the claimed order on the central interval $[-\tfrac14\log\log n, (\log n)^{1/4}]$ with negligible contributions outside it; if those error terms are actually larger than asserted, the exact constants in Theorem 1 change.
Editorial extensions
If this is right
- With the classical norming constants, the Berry–Esseen constant for Gaussian maxima is exactly $(\log\log n)^2/(16e\log n)$, upgrading the order-of-magnitude bound of [5] to a sharp asymptotic for this normalization.
- The five distances separate cleanly: the uniform and total variation constants are $1/e$ and $2/e$ multiples of the $W_1$ scale, while Kullback–Leibler and Fisher information converge faster, at order $(\log\log n)^4/\log^2 n$, because they respond to the squared density correction.
- Changing the centering constant changes the order of convergence: the centering of [5] gives $1/\log n$, the classical centering gives $(\log\log n)^2/\log n$, and an intermediate choice gives $\log\log n/\log n$; no choice beats $1/\log n$ in the limit, consistent with the sharpness result of [5].
- The classical upper-bound constant $c_2 \le 3/\log n$ from [5] is replaced by the explicit constant $d_1/4 \approx 0.326$ for the uniform distance under that centering, with analogous explicit constants $d_3 \approx 2.6$, $d_4 \approx 30.8$, and $d_5 \approx 15.4$ for total variation, Kullback–Leibler, and Fisher information.
Reading between the lines
- If the sixth-order tail expansion holds for other light-tailed parents in the Gumbel domain of attraction (Weibull-type tails with shape parameter greater than 1, say), the same truncation-and-cancel scheme should give exact constants with the same structure; the paper itself treats only Gaussian tails.
- The mechanism behind the $(\log\log n)^4/\log^2 n$ rate for KL and Fisher is that the leading $1/\log n$ corrections cancel in the relevant expectations; this suggests the general rule that information-style divergences scale as the square of distribution-function distances in smooth parametric limits, a rule the paper does not state.
- Reading Theorem 1 against the $d_1/4$ constant of Theorem 2, the classical centering gives a smaller uniform error than the centering of [5] for every sample size below roughly $10^{19}$; only beyond that astronomical size does the asymptotically optimal $1/\log n$ rate win, a concrete comparison the paper does not draw.
- The identity $\mathbb{E}\log F(X_{(n)}) = -1/n$ used to collapse the KL computation actually holds for any continuous parent distribution $F$ by the change of variable $u = F(x)$, so the trick transfers verbatim to other extreme-value problems; the paper states it only in the Gaussian setting.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper studies the rate of convergence of the linearly normalized maximum of i.i.d. standard normals, Y_n = a_n(X_(n) - b_n), to the standard Gumbel distribution, with a_n = sqrt(2 log n) and b_n = sqrt(2 log n) - (log log n + log(4 pi))/(2 sqrt(2 log n)). Theorem 1 claims exact asymptotic constants for the Kolmogorov distance, the W1 Wasserstein distance, total variation, Kullback-Leibler divergence, and Fisher information, all of order (log log n)^2/log n or (log log n)^4/log^2 n. Sections 2 and 3 derive these results from Mills-ratio expansions for the tail, distribution, and density of Y_n, together with Gamma-integral evaluations. Section 4 quantifies how the choice of norming constants changes the rates (Theorems 2 and 3). The paper is largely self-contained; the self-citations [9]-[12] supply methodology rather than the target constants.
Significance. The constants in Theorem 1, if correct, would sharpen Hall's 1/log n bounds to exact asymptotics and extend them to strong metrics; the TV, KL, and Fisher information results appear to be new. The derivations are detailed and do not rely on numerical fitting: all constants come from explicit Mills-ratio expansions and Gamma identities, and the self-citations are methodological only, so there is no circularity. I also checked the uniformity concern about Lemmas 2.1 and 2.3 on [-(1/4)log log n, (log n)^{1/4}]: at the endpoints the quoted error terms are o(1), so I do not see a gap there. The central defect is a factor-of-two error in the Fisher information constant, which is localized and repairable.
major comments (1)
- [Section 3.5, Eq. (3.8); Theorem 1] The Fisher information constant in Theorem 1 is off by a factor of 2. The chain after Eq. (3.8) evaluates the integral J = integral_{-infty}^{infty} e^{-3x-e^{-x}} dx as if it were 1, but with u = e^{-x} one has J = integral_0^infty u^2 e^{-u} du = Gamma(3) = 2. Therefore the displayed conclusion should be I(L(Y_n)|Lambda) = c_n^4/(32 log^2 n) (1+o(1)) = (log log n)^4/(512 log^2 n) (1+o(1)), not (log log n)^4/(1024 log^2 n). Since Theorem 1 asserts exact constants, the fifth line of Theorem 1 is false as written; the other four assertions in Theorem 1 are not affected by this arithmetic slip.
minor comments (4)
- [Section 2, after Eq. (2.1)] The notation 'c_n = log(sqrt(2 pi a_n))' is ambiguous and appears to mean 'log(sqrt(2 pi) a_n)' = log sqrt(4 pi log n), which is the value used later in Section 3. Please correct the parenthesis placement so that the two definitions agree.
- [Lemma 2.4] The statement of Lemma 2.4 gives E[Y_n] = gamma - c_n^2/(4 log n) + gamma c_n/(2 log n), but the proof concludes with (gamma+1) c_n/(2 log n). This quantity is not used in the proof of Theorem 1, so it does not affect the main results, but the statement and proof should be made consistent.
- [Section 3.1, around Eq. (3.1)] In the bound for sup_{x <= -ell_1(n)} |F_n(x) - Lambda(x)|, the second term should be e^{-e^{ell_1(n)}}, not e^{e^{ell_1(n)}}; the following line correctly uses the exponentially decaying expression.
- [Throughout] There are several typographical issues: 'Premililaries' in the Section 2 heading, 'the the scaling' in Section 4, and 'forth section' in the introduction. These should be corrected.
Circularity Check
No significant circularity: the convergence-rate constants are derived from explicit Mills-ratio expansions and integral evaluations, not from the claimed conclusions.
full rationale
The paper's claimed rates are obtained by direct asymptotic analysis. Lemmas 2.1–2.3 start from the Gaussian Mills ratio with the stated normalization an = sqrt(2 log n) and bn = sqrt(2 log n) − (log log n + log(4π))/(2 sqrt(2 log n)), expand the resulting tail and density expressions in powers of t_n(x) = x − c_n, and then evaluate the residual integrals against the Gumbel density e^{−x−e^{−x}}. The leading constants (log log n)^2/(16e log n), (log log n)^2/(16 log n), (log log n)^2/(8e log n), and (log log n)^4/(512 log^2 n) emerge from integrals such as ∫ e^{−x−e^{−x}} dx = 1, ∫ e^{−x−e^{−x}}|e^{−x}−1|dx = 2/e, and Γ-function identities; they are not inserted as inputs. The self-citations [9]–[12] are used only to choose the interval cutoffs ℓ1(n), ℓ2(n) and to borrow the proof strategy for the Berry-Esseen bound and W1 distance; they do not supply the target constants or any uniqueness result that forces the conclusions. The discrepancy between the statement and proof of Lemma 2.4 for E[Y_n] (coefficient γ versus γ+1) is not used in Theorem 1, and the KL computation relies on a separately stated and proved expectation E[e^{−Y_n} − t_n^2(Y_n)/(4 log n)]. The factor-of-2 issue in Section 3.5, where ∫ e^{−3x−e^{−x}} dx = 2 rather than 1, would make the printed Fisher constant incorrect, but that is an arithmetic/correctness error, not circularity: the displayed computation does not reduce to its own conclusion by construction. Overall, the derivation is self-contained against standard Mills-ratio expansions and external benchmark results such as Leadbetter et al.'s (1.3).
Assumptions & free parameters
assumptions (4)
- standard math Mills ratio asymptotic expansion of the standard normal tail, Psi(z) = phi(z)/z (1 + O(1/z^2)).
- standard math Uniform Taylor expansions of exp and log on the interval |x| <= (log n)^{1/4}, with t_n^2(x)/(log n) tending to 0.
- standard math Gamma integral identities such as integral_0^inf t^{s-1} e^{-t} dt = Gamma(s) and log-moment values for Gumbel moments.
- domain assumption The sample variables X_i are i.i.d. standard Gaussian and the normalization an, bn is the one stated in Theorem 1.
Cite this review
Pith. "Pith review of Revisit on the convergence rate of normal extremes." pith.science (2026). https://pith.science/paper/JW7AP6UI
@misc{pith2026250709496,
author = {Pith},
title = {Pith review of: Revisit on the convergence rate of normal extremes},
year = {2026},
howpublished = {\url{https://pith.science/paper/JW7AP6UI}},
note = {Machine review of arXiv:2507.09496}
}
abstract
Let $(X_i)_{1 \le i \le n}$ be independent and identically distributed (i.i.d.) standard Gaussian random variables, and denote by $X_{(n)} = \max_{1 \le i \le n} X_i$ the maximum order statistic. It is well-known in extreme value theory that the linearly normalized maximum $ Y_n = a_n(X_{(n)} - b_n), $ converges weakly to the standard Gumbel distribution $\Lambda$ as $n \to \infty$, where $a_n > 0$ and $b_n$ are appropriate scaling and centering constants. In this note, choosing $$a_n=\sqrt{2\log n}\quad \text{and}\quad b_n = \sqrt{2 \log n} - \frac{\log \log n + \log (4\pi)}{2 \sqrt{2 \log n}},$$ we provide the exact order of this convergence under several distances including Berry-Esseen bound, $W_1$ distance, total variation distance, Kullback-Leibler divergence and Fisher information. We also show how the orders of these convergence are influenced by the choice of $b_n$ and $a_n.$
Reference graph
Works this paper leans on
-
[9]
Y. T. Ma and X. Hu. Universality of convergence rate of rightmost eigenvalue of complex IID random matrices. arXiv:2506.04560, 2025
arXiv 2025
-
[12]
Y. T. Ma and S. Wang. Optimal W1 and Berry-Esseen bound between the spectral radius of large Chiral non-Hermitian random matrices and Gumbel. arXiv:2501.08661, 2025
arXiv 2025
-
[1]
Abramowitz, M. and Stegun, I. A. Handbook of mathematical functions. 2nd ed.(Dover, New York, 1972)
work page 1972
-
[2]
Cram´ er, H. (1946). Mathematical Methods of Statistics. Princeton Univ. Press
work page 1946
-
[3]
Fisher, R. A. and Tippett, L. H. C. (1928). Limiting forms of the frequency distribution of the largest or smallest member of a sample. Proc. Cambridge Philos. Soc
work page 1928
-
[4]
Gnedenko, B. (1943). Sur la distribution limite du terme maximum d’une s´ erie al´ eatoire. Ann. of Math., 44(3), 423-453
work page 1943
-
[5]
(1979) On the rate of convergence of normal extremes
Hall, P. (1979) On the rate of convergence of normal extremes. J. Appl. Probab., 19(2), 433-439
work page 1979
-
[6]
Hall, P. (1980). Estimating probabilities for normal extremes. Adv. Appl. Probab., 12, 491-500
work page 1980
Show all 14 references
-
[7]
Leadbetter, M. R. (1971). Extreme value theory for stochastic processes. Proc. 5th Princeton Systems Symposium
1971
-
[8]
R., Lindgren, G., and Rootz´ en, H
Leadbetter, M. R., Lindgren, G., and Rootz´ en, H. (1983). Extremes and Related Prop- erties of Random Sequences and Processes. Springer-Verlag
1983
-
[10]
Y. T. Ma and X. Meng. Exact convergence rate of spectral radius of complex Ginibre to Gumbel distribution. arXiv:2501.08039, 2025
2025 arXiv
-
[11]
Y. T. Ma and X. Meng. How fast does spectral radius of truncated circular unitary ensemble converge? arXiv:2506.16967, 2025
2025 arXiv
-
[13]
Rootz´ en, H. (1982). The rate of convergence of extremes of stationary normal sequences. Adv. Appl. Probab., 15(1), 54-80
1982
-
[14]
C. Villani. Optimal Transport: Old and New. Springer-Verlag, Berlin, 2000. Yutao MA, School of Mathematical Sciences& Laboratory of Mathematics& Complex Systems of Ministry of Education, Beijing Normal University, 100875 Bei- jing, China. Email address: mayt@bnu.edu.cn Bingjie...
2000
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.