Pith. sign in

REVIEW 1 major objections 4 minor 14 references

Revisit on the convergence rate of normal extremes

T0 review · 1 major / 4 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read The paper determines the exact convergence rate of the normalized Gaussian maximum to the Gumbel distribution, giving sharp constants for five distances and showing how the centering choice changes the speed.

desk verdict Exact rates for four of five distances are right and new; the Fisher-information constant in Theorem 1 is off by a factor of two, so the paper needs a small but load-bearing correction before the theorem is quoted. read the letter →

arxiv 2507.09496 v1 pith:JW7AP6UI submitted 2025-07-13 math.PR

classification math.PR MSC 60G7060B2060B10
keywords extremevaluetheoryGumbeldistributionnormalextremesBerry-EsseenboundWassersteindistancetotalvariationKullback-LeiblerdivergenceFisherinformation
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper establishes the exact speed at which the largest of $n$ independent standard normal draws approaches the Gumbel distribution after classical linear rescaling. With $a_n = \sqrt{2\log n}$ and $b_n = \sqrt{2\log n} - \frac{\log\log n + \log(4\pi)}{2\sqrt{2\log n}}$, it derives precise asymptotic constants for five measures of discrepancy: the uniform (Berry–Esseen) bound, the $W_1$ Wasserstein distance, total variation, Kullback–Leibler divergence, and Fisher information. The rates are powers of $\log\log n$ over powers of $\log n$ with explicit coefficients — for instance $(\log\log n)^2/(16e\log n)$ for the uniform bound and $(\log\log n)^4/(512\log^2 n)$ for the Kullback–Leibler divergence. The paper also shows that alternative centering constants produce rates proportional to $1/\log n$ or $\log\log n/\log n$ with explicit constants, refining the classical order-of-magnitude picture. This matters because the Gaussian maximum is the canonical example in extreme value theory, and the exact distance-by-distance picture with explicit constants is new.

What carries the argument

The workhorse is a pair of refined Mills-ratio expansions of the Gaussian tail — Lemma 2.1 for the survival function and Lemma 2.3 (with the sixth-order version (2.6)) for the induced density — valid uniformly on the central interval $[-\ell_1(n), \ell_2(n)] = [-\tfrac14\log\log n, (\log n)^{1/4}]$. The interval is chosen so that the leading correction $\frac{e^{-x}(t_n(x)^2 + 2t_n(x) + 2)}{4\log n}$ is $O((\log n)^{-1/4})$, small enough for a single Taylor expansion to capture the rate, while contributions outside the interval are shown to be only $O(e^{-(\log n)^{1/4}})$. Here $t_n(x) = x - c_n$ with $c_n = \log\sqrt{4\pi\log n}$, so the dominant coefficient is $c_n^2 \approx \tfrac14(\log\log n)^2$, which produces the $(\log\log n)^2/\log n$ order. For the distribution-function distances the sharp constants come from three classical integrals: $\sup_x e^{-e^{-x}-x} = 1/e$, $\int_{-\infty}^{\infty} e^{-e^{-x}-x}dx = 1$, and $\int_{-\infty}^{\infty} e^{-x-e^{-x}}|e^{-x}-1|dx = 2/e$. The Kullback–Leibler and Fisher computations need the sixth-order density expansion because the $1/\log n$ coefficients integrate to zero; the surviving $(\log n)^{-2}$ term is proportional to $c_n^4$, giving the $(\log\log n)^4/\log^2 n$ order. A final device is the exact identity $\mathbb{E}\log\Phi(X_{(n)}) = -1/n$, which collapses the Kullback–Leibler computation to the moment calculation of Lemma 2.4.

What would settle it

Recompute the Fisher-information constant from the paper's own Section 3.5: with $h_n(x) \sim -\frac{t_n(x)^2}{4\log n}$ and $\int_{-\infty}^{\infty} e^{-3x - e^{-x}}dx = \int_0^\infty t^2 e^{-t}dt = 2$, the displayed integral in (3.8) evaluates to $\frac{(\log\log n)^4}{512\log^2 n}(1+o(1))$, which disagrees with the $1024$ stated in Theorem 1; a numerical evaluation of all five distances for $n$ up to $10^{15}$ — feasible because $F_n(x) = \Phi^n(a_n + (x - c_n)/a_n)$ is explicit — would settle which coefficient is right.

Watch

Extended reading notes

Core claim

The central claim, Theorem 1, is that for i.i.d. standard normal variables and the normalized maximum $Y_n = a_n(X_{(n)} - b_n)$ with $a_n = \sqrt{2\log n}$ and $b_n = \sqrt{2\log n} - \frac{\log\log n + \log(4\pi)}{2\sqrt{2\log n}}$, the distribution of $Y_n$ approaches the standard Gumbel law $\Lambda$ at the exact rates $\sup_x |P(Y_n \le x) - \Lambda(x)| = \frac{(\log\log n)^2}{16e\log n}(1+o(1))$, $W_1 = \frac{(\log\log n)^2}{16\log n}(1+o(1))$, total variation $= \frac{(\log\log n)^2}{8e\log n}(1+o(1))$, Kullback–Leibler $= \frac{(\log\log n)^4}{512\log^2 n}(1+o(1))$, and Fisher information $= \frac{(\log\log n)^4}{1024\log^2 n}(1+o(1))$. This upgrades the pointwise distribution expansion known previously at each fixed $x$ (the expression (1.3) from [8]) to a uniform statement with the sharp constant, and extends the refinement to the density, which is what the integral and derivative distances require. Theorems 2 and 3 then compare centering schemes: the centering defined in [5] by $2\pi b^2 e^{b^2} = n^2$ gives pure $1/\log n$ rates with explicit constants such as $d_1/(4\log n)$ with $d_1 \approx 1.305$, while an intermediate second-order centering gives rates of order $\log\log n/\log n$, showing that the order of convergence is governed by how accurately the centering absorbs the $\log\log n$ term.

Load-bearing premise

The load-bearing premise is that the refined Gaussian-tail and density expansions (Lemmas 2.1–2.3, including the sixth-order version (2.6)) are uniformly accurate to the claimed order on the central interval $[-\tfrac14\log\log n, (\log n)^{1/4}]$ with negligible contributions outside it; if those error terms are actually larger than asserted, the exact constants in Theorem 1 change.

Editorial extensions

If this is right

  • With the classical norming constants, the Berry–Esseen constant for Gaussian maxima is exactly $(\log\log n)^2/(16e\log n)$, upgrading the order-of-magnitude bound of [5] to a sharp asymptotic for this normalization.
  • The five distances separate cleanly: the uniform and total variation constants are $1/e$ and $2/e$ multiples of the $W_1$ scale, while Kullback–Leibler and Fisher information converge faster, at order $(\log\log n)^4/\log^2 n$, because they respond to the squared density correction.
  • Changing the centering constant changes the order of convergence: the centering of [5] gives $1/\log n$, the classical centering gives $(\log\log n)^2/\log n$, and an intermediate choice gives $\log\log n/\log n$; no choice beats $1/\log n$ in the limit, consistent with the sharpness result of [5].
  • The classical upper-bound constant $c_2 \le 3/\log n$ from [5] is replaced by the explicit constant $d_1/4 \approx 0.326$ for the uniform distance under that centering, with analogous explicit constants $d_3 \approx 2.6$, $d_4 \approx 30.8$, and $d_5 \approx 15.4$ for total variation, Kullback–Leibler, and Fisher information.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the sixth-order tail expansion holds for other light-tailed parents in the Gumbel domain of attraction (Weibull-type tails with shape parameter greater than 1, say), the same truncation-and-cancel scheme should give exact constants with the same structure; the paper itself treats only Gaussian tails.
  • The mechanism behind the $(\log\log n)^4/\log^2 n$ rate for KL and Fisher is that the leading $1/\log n$ corrections cancel in the relevant expectations; this suggests the general rule that information-style divergences scale as the square of distribution-function distances in smooth parametric limits, a rule the paper does not state.
  • Reading Theorem 1 against the $d_1/4$ constant of Theorem 2, the classical centering gives a smaller uniform error than the centering of [5] for every sample size below roughly $10^{19}$; only beyond that astronomical size does the asymptotically optimal $1/\log n$ rate win, a concrete comparison the paper does not draw.
  • The identity $\mathbb{E}\log F(X_{(n)}) = -1/n$ used to collapse the KL computation actually holds for any continuous parent distribution $F$ by the change of variable $u = F(x)$, so the trick transfers verbatim to other extreme-value problems; the paper states it only in the Gaussian setting.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

1 major / 4 minor

Summary. This paper studies the rate of convergence of the linearly normalized maximum of i.i.d. standard normals, Y_n = a_n(X_(n) - b_n), to the standard Gumbel distribution, with a_n = sqrt(2 log n) and b_n = sqrt(2 log n) - (log log n + log(4 pi))/(2 sqrt(2 log n)). Theorem 1 claims exact asymptotic constants for the Kolmogorov distance, the W1 Wasserstein distance, total variation, Kullback-Leibler divergence, and Fisher information, all of order (log log n)^2/log n or (log log n)^4/log^2 n. Sections 2 and 3 derive these results from Mills-ratio expansions for the tail, distribution, and density of Y_n, together with Gamma-integral evaluations. Section 4 quantifies how the choice of norming constants changes the rates (Theorems 2 and 3). The paper is largely self-contained; the self-citations [9]-[12] supply methodology rather than the target constants.

Significance. The constants in Theorem 1, if correct, would sharpen Hall's 1/log n bounds to exact asymptotics and extend them to strong metrics; the TV, KL, and Fisher information results appear to be new. The derivations are detailed and do not rely on numerical fitting: all constants come from explicit Mills-ratio expansions and Gamma identities, and the self-citations are methodological only, so there is no circularity. I also checked the uniformity concern about Lemmas 2.1 and 2.3 on [-(1/4)log log n, (log n)^{1/4}]: at the endpoints the quoted error terms are o(1), so I do not see a gap there. The central defect is a factor-of-two error in the Fisher information constant, which is localized and repairable.

major comments (1)
  1. [Section 3.5, Eq. (3.8); Theorem 1] The Fisher information constant in Theorem 1 is off by a factor of 2. The chain after Eq. (3.8) evaluates the integral J = integral_{-infty}^{infty} e^{-3x-e^{-x}} dx as if it were 1, but with u = e^{-x} one has J = integral_0^infty u^2 e^{-u} du = Gamma(3) = 2. Therefore the displayed conclusion should be I(L(Y_n)|Lambda) = c_n^4/(32 log^2 n) (1+o(1)) = (log log n)^4/(512 log^2 n) (1+o(1)), not (log log n)^4/(1024 log^2 n). Since Theorem 1 asserts exact constants, the fifth line of Theorem 1 is false as written; the other four assertions in Theorem 1 are not affected by this arithmetic slip.
minor comments (4)
  1. [Section 2, after Eq. (2.1)] The notation 'c_n = log(sqrt(2 pi a_n))' is ambiguous and appears to mean 'log(sqrt(2 pi) a_n)' = log sqrt(4 pi log n), which is the value used later in Section 3. Please correct the parenthesis placement so that the two definitions agree.
  2. [Lemma 2.4] The statement of Lemma 2.4 gives E[Y_n] = gamma - c_n^2/(4 log n) + gamma c_n/(2 log n), but the proof concludes with (gamma+1) c_n/(2 log n). This quantity is not used in the proof of Theorem 1, so it does not affect the main results, but the statement and proof should be made consistent.
  3. [Section 3.1, around Eq. (3.1)] In the bound for sup_{x <= -ell_1(n)} |F_n(x) - Lambda(x)|, the second term should be e^{-e^{ell_1(n)}}, not e^{e^{ell_1(n)}}; the following line correctly uses the exponentially decaying expression.
  4. [Throughout] There are several typographical issues: 'Premililaries' in the Section 2 heading, 'the the scaling' in Section 4, and 'forth section' in the introduction. These should be corrected.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the convergence-rate constants are derived from explicit Mills-ratio expansions and integral evaluations, not from the claimed conclusions.

full rationale

The paper's claimed rates are obtained by direct asymptotic analysis. Lemmas 2.1–2.3 start from the Gaussian Mills ratio with the stated normalization an = sqrt(2 log n) and bn = sqrt(2 log n) − (log log n + log(4π))/(2 sqrt(2 log n)), expand the resulting tail and density expressions in powers of t_n(x) = x − c_n, and then evaluate the residual integrals against the Gumbel density e^{−x−e^{−x}}. The leading constants (log log n)^2/(16e log n), (log log n)^2/(16 log n), (log log n)^2/(8e log n), and (log log n)^4/(512 log^2 n) emerge from integrals such as ∫ e^{−x−e^{−x}} dx = 1, ∫ e^{−x−e^{−x}}|e^{−x}−1|dx = 2/e, and Γ-function identities; they are not inserted as inputs. The self-citations [9]–[12] are used only to choose the interval cutoffs ℓ1(n), ℓ2(n) and to borrow the proof strategy for the Berry-Esseen bound and W1 distance; they do not supply the target constants or any uniqueness result that forces the conclusions. The discrepancy between the statement and proof of Lemma 2.4 for E[Y_n] (coefficient γ versus γ+1) is not used in Theorem 1, and the KL computation relies on a separately stated and proved expectation E[e^{−Y_n} − t_n^2(Y_n)/(4 log n)]. The factor-of-2 issue in Section 3.5, where ∫ e^{−3x−e^{−x}} dx = 2 rather than 1, would make the printed Fisher constant incorrect, but that is an arithmetic/correctness error, not circularity: the displayed computation does not reduce to its own conclusion by construction. Overall, the derivation is self-contained against standard Mills-ratio expansions and external benchmark results such as Leadbetter et al.'s (1.3).

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

The paper introduces no free parameters and no new entities. It relies on standard analytic facts: the Mills ratio expansion for the normal tail, Taylor expansions with controlled remainders, Gamma integral identities for Gumbel moments, and the assumption of i.i.d. standard Gaussian variables.

assumptions (4)
  • standard math Mills ratio asymptotic expansion of the standard normal tail, Psi(z) = phi(z)/z (1 + O(1/z^2)).
    Invoked in Lemma 2.1 to expand the tail at z = an + (x-cn)/an.
  • standard math Uniform Taylor expansions of exp and log on the interval |x| <= (log n)^{1/4}, with t_n^2(x)/(log n) tending to 0.
    Used repeatedly in Lemmas 2.1-2.3 and Section 3 to extract leading constants.
  • standard math Gamma integral identities such as integral_0^inf t^{s-1} e^{-t} dt = Gamma(s) and log-moment values for Gumbel moments.
    Used in Lemma 2.4 for E[Y_n] and E[e^{-Y_n} - t_n^2(Y_n)/(4 log n)].
  • domain assumption The sample variables X_i are i.i.d. standard Gaussian and the normalization an, bn is the one stated in Theorem 1.
    The entire computation is under this model, stated in the abstract and Section 1.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Revisit on the convergence rate of normal extremes." pith.science (2026). https://pith.science/paper/JW7AP6UI

@misc{pith2026250709496,
  author       = {Pith},
  title        = {Pith review of: Revisit on the convergence rate of normal extremes},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/JW7AP6UI}},
  note         = {Machine review of arXiv:2507.09496}
}
abstract

Let $(X_i)_{1 \le i \le n}$ be independent and identically distributed (i.i.d.) standard Gaussian random variables, and denote by $X_{(n)} = \max_{1 \le i \le n} X_i$ the maximum order statistic. It is well-known in extreme value theory that the linearly normalized maximum $ Y_n = a_n(X_{(n)} - b_n), $ converges weakly to the standard Gumbel distribution $\Lambda$ as $n \to \infty$, where $a_n > 0$ and $b_n$ are appropriate scaling and centering constants. In this note, choosing $$a_n=\sqrt{2\log n}\quad \text{and}\quad b_n = \sqrt{2 \log n} - \frac{\log \log n + \log (4\pi)}{2 \sqrt{2 \log n}},$$ we provide the exact order of this convergence under several distances including Berry-Esseen bound, $W_1$ distance, total variation distance, Kullback-Leibler divergence and Fisher information. We also show how the orders of these convergence are influenced by the choice of $b_n$ and $a_n.$

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

14 extracted references · 10 canonical work pages

  1. [9]

    Y. T. Ma and X. Hu. Universality of convergence rate of rightmost eigenvalue of complex IID random matrices. arXiv:2506.04560, 2025

  2. [12]

    Y. T. Ma and S. Wang. Optimal W1 and Berry-Esseen bound between the spectral radius of large Chiral non-Hermitian random matrices and Gumbel. arXiv:2501.08661, 2025

  3. [1]

    and Stegun, I

    Abramowitz, M. and Stegun, I. A. Handbook of mathematical functions. 2nd ed.(Dover, New York, 1972)

  4. [2]

    Cram´ er, H. (1946). Mathematical Methods of Statistics. Princeton Univ. Press

  5. [3]

    Fisher, R. A. and Tippett, L. H. C. (1928). Limiting forms of the frequency distribution of the largest or smallest member of a sample. Proc. Cambridge Philos. Soc

  6. [4]

    Gnedenko, B. (1943). Sur la distribution limite du terme maximum d’une s´ erie al´ eatoire. Ann. of Math., 44(3), 423-453

  7. [5]

    (1979) On the rate of convergence of normal extremes

    Hall, P. (1979) On the rate of convergence of normal extremes. J. Appl. Probab., 19(2), 433-439

  8. [6]

    Hall, P. (1980). Estimating probabilities for normal extremes. Adv. Appl. Probab., 12, 491-500

Show all 14 references
  1. [7]

    Leadbetter, M. R. (1971). Extreme value theory for stochastic processes. Proc. 5th Princeton Systems Symposium

  2. [8]

    R., Lindgren, G., and Rootz´ en, H

    Leadbetter, M. R., Lindgren, G., and Rootz´ en, H. (1983). Extremes and Related Prop- erties of Random Sequences and Processes. Springer-Verlag

  3. [10]

    Y. T. Ma and X. Meng. Exact convergence rate of spectral radius of complex Ginibre to Gumbel distribution. arXiv:2501.08039, 2025

  4. [11]

    Y. T. Ma and X. Meng. How fast does spectral radius of truncated circular unitary ensemble converge? arXiv:2506.16967, 2025

  5. [13]

    Rootz´ en, H. (1982). The rate of convergence of extremes of stationary normal sequences. Adv. Appl. Probab., 15(1), 54-80

  6. [14]

    C. Villani. Optimal Transport: Old and New. Springer-Verlag, Berlin, 2000. Yutao MA, School of Mathematical Sciences& Laboratory of Mathematics& Complex Systems of Ministry of Education, Beijing Normal University, 100875 Bei- jing, China. Email address: mayt@bnu.edu.cn Bingjie...

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.