Pith. sign in

REVIEW 4 minor 25 references

This paper proves a sharp, mean-dependent moderate-deviation exponent for Gaussian maxima and derives the critical Sherrington–Kirkpatrick free-energy variance from entropy.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-01 07:38 UTC pith:UORSLCDA

load-bearing objection Theorem 1 is a clean, likely correct resolution of the Ding–Eldan–Zhai question; the SK half is a credible alternative proof of Du–Huang, with its only real fragility sitting in two external critical-window estimates.

arxiv 2607.21392 v2 pith:UORSLCDA submitted 2026-07-23 math.PR

Moderate Deviations for Gaussian Maxima and an Entropy Proof of Critical SK Free Energy Fluctuations

classification math.PR MSC 60G1582B4460F1060K3594A17
keywords Gaussian fieldsmoderate deviationsGaussian maximaSherrington-Kirkpatrick modelfree energy fluctuationsentropyinformation percolationGOE
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

This paper answers a question about how far a Gaussian maximum can run past its expected value. For a centered Gaussian vector with variances at most 1, if the expected maximum lies between α√log N and √(2 log N)−κ√log N, then the probability of exceeding the mean by κ√log N is at most N^{−κ²/(2−α²)+o(1)}; an equicorrelated example shows no larger exponent is possible. The paper also proves that the free energy of the Sherrington–Kirkpatrick spin glass at its critical temperature has variance (1/6) log N + O(1), matching the value predicted from the near-critical cutoff. The proof is organized around entropy: the variance is written as a relative entropy of a planted Gaussian synchronization model, differentiated via the I-MMSE identity, and controlled by information-percolation and critical random-graph estimates. A separate lower bound uses Gaussian convexity at a negative replica parameter together with spherical and GOE eigenvalue estimates.

Core claim

The paper claims two sharp results. First, for a centered Gaussian vector with individual variances at most 1 and expected maximum E m_N, if E m_N is at least α√log N and E m_N + κ√log N ≤ √(2 log N), then P(max ≥ E m_N + κ√log N) ≤ N^{−κ²/(2−α²)+o(1)}. The exponent comes from optimizing the ratio (t²−2)/(t−α)² at t=2/α, and an equicorrelated Gaussian field shows no better exponent is possible. Second, the free energy of the Sherrington–Kirkpatrick model at β_c=1/√2 has variance (1/6) log N + O(1). The upper bound is an entropy calculation: Gaussian convexity bounds the variance by a tilted entropy, which is identified with the relative entropy of a planted Gaussian synchronization model; th

What carries the argument

The load-bearing identity is the concavity of the Gaussian quantile of the maximum's distribution: the map t → Φ^{-1}(P(max X_i ≤ t)) is concave, allowing the quantile at the moderate-deviation level to be interpolated from the median and a high deterministic level t√log N. A union bound plus two-sided Gaussian tail bounds supplies the quantile at the high level, and optimizing h(t)=(t²−2)/(t−α)² over t>√2 at t=2/α yields the exponent 2/(2−α²). For the spin-glass upper bound, the machinery is a chain: Gaussian convexity of the log-partition function bounds variance by an entropy; a diagonal/off-diagonal split reduces it to the off-diagonal entropy; that entropy is identified with the relativ

Load-bearing premise

The spin-glass upper bound rests on a quoted sharp estimate at the critical random-graph window—two fixed vertices must be connected with probability at most order N^{−2/3}—and on an information-percolation comparison, and if either degraded by a polynomial factor the O(1) transfer from the critical window would fail and the variance constant would not close.

What would settle it

Compute the two-point connection probability P(1↔2) in the critical random graph with edge probability 1/N + A N^{−4/3}; the upper-bound chain needs this to be O(N^{−2/3}). If simulation or rigorous estimate showed a slower decay, the entropy derivative bound would exceed O(N^{1/3}) and the variance constant would not close. For the moderate-deviation half, the equicorrelated example already saturates the exponent, so the falsifying observation would be any Gaussian vector satisfying the hypotheses with a strictly larger tail than N^{−κ²/(2−α²)+o(1)}.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • For every Gaussian vector satisfying the stated scale conditions, the moderate-deviation tail is at most N^{−κ²/(2−α²)+o(1)}, and the equicorrelated field shows this exponent is best possible.
  • The Sherrington–Kirkpatrick free-energy variance at β_c=1/√2 is (1/6) log N + O(1), confirming the coefficient predicted by the critical-window cutoff through an independent entropy route.
  • The upper-bound chain shows the variance at criticality is governed by an entropy derivative of size O(N^{1/3}) over the critical window [1−N^{−1/3},1]; the O(1) cost of moving to the critical point relies on two-point connection probabilities O(N^{−2/3}) in the critical random graph.
  • The lower bound exhibits a matching N^{1/2}e^{−N/2} inverse second moment for the normalized SK partition function, which is what produces the 1/6 coefficient when combined with the entropy upper bound.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • If the entropy route extends to the near-critical window, it should recover the full divergence profile −(1/2) log(1−β²/β_c²) as β approaches β_c, not just the critical constant.
  • The analytic optimization used for the moderate-deviation exponent is not Gaussian-specific in shape; a similar tail-sharpening pattern might hold for any process whose upper tail at moderate deviations matches the Gaussian and whose expected maximum sits at the same scale.
  • The critical random graph's N^{−2/3} two-point connection probability is what sets the N^{−1/3} critical window; replacing the graph comparison by another percolation model would probe whether the 1/6 coefficient is universal across mean-field spin glasses.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

0 major / 4 minor

Summary. The paper contains two main results. Theorem 1 gives a moderate-deviation upper bound for the maximum of an N-dimensional centered Gaussian vector with variances at most one: under the assumptions E max >= alpha sqrt(log N) and E max + kappa sqrt(log N) <= sqrt(2 log N), the probability that the maximum exceeds its mean by kappa sqrt(log N) decays at most as N^{-kappa^2/(2-alpha^2)+o(1)}. The exponent is shown to be sharp by an equicorrelated Gaussian field. Theorem 3 establishes Var(F_N(beta_c)) = (1/6) log N + O(1) for the Sherrington--Kirkpatrick free energy at beta_c = 1/sqrt(2). The proof is based on an entropy/tilting argument: the upper bound identifies the variance with a Kullback--Leibler divergence of a Gaussian synchronization model and controls its derivative via the I--MMSE formula, information percolation, and critical random-graph susceptibility; the lower bound combines Gaussian convexity at a negative replica parameter with spherical inverse moments and GOE eigenvalue identities.

Significance. Both theorems are substantial. Theorem 1 answers the question of Ding--Eldan--Zhai with a sharp, explicit exponent and a short proof. Theorem 3 gives a genuinely different route to a recently proved critical SK result: even though Du--Huang already established the asymptotic, the entropy perspective is of independent interest and likely transferable to related models. I checked the key internal computations -- the t* = 2/alpha optimization, the exact second-moment calculation at lambda_0, the diagonal split B_N = beta^2/2 + Btilde_N, the KL representation, and the spherical inverse-moment/contour chain -- and found no internal inconsistency. The argument is not circular: the SK upper bound is anchored at the critical-window endpoint lambda_0 = 1 - N^{-1/3} and the lower bound does not assume the target variance. The main external dependence is [1, Thm 3.6] and [15, Cor 5.3]; the application in Lemma 11 is correct provided those results have exactly the quoted normalization and uniformity.

minor comments (4)
  1. [Section 2, Eq. (11) and Lemma 3] The symbol b_N in Eq. (11) is undefined; it should be m_N. The same typo appears in Lemma 3 when invoking (11).
  2. [Theorem 1] The assumption 'Emax + kappa sqrt(log N) <= sqrt(2) log N' should read 'Emax + kappa sqrt(log N) <= sqrt(2 log N)', as in the abstract. If taken literally as (sqrt(2)) log N, the statement is false and inconsistent with the proof, which requires the shifted maximum to be below t sqrt(log N) for a fixed t > sqrt(2).
  3. [Lemma 11, Eq. (45)] The closing O(1) bound in Lemma 11 is load-bearing and rests entirely on the exact forms of [1, Thm 3.6] and [15, Cor 5.3]. For transparency, please state explicitly the version of [15, Cor 5.3] used, namely the uniform estimate E sum_C |C|^2 = O_A(N^{4/3}) for p = 1/N + A N^{-4/3}, and the exact information-percolation inequality used from [1]. I do not regard this as a gap, but as written the reader cannot verify the absence of extra polynomial factors without consulting the cited papers.
  4. [Lemma 15 / Proposition 4] The soft-edge input from [19] is standard, but the statement 'joint soft edge convergence ... up to a fixed positive deterministic scaling constant' is terse. Since the lower bound requires a positive-probability event involving the first and fourth eigenvalues, it would help to spell out the limiting joint law and the choice of continuity points q_-, q_+, L.

Circularity Check

0 steps flagged

No significant circularity: the derivations are self-contained and rely on independent external theorems, not on their own conclusions.

full rationale

I walked both derivation chains and found no step in which a claimed prediction reduces by construction to an input, nor any load-bearing self-citation. For Theorem 1, the proof uses Ehrhard's concavity, Gaussian concentration around a median, and standard two-sided Mills bounds; the quantile interpolation and the optimization over t are explicit and do not assume the desired tail exponent. For the SK upper bound, the variance is only bounded by the entropy quantity B_N via convexity, and B_N is then computed directly at lambda0 by exact second-moment identities and propagated via the I-MMSE derivative, the Abbe--Boix-Adsera information-percolation inequality, and the Janson--Spencer critical-window susceptibility estimate. These are external, parameter-free results with stated general assumptions; no fitted value from the present paper enters them. For the lower bound, Gaussian convexity at s=-2 reduces the problem to an inverse moment and the already-proved entropy bound; using the entropy upper bound inside the lower bound is a legitimate reuse of an independent lemma, not an assumption of the target variance. The spherical inverse moment is obtained through Haar averaging and GOE soft-edge/dimension-shift identities. I also checked the acknowledgements and references: the paper does not cite itself, and the cited external theorems are not author-uniqueness claims or ansatz-based rescaling imported from the author's own prior work. The only fragility noted by a skeptical reader is the exact normalization of [1, Thm 3.6] and [15, Cor 5.3]; if either had hidden polynomial losses the O(1) transfer in Lemma 11 would fail. But that is a dependence on external inputs, not circularity, and the manuscript itself does not assert those theorems or claim they are proved here.

Axiom & Free-Parameter Ledger

1 free parameters · 9 axioms · 0 invented entities

No new postulated entities: the Gaussian synchronization model is a reformulation of the tilted SK measure, and the entropy B̃_N is derived from the model. The only hand-chosen scale is λ_0 = 1 − N^{−1/3}, dictated by the critical-window balance rather than fit to the target result. All other inputs are external theorems from the cited literature.

free parameters (1)
  • λ_0 (critical-window endpoint 1 − N^{−1/3}) = 1 − N^{-1/3}
    Chosen by hand as the entry point of the critical window. Not fit to data: the exponent 1/3 is forced by the critical-window scaling and by the O(1) balance in Lemma 11 (N^{1/3} derivative cost × N^{−1/3} interval length); the exact O(1) position is absorbed by the final O(1) constant. Listed for transparency because the coefficient (1/12)log N in Proposition 2 is sensitive to this choice: a gap N
axioms (9)
  • standard math Ehrhard's concavity: Φ^{-1}(P(max_{i≤N} X_i ≤ t)) is concave in t (Lemma 1, cited to [12]).
    Used in Lemma 3 to interpolate the Gaussian quantile between the median level m_N and the deterministic level t√(log N); classical theorem.
  • standard math Gaussian concentration around the median: P(|max X_i − m_N| ≥ u) ≤ 2e^{−u²/2} (Lemma 2, cited to [25]).
    Gives the universal constant C_med = √(2π) controlling |E max − m_N|, used in Lemma 3's ratio estimate (16).
  • domain assumption Wei-Kuo Chen's Gaussian convexity: for convex Ψ, K_Ψ(s) = (1/s) log E e^{sΨ(G)} is convex on R (Theorem 4, cited to [5]).
    Load-bearing for both SK bounds: Lemma 5 (upper bound) and Lemma 12 (convexity on R, negative replica parameter s = −2). [5] is a 2026 journal article; statement is used as a black box.
  • standard math I-MMSE formula: (d/dγ) I(U; Y_γ) = ½ E‖U − E[U|Y_γ]‖² (Theorem 5, [14]).
    Converts the derivative of the KL divergence into squared posterior edge correlations in Proposition 3.
  • domain assumption Abbe–Boix-Adserà information-percolation bound: I₂(θ_u;θ_v|Y) ≤ P(u↔v in bond percolation with p_e = I₂(θ_i;θ_j|Y_e)) (Theorem 6, [1, Thm 3.6]).
    The linchpin of Corollary 1: bounds posterior correlations by ER connection probabilities; the sharp O(1) window transfer in Lemma 11 depends on it.
  • domain assumption Janson–Spencer critical-window component-size moments: at p = 1/N + O(N^{−4/3}), E Σ_C |C|² = O(N^{4/3}) ([15, Cor 5.3]).
    Yields P(1↔2) = O(N^{−2/3}) in Lemma 11, converting the N^{1/3} derivative cost into an O(1) loss across the window.
  • domain assumption β-ensemble soft-edge / stochastic Airy spectrum universality: the top eigenvalues of GOE centered at 2 and scaled N^{2/3} converge jointly ([19]).
    Basis of Lemma 15's existence of γ_N with r_N(γ_N) ≥ c N^{2/3}; quoted tersely as joint convergence 'up to a fixed positive deterministic scaling constant'.
  • standard math Mehta integral formula C_{n,a} = (2/a)^{n(n+1)/4} (2π)^{n/2} Π Γ(1+j/2)/Γ(3/2) (eq. (74), [13]).
    Used to evaluate the constant ratio C_{N+1,N}/((N+1)C_{N,N}) ≥ c N^{−1/2} e^{−N/2} in the dimension-shift identity (Lemma 16).
  • standard math Bromwich/contour inversion for the spherical partition function Z^{sph}_N = (Γ(N/2)/(2πi(N/2)^{N/2−1})) ∫ e^{Nz/2} det(zI−W)^{−1/2} dz (Lemma 14, following [3]).
    The contour representation on which the lower-bound estimate (66)–(70) rests; the inversion argument is included in the text.

pith-pipeline@v1.3.0-alltime-deepseek · 20891 in / 49828 out tokens · 419796 ms · 2026-08-01T07:38:07.574129+00:00 · methodology

0 comments
read the original abstract

In this paper, we study two problems concerning Gaussian maxima. First, let $(X_1,\ldots,X_N)$ be a centered Gaussian vector with $\operatorname{Var}(X_i)\leq 1$. Suppose that, for fixed $\alpha\in(0,\sqrt 2)$ and $\kappa>0$, $\mathbb{E}\max_iX_i\geq\alpha\sqrt{\log N}$ and $\mathbb{E}\max_iX_i+\kappa\sqrt{\log N}\leq\sqrt{2\log N}$. We prove that $\mathbb{P}\left(\max_iX_i\geq \mathbb{E}\max_iX_i+\kappa\sqrt{\log N}\right) \leq N^{-\kappa^2/(2-\alpha^2)+o(1)}$. This answers a question of Ding, Eldan and Zhai. The exponent is sharp, as witnessed by an equicorrelated Gaussian field. Second, for the Sherrington--Kirkpatrick model at the critical inverse temperature $\beta_c=1/\sqrt2$, we prove $\operatorname{Var}\bigl(F_N(\beta_c)\bigr)=\frac16\log N+O(1)$. Our argument establishes the variance asymptotics at the critical temperature from an entropy perspective, via a route distinct from that of Du and Huang. For the upper bound, we express the variance as an entropy under exponential tilting and identify this entropy with the Kullback--Leibler divergence of a Gaussian synchronization model. Its derivative is then bounded using the I-MMSE formula, information percolation, and estimates for the susceptibility of the critical Erd\H{o}s--R\'enyi random graph. For the lower bound, we combine Gaussian convexity applied at the replica parameter with an estimate for inverse moments on the sphere and an identity relating GOE eigenvalue densities in consecutive dimensions.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

25 extracted references · 3 linked inside Pith

  1. [1]

    EmmanuelAbbeandEnricBoix-Adser‘a.Aninformation-percolationboundforspinsynchronizationongeneralgraphs. Ann. Appl. Probab., 30(3):1066–1090, 2020

  2. [2]

    Free-energy fluctuations and chaos in the sherrington-kirkpatrick model.Phys

    Timo Aspelmeier. Free-energy fluctuations and chaos in the sherrington-kirkpatrick model.Phys. Rev. Lett., 100(11):117205, 2008

  3. [3]

    Fluctuations of the free energy of the spherical sherrington–kirkpatrick model.J

    Jinho Baik and Ji Oon Lee. Fluctuations of the free energy of the spherical sherrington–kirkpatrick model.J. Stat. Phys., 165(2):185–224, 2016

  4. [4]

    An error bound in the Sudakov–Fernique inequality.arXiv preprint arXiv:math/0510424, 2005

    Sourav Chatterjee. An error bound in the Sudakov–Fernique inequality.arXiv preprint arXiv:math/0510424, 2005

  5. [5]

    A Gaussian convexity for logarithmic moment generating functions with applications in spin glasses

    Wei-Kuo Chen. A Gaussian convexity for logarithmic moment generating functions with applications in spin glasses. Ann. Inst. Henri Poincaré Probab. Stat., 62(1), 2026

  6. [6]

    Order of fluctuations of the free energy in the SK model at critical temperature

    Wei-Kuo Chen and Wai-Kit Lam. Order of fluctuations of the free energy in the SK model at critical temperature. ALEA Lat. Am. J. Probab. Math. Stat., 16(1):809–816, 2019

  7. [7]

    Gaussian approximations and multiplier bootstrap for maxima of sums of high-dimensional random vectors.Ann

    Victor Chernozhukov, Denis Chetverikov, and Kengo Kato. Gaussian approximations and multiplier bootstrap for maxima of sums of high-dimensional random vectors.Ann. Statist., 41(6):2786–2819, 2013

  8. [8]

    Comparison and anti-concentration bounds for maxima of Gaussian random vectors.Probab

    Victor Chernozhukov, Denis Chetverikov, and Kengo Kato. Comparison and anti-concentration bounds for maxima of Gaussian random vectors.Probab. Theory Related Fields, 162(1–2):47–70, 2015

  9. [9]

    Fluctuations for the sherrington–kirkpatrick spin glass model near the critical tem- perature.arXiv preprint arXiv:2603.05636, 2026

    Partha S Dey and Taegu Kang. Fluctuations for the sherrington–kirkpatrick spin glass model near the critical tem- perature.arXiv preprint arXiv:2603.05636, 2026

  10. [10]

    On multiple peaks and moderate deviations for the supremum of a Gaussian field.Ann

    Jian Ding, Ronen Eldan, and Alex Zhai. On multiple peaks and moderate deviations for the supremum of a Gaussian field.Ann. Probab., 43(6):3468–3493, 2015

  11. [11]

    Fluctuations of the Sherrington–Kirkpatrick free energy at critical temperature.arXiv preprint arXiv:2607.02172, 2026

    Hang Du and Brice Huang. Fluctuations of the Sherrington–Kirkpatrick free energy at critical temperature.arXiv preprint arXiv:2607.02172, 2026

  12. [12]

    Symétrisation dans l’espace de Gauss.Math

    Antoine Ehrhard. Symétrisation dans l’espace de Gauss.Math. Scand., 53(2):281–301, 1983

  13. [13]

    Forrester.Log-Gases and Random Matrices, volume 34 ofLondon Mathematical Society Monographs Series

    Peter J. Forrester.Log-Gases and Random Matrices, volume 34 ofLondon Mathematical Society Monographs Series. Princeton University Press, Princeton, NJ, 2010

  14. [14]

    Mutual information and minimum mean-square error in gaussian channels.IEEE Trans

    Dongning Guo, Shlomo Shamai, and Sergio Verd’u. Mutual information and minimum mean-square error in gaussian channels.IEEE Trans. Inform. Theory, 51(4):1261–1282, 2005

  15. [15]

    A point process describing the component sizes in the critical window of the random graph evolution.Combin

    Svante Janson and Joel Spencer. A point process describing the component sizes in the critical window of the random graph evolution.Combin. Probab. Comput., 16(4):631–658, 2007

  16. [16]

    Virasoro.Spin Glass Theory and Beyond: An Introduction to the Replica Method and Its Applications, volume 9 ofWorld Scientific Lecture Notes in Physics

    Marc Mézard, Giorgio Parisi, and Miguel A. Virasoro.Spin Glass Theory and Beyond: An Introduction to the Replica Method and Its Applications, volume 9 ofWorld Scientific Lecture Notes in Physics. World Scientific Publishing Co., Inc., Teaneck, NJ, 1987

  17. [17]

    Springer Monographs in Mathematics

    Dmitry Panchenko.The Sherrington–Kirkpatrick Model. Springer Monographs in Mathematics. Springer, New York, NY, 2013

  18. [18]

    Phase diagram and large deviations in the free energy of mean-field spin-glasses

    Giorgio Parisi and Tommaso Rizzo. Phase diagram and large deviations in the free energy of mean-field spin-glasses. Phys. Rev. B, 79(13):134205, 2009

  19. [19]

    Ramírez, Brian Rider, and Bálint Virág

    José A. Ramírez, Brian Rider, and Bálint Virág. Beta ensembles, stochastic airy spectrum, and a diffusion.J. Amer. Math. Soc., 24(4):919–944, 2011

  20. [20]

    The order of free energy fluctuations in the critical Sherrington–Kirkpatrick model revisited.arXiv preprint arXiv:2606.21360, 2026

    Adrien Schertzer. The order of free energy fluctuations in the critical Sherrington–Kirkpatrick model revisited.arXiv preprint arXiv:2606.21360, 2026

  21. [21]

    Solvable model of a spin-glass.Phys

    David Sherrington and Scott Kirkpatrick. Solvable model of a spin-glass.Phys. Rev. Lett., 35(26):1792–1796, 1975

  22. [22]

    Michel Talagrand.Spin Glasses: A Challenge for Mathematicians: Cavity and Mean Field Models, volume 46 of Ergebnisse der Mathematik und ihrer Grenzgebiete. 3. Folge. A Series of Modern Surveys in Mathematics. Springer, Berlin, Heidelberg, 2003

  23. [23]

    Volume I: Basic Examples, volume 54 ofErgebnisse der Math- ematik und ihrer Grenzgebiete

    Michel Talagrand.Mean Field Models for Spin Glasses. Volume I: Basic Examples, volume 54 ofErgebnisse der Math- ematik und ihrer Grenzgebiete. 3. Folge. A Series of Modern Surveys in Mathematics. Springer, Berlin, Heidelberg, 2011

  24. [24]

    Volume II: Advanced Replica-Symmetry and Low Temperature, volume55ofErgebnisse der Mathematik und ihrer Grenzgebiete

    MichelTalagrand.Mean Field Models for Spin Glasses. Volume II: Advanced Replica-Symmetry and Low Temperature, volume55ofErgebnisse der Mathematik und ihrer Grenzgebiete. 3. Folge. A Series of Modern Surveys in Mathematics. Springer, Berlin, Heidelberg, 2011

  25. [25]

    Cambridge University Press, Cambridge, 2018

    Roman Vershynin.High-dimensional probability, volume 47 ofCambridge Series in Statistical and Probabilistic Math- ematics. Cambridge University Press, Cambridge, 2018. An introduction with applications in data science, With a foreword by Sara van de Geer. School of Mathematical Sciences, Peking University Email address:ymchenmath@math.pku.edu.cn