REVIEW 2 major objections 4 minor 40 references
Upper Bounds for Local Learning Coefficients of Three-Layer Neural Networks
T0 review · 2 major / 4 minor · reviewed 2026-07-14 · grok-4.5
Pith's one-line read An upper bound for local learning coefficients at singular points of three-layer nets is given by a budget-demand-supply counting rule on the Taylor expansion of the log-likelihood ratio.
desk verdict Solid upper-bound formula for local RLCT at singular points of three-layer nets; tight when N=1, sometimes loose otherwise, with independence caveats already flagged by the author. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
Main Theorem (equation 2.1): a counting rule under budget-demand-supply constraints obtained by successive blow-ups that produce a normal-crossing form of the Kullback-Leibler divergence. The six integers (r, α, β, γ, (m_s), (n_s)) completely determine the bound.
What would settle it
Take any three-layer net with N=1 whose exact learning coefficient is already known (e.g., tanh or exponential activations). Compute the right-hand side of the new bound at P1 or P2; if it differs from the known exact value, the Main Theorem is false for that case.
Extended reading notes
Core claim
Under four explicit conditions on the Taylor expansion of the log-likelihood ratio at a realization parameter P, the local learning coefficient satisfies λ_P ≤ r/2 plus a closed-form expression that counts the maximum number of “items” purchasable under budget β, demand α and successive prices m_s with inventories n*_s. The multiplicity is 2 precisely when the budget is exhausted exactly at a shelf boundary, and 1 otherwise.
Load-bearing premise
The random variables built from activation values, first derivatives times inputs, and higher monomials must be linearly independent almost surely; if that independence fails the normal-crossing analysis and the stated upper bound do not apply.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper derives an upper-bound formula (Main Theorem, Eq. 2.1) for the local learning coefficient (real log canonical threshold) at a class of singular realization parameters of statistical models whose log-likelihood ratio admits a Taylor expansion of a specified form. The bound is expressed combinatorially via quantities (L, n*_s, K) that count the maximum number of items purchasable under budget β, demand α and shelf prices/inventories (m_s, n_s). The formula is obtained by an explicit four-step sequence of blow-ups that produces a normal-crossing form of the Kullback–Leibler divergence. It is then specialized to three-layer neural networks with real-analytic activations, yielding concrete upper bounds (3.2) and (3.4) at two singular strata P1 and P2 for non-polynomial activations (modulo linear-independence hypotheses) and, for polynomial activations, only when the true distribution has no hidden units. When the input dimension is one the numerical values recover previously known exact learning coefficients; for higher input dimension the bounds are consistent with earlier upper bounds of Aoyagi but can be strict (e.g., reduced-rank regression case 1).
Significance. If the Main Theorem and its applications hold, the work supplies the first broadly applicable upper-bound formula for local learning coefficients at singular points of three-layer networks, covering activations such as swish and (under H*=0) polynomials, and thereby extends the exact results of Aoyagi for Vandermonde-type and ReLU singularities as well as the author’s earlier semiregular (nonsingular-point) formula. The budget–demand–supply interpretation and the systematic accounting of how the numbers of weight parameters (r, α, β) and the orders (m_s) enter the coefficient give a transparent geometric picture that is useful for model selection via sBIC and for understanding Bayesian asymptotics of over-parametrized networks. The detailed chart-by-chart blow-up analysis (Appendix C), the genericity lemma for Vandermonde-type Jacobians (Lemma A.1), and the complete worked example (Section 4) constitute solid technical contributions that can be reused for deeper architectures.
major comments (2)
- Abstract and §3.1 claim that the formula “applies in general settings” for non-polynomial analytic activations, yet the load-bearing linear-independence hypotheses (Main Theorem (iii) and concrete conditions (3.1)/(3.3)) can fail even under Assumption 1, as the paper itself records for swish-type activations with certain true weights (Remark 3.2(3)). The abstract and introduction should state the independence requirement with the same prominence given to the H*=0 restriction for polynomials, so that the scope of the upper bounds (3.2) and (3.4) is not overstated.
- §3.2.1 (reduced-rank regression): after the coordinate change the Main Theorem recovers only three of the four cases of the exact learning coefficient of Aoyagi–Watanabe (2005). The missing case (case 1) shows that the inequality in (2.1) can be strict. Remark 2.2 already notes that equality holds when the Jacobian of condition (ii) is nonsingular for every b eq0; a short additional paragraph quantifying how often this occurs for the reduced-rank stratum would clarify when the bound is tight versus merely an upper bound.
minor comments (4)
- Figure 2 caption: “Uppe bound of λ” is missing the letter “r”.
- Notation for multi-indices and the re-indexing of (h,k) into a single index n in Appendix A is dense; a short table summarizing the correspondence between (r,α,β,γ,m_s,n_s) and network dimensions for P1 versus P2 would help the reader.
- In the statement of the Main Theorem the case γ=∞ is handled by a footnote; moving the definition of L into the main text would improve readability.
- Several self-citations to the author’s semiregular papers [19,20] are essential for the nonsingular baseline, but a one-sentence reminder of the precise statement of the earlier formula would make the comparison self-contained.
Circularity Check
No significant circularity: the Main Theorem upper bound is derived from an explicit four-step blow-up sequence under stated analytic and linear-independence hypotheses, not by construction from the target coefficient or a load-bearing self-citation.
full rationale
The derivation chain is self-contained mathematical analysis. The Main Theorem (eq. 2.1) is obtained by performing coordinate transformations (blow-ups CT1–CT9 and the a↦a' change of Step 2) on the Taylor expansion of the log-likelihood ratio f, producing normal-crossing forms of K whose real log canonical thresholds are bounded by the budget–demand–supply expression; the proof outline (Section 5) and full chart-by-chart calculation (Appendix C) do not presuppose the value of λ_P. Conditions (i)–(iv) and the concrete linear-independence hypotheses (3.1)/(3.3) are assumptions under which the bound holds; the paper itself records the cases in which they fail (Remark 3.2(3), H*=0 restriction for polynomials) and the cases in which the inequality is strict (reduced-rank case 1). Self-citations to the author’s semiregular papers [19,20] supply the nonsingular baseline being improved and some preparatory lemmas (e.g., Lemma B.1), but are not used to force the singular-point formula. Agreement with Aoyagi’s exact N=1 results is an external consistency check, not a circular reduction. No fitted parameters, no uniqueness theorem imported from the same authors, and no renaming of a known empirical pattern appear. The single minor self-citation does not raise the score above 1.
Assumptions & free parameters
assumptions (6)
- standard math Hironaka resolution of singularities (Theorem 1.1): real-analytic K admits a proper analytic map to normal-crossing form.
- domain assumption f and K are real-analytic near realization parameters; expectation and θ-derivatives may be interchanged.
- domain assumption Model is realizable: there exists θ* with q = p(·|θ*) a.s.; prior positive at realization parameters.
- ad hoc to paper Main Theorem conditions (i)–(iv): g_{s,n}(a=0,·)=0; Jacobian rank of lowest-degree terms equals ∑ n*_s for some b≠0; linear independence of (Z_{s,n}) and of Fisher score directions; higher terms h_s lie in the ideal generated by the g_{s,n}.
- ad hoc to paper For polynomial activations, the true distribution has no hidden units (H*=0).
- domain assumption Linear independence of activation values, first derivatives times inputs, and monomials of degrees m_s (conditions (3.1), (3.3)); Proposition A.1 gives sufficient conditions.
invented entities (2)
-
Budget–demand–supply counting quantities (L, n*_s, K) and the associated upper-bound formula (2.1)
-
Singular realization strata P1 and P2 for three-layer nets
independent evidence
Cite this review
Pith. "Pith review of Upper Bounds for Local Learning Coefficients of Three-Layer Neural Networks." pith.science (2026). https://pith.science/paper/3TXO7I7C
@misc{pith2026260312785,
author = {Pith},
title = {Pith review of: Upper Bounds for Local Learning Coefficients of Three-Layer Neural Networks},
year = {2026},
howpublished = {\url{https://pith.science/paper/3TXO7I7C}},
note = {Machine review of arXiv:2603.12785}
}
read the original abstract
Three-layer neural networks are known to form singular learning models, and their Bayesian asymptotic behavior is governed by the learning coefficient, or real log canonical threshold. Although this quantity has been clarified for regular models and for some special singular models, broadly applicable methods for evaluating it in neural networks remain limited. Recently, a formula for the local learning coefficient of semiregular models was proposed, yielding an upper bound on the learning coefficient. However, this formula applies only to nonsingular points in the set of realization parameters and cannot be used at singular points. In particular, for three-layer neural networks, the resulting upper bound has been shown to differ substantially from learning coefficient values already known in some cases. In this paper, we derive a formula for an upper bound on local learning coefficients at a class of singular realization parameters in three-layer neural networks. This formula can be interpreted as a counting rule under budget, demand, and supply constraints. In the non-polynomial real-analytic case, the formula applies in general settings, whereas in the polynomial case it applies under the restriction that the true distribution has no hidden units. In particular, our result covers activation functions such as the swish function and also includes polynomial activation functions under the above restriction, thereby extending previous results to a broader class of activation functions. We further show that, when the input dimension is one, the numerical value given by the right-hand side of our upper-bound formula agrees with the previously known learning coefficient, thereby providing a useful comparison with known exact results. Our result also provides a systematic perspective on how the weight parameters of three-layer neural networks affect the learning coefficient.
Reference graph
Works this paper leans on
-
[1]
Akaike, H. (1974). A new look at the statistical model ide ntification. IEEE Transactions on Automatic Control , 19, 716-723
1974
-
[2]
Aoyagi, M. (2006). The zeta function of learning theory a nd generalization error of three layered neural perceptron. RIMS Kokyuroku, Recent Topics on Real and Complex Singularities , 1501, 153-167
2006
-
[3]
Aoyagi, M. (2009). Log canonical threshold of Vandermon de matrix type singularities and generalization error of a three-layered neural network in Bayesian estimation. International Journal of Pure and Applied Mathemat- ics, 52, 177-204
2009
-
[4]
Aoyagi, M. (2010). A Bayesian learning coefficient of gene ralization error and Vandermonde matrix-type singularities. Communications in Statistics—Theory and Methods , 39, 2667-2687
2010
-
[5]
Aoyagi, M. (2013a). Consideration on singularities in l earning theory and the learning coefficient. Entropy, 15, 3714-3733
-
[6]
Aoyagi, M. (2013). Learning coefficient in Bayesian estim ation of restricted Boltzmann machine. Journal of Algebraic Statistics , 4, 30-57
2013
-
[7]
Aoyagi, M. (2019a). Learning coefficient of Vandermonde m atrix-type sin- gularities in model selection. Entropy, 21, 1-12
-
[8]
Aoyagi, M. (2019b). Learning coefficients and informatio n criteria. Fron- tiers in Artificial Intelligence and Applications , 351-362
Show all 40 references
-
[9]
Aoyagi, M. (2024). Consideration on the learning efficien cy of multiple- layered neural networks with linear units. Neural Networks, 172, 106132
2024
-
[10]
Aoyagi, M. (2025). Singular learning coefficients and effi ciency in learning theory. arXiv:2501.12747
2025 arXiv
-
[11]
Aoyagi, M., & Watanabe, S. (2005). Resolution of singul arities and the generalization error with Bayesian estimation for layered neural network. IEICE Transactions J88-D-II , 10, 2112-2124
2005
-
[12]
Aoyagi, M., & Watanabe, S. (2005). Stochastic complexi ties of reduced rank regression in Bayesian estimation. Neural Networks, 18, 924-933
2005
-
[13]
W., & Salamon, D
Robbin, J. W., & Salamon, D. A. (2000). The exponential V andermonde matrix. Linear Algebra and its Applications , 317, 225-226. 23
2000
-
[14]
Drton, M., & Plummer, M. (2017). A Bayesian information criterion for singular models. Journal of the Royal Statistical Society Series B: Statistica l Methodology, 79, 323-380
2017
-
[15]
Drton, M., Lin, S., Weihs, L., & Zwiernik, P. (2017). Mar ginal likelihood and model selection for Gaussian latent tree and forest mode ls. Bernoulli, 23, 1202-1232
2017
-
[16]
Hironaka, H. (1964). Resolution of singularities of an algebraic variety over a field of characteristic zero. Annals of Math , 79, 109-326
1964
-
[17]
Imai, T. (2019). Estimating real log canonical thresho lds. arXiv:1906.01341
2019 arXiv
-
[18]
Kashiwara, M. (1976). B-functions and holonomic syste ms. Inventiones Mathematicae, 38, 33-53
1976
-
[19]
Kurumadani, Y. (2025). Learning coefficients in semireg ular models I: prop- erties. Japanese Journal of Statistics and Data Science , 8, 1051-1079
2025
-
[20]
Kurumadani, Y. (2025). Learning coefficients in semireg ular models II: ex- tensions. Japanese Journal of Statistics and Data Science
2025
-
[21]
Lau, E., Furman, Z., Wang, G., Murfet, D., & Wei, S. (2025 ). The lo- cal learning coefficient: A singularity-aware complexity me asure. In Pro- ceedings of the 28th International Conference on Artificial Inte lligence and Statistics (AISTATS 2025) , Proceedings of Machine Learni...
2025
-
[22]
Mustata, M. (2002). Singularities of pairs via jet sche mes. Journal of the American Mathematical Society , 15, 599-615
2002
-
[23]
Rusakov, D., & Geiger, D. (2002). Asymptotic model sele ction for naive Bayesian networks. In Proceedings of the Eighteenth Conference on Uncer- tainty in Artificial Intelligence , 438-445
2002
-
[24]
Rusakov, D., & Geiger, D. (2005). Asymptotic model sele ction for naive Bayesian networks. Journal of Machine Learning Research , 6, 1-35
2005
-
[25]
Sato, K., & Watanabe, S. (2019). Bayesian generalizati on error of Poisson mixture and simplex Vandermonde matrix type singul arity. arXiv:1912.13289
2019 arXiv
-
[26]
Schwarz, G. (1978). Estimating the dimension of a model . The Annals of Statistics, 6, 461-464
1978
-
[27]
Takeuchi, K. (1976). Distribution of an information st atistic and the crite- rion for the optimal model. Mathematical Science, 153, 12-18
1976
-
[28]
Watanabe, S. (2001a). Algebraic analysis for nonident ifiable learning ma- chines. Neural Computation, 13, 899-933. 24
-
[29]
Watanabe, S. (2001b). Algebraic geometrical methods f or hierarchical learning machines. Neural Networks, 14, 1049-1060
-
[30]
Watanabe, S. (2001c). Algebraic geometry of learning m achines with sin- gularities and their prior distributions. Journal of Japanese Society of Ar- tificial Intelligence , 16, 308-315
-
[31]
Watanabe, S., Yamazaki, K., & Aoyagi, M. (2004). Kullba ck information of normal mixture is not an analytic function. Technical Report of IEICE (in Japanese)
2004
-
[32]
Watanabe, S. (2009). Algebraic geometry and statistical learning theory, vol. 25 . New York, USA: Cambridge University Press
2009
-
[33]
Watanabe, S. (2010). Equations of states in singular st atistical estimation. Neural Networks , 23, 20-34
2010
-
[34]
Watanabe, S. (2013). A widely applicable Bayesian info rmation criterion. Journal of Machine Learning Research , 14, 867-897
2013
-
[35]
Watanabe, S. (2018). Mathematical theory of Bayesian statistics . Boca Ra- ton, FL, USA: CRC Press
2018
-
[36]
Watanabe, T., & Watanabe, S. (2022). Asymptotic behavi or of Bayesian generalization error in multinomial mixtures. arXiv:2203 .06884
2022
-
[37]
Zwiernik, P. (2011). An asymptotic behavior of the marg inal likelihood for general Markov models. Journal of Machine Learning Research , 12, 3283- 3310. Appendix A. V erification of the Assumptions of the Main Theor em In this section, we verify that the three-layer neural net...
2011
-
[38]
, θ ′′ r , a ′′ 1, 1,
We now denote the transformed coordinates θ′′ 1 , . . . , θ ′′ r , a ′′ 1, 1, . . . , a ′′ 1,n ∗ 1 again by θ′ 1, . . . , θ ′ r, a ′ 1, 1, . . . , a ′ 1,n ∗ 1 . Applying CT5, {θ′ j → θ′ 1θ′′ j , a ′ 1,n → θ′ 1a′′ 1,n , b 1 → θ′ 1b′ 1 |2 ≤ j ≤ r, 1 ≤ n ≤ n∗ 1}, we obtain f =θ′m...
-
[39]
The above argument was given for CT5, but the same conclusion holds when CT6 is applied instead
In this case, the multiplicity is m = 2; otherwise, m = 1. The above argument was given for CT5, but the same conclusion holds when CT6 is applied instead. 40 Apply CT7 m2 − m1 times Proceeding as above, we apply the coordinate transformatio n π = {θ′ j → bm2− m1 1 θ′′ j , a ′...
-
[40]
This case will be considered in the next step. Summarizing, among the normal crossings obtained in Step 3- 1, we have inf Q min j h(Q) j + 1 k(Q) j = min { m1r + β 2m1 , m2r + β + (m2 − m1)n∗ 1 2m2 } , and the multiplicity is m = 2 when β = m1n∗ 1, and m = 1 otherwise. We cont...
Reviewed July 14, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.