REVIEW 2 major objections 1 minor 3 cited by
Non-differentiability of ReLU supplies a complete classification of all parameter symmetries for shallow networks.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
Shallow ReLU networks admit a complete classification of parameter symmetries obtained by exploiting ReLU non-differentiability rather than analytic activation assumptions.
T0 review reviewed 2026-07-12 challenge →
load-bearing objection The supplied PDF body is the wrong paper (number theory, 2604.14036); the ReLU symmetry classification cannot be checked at all. the 2 major comments →
A Complete Symmetry Classification of Shallow ReLU Networks
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
Core claim
By exploiting the non-differentiability of ReLU rather than requiring analyticity, every continuous and discrete symmetry of shallow ReLU networks can be listed completely; no further redundancies remain outside the classified families.
What carries the argument
The combinatorial arrangement of ReLU kink hyperplanes and the induced piecewise-linear partition of input space, which generate the algebraic relations that force any two parameter vectors realizing the same function to differ by one of the listed symmetries.
Load-bearing premise
That the constraints coming from kink locations and linear pieces are strong enough to capture every possible continuous family or discrete action, with nothing exotic left over.
What would settle it
An explicit pair of distinct shallow ReLU parameter vectors that induce identical input-output maps yet cannot be transformed into each other by any symmetry appearing in the classification.
If this is right
- The neuromanifold of shallow ReLU nets is now a fully described geometric object rather than an unknown quotient.
- Optimization analyses that depend on the geometry of the neuromanifold become rigorous for ReLU.
- Reverse-engineering and parameter-identifiability questions for shallow ReLU architectures are settled.
- Any claim that two shallow ReLU nets compute the same function can be checked by testing membership in the listed symmetry classes.
Where Pith is reading between the lines
- The same non-smoothness that complicates gradient-based training is precisely what makes a complete symmetry list possible.
- Symmetry-aware reparameterizations or optimizers could eliminate redundant degrees of freedom during training.
- The kink-locus technique may supply a template for attacking deeper ReLU architectures where analytic methods still fail.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript claims a complete classification of parameter-space symmetries for shallow ReLU networks: distinct parameters that realize the same function. It positions the result against earlier reverse-engineering / identifiability literature and against analytic-activation classifications that excluded ReLU, asserting that the non-differentiability (kink structure) of ReLU supplies the combinatorial and algebraic constraints needed for a full classification of continuous and discrete symmetries of the shallow architecture, with consequences for the geometry of the neuromanifold and for optimization.
Significance. A genuine complete classification of shallow ReLU symmetries would be a substantial contribution to the geometric theory of neural networks. It would close a long-standing gap left by analytic-activation results, give a precise description of the neuromanifold for the most widely used activation, and supply a rigorous foundation for claims about how parameter-space symmetries affect optimization dynamics. If the proofs are correct, the work would be of lasting interest to both the theory and optimization communities.
major comments (2)
- The document supplied under the title and arXiv identifier of the ReLU paper is in fact an unrelated number-theory manuscript ("Distribution modulo one of linear recurrent sequences," arXiv:2604.14036). No definitions of the neuromanifold, no statements of the symmetry groups or equivalence relations, and no proofs appear. Consequently the central claim of a complete classification cannot be verified from the provided text.
- Because the body is missing, the load-bearing assertion that kink-locus combinatorics and linear-piece algebra exhaust every continuous and discrete symmetry of the shallow ReLU architecture remains an uncheckable assertion rather than a demonstrated theorem. Without the actual arguments it is impossible to assess completeness or to locate any gaps.
minor comments (1)
- The abstract alone is well written and correctly situates the problem historically, but it cannot substitute for the missing technical development.
Circularity Check
No circularity: body is an unrelated number-theory paper (arXiv:2604.14036), so the claimed ReLU symmetry classification has no inspectable derivation chain at all.
full rationale
The abstract and paper_id claim a complete classification of parameter-space symmetries of shallow ReLU networks obtained by exploiting non-differentiability. The supplied full manuscript text is instead the entirely different paper “Distribution modulo one of linear recurrent sequences” (arXiv:2604.14036). That text contains theorems, corollaries and proofs about fractional parts of linear recurrent sequences (using well-distribution, Morse–Hedlund complexity, and refinements of Dubickas’ arguments). None of those arguments, equations or citations concern neural networks, ReLU, neuromanifolds or parameter symmetries. Consequently there is no derivation chain for the claimed classification that could reduce to its own inputs by construction, by self-citation, by fitted parameters, or by any of the enumerated circular patterns. The abstract alone states a pure classification theorem with no free parameters or self-referential definitions. Score is therefore 0 with empty steps; the mismatch is a content-supply failure, not circularity inside a coherent argument.
Axiom & Free-Parameter Ledger
axioms (3)
- domain assumption Distinct parameter vectors of a neural architecture can realize identical functions (parameter non-identifiability).
- domain assumption Prior complete symmetry classifications require analytic activations and therefore exclude ReLU.
- ad hoc to paper Non-differentiability / kink structure of ReLU supplies enough constraints for a complete shallow classification.
Cite this review
Pith. "Pith review of A Complete Symmetry Classification of Shallow ReLU Networks." pith.science (2026). https://pith.science/paper/ENI3YBMC
@misc{pith2026260414037,
author = {Pith},
title = {Pith review of: A Complete Symmetry Classification of Shallow ReLU Networks},
year = {2026},
howpublished = {\url{https://pith.science/paper/ENI3YBMC}},
note = {Machine review of arXiv:2604.14037}
}
read the original abstract
Parameter space is not function space for neural network architectures. This fact, investigated as early as the 1990s under terms such as ``reverse engineering," or ``parameter identifiability", has led to the natural question of parameter space symmetries\textemdash the study of distinct parameters in neural architectures which realize the same function. Indeed, the quotient space obtained by identifying parameters giving rise to the same function, called the \textit{neuromanifold}, has been shown in some cases to have rich geometric properties, impacting optimization dynamics. Thus far, techniques towards complete classifications have required the analyticity of the activation function, notably excising the important case of ReLU. Here, in contrast, we exploit the non-differentiability of the ReLU activation to provide a complete classification of the symmetries in the shallow case.
Figures
Forward citations
Cited by 3 Pith papers
-
Most ReLU Networks Admit Identifiable Parameters
For ReLU networks with input and hidden widths at least 2, most parameters are identifiable up to symmetry, so the functional dimension equals the parameter count minus the number of hidden neurons.
-
Most ReLU Networks Admit Identifiable Parameters
For ReLU networks with width at least two in input and hidden layers, an open set of parameters is identifiable, implying functional dimension equals parameter count minus hidden neurons.
-
Implicit Bias of SGD in Multivariate ReLU Networks: Effective Width Collapse
Noisy SGD in the mean-field regime forces wide multivariate ReLU networks to an effective width of at most 2P-1, yielding a continuous piecewise-affine predictor whose hyperplanes are non-redundant with respect to the...
Reference graph
Works this paper leans on
-
[1]
M. Astorg and L. Boc Thaler,Dynamics of skew-products tangent to the identity, J. Eur. Math. Soc. (JEMS)��(2026), no. 2, 559–618, DOI 10.4171/jems/1566. MR5031425↑4, 8
-
[2]
Bugeaud,Distribution modulo one and Diophantine approximation, Cambridge Tracts in Math- ematics, vol
Y. Bugeaud,Distribution modulo one and Diophantine approximation, Cambridge Tracts in Math- ematics, vol. 193, Cambridge University Press, Cambridge, 2012. MR2953186↑8, 9, 11
2012
-
[3]
Z. Chen, Z. Ye, and W. Zheng,On a question of Astorg and Boc Thaler(2026). arXiv:2511.21324. ↑4, 5
Pith/arXiv arXiv 2026
-
[4]
Dubickas,Arithmetical properties of powers of algebraic numbers, Bull
A. Dubickas,Arithmetical properties of powers of algebraic numbers, Bull. London Math. Soc.�� (2006), no. 1, 70–80, DOI 10.1112/S0024609305017728. MR2201605↑1, 4 11
-
[5]
,On the distance from a rational power to the nearest integer, J. Number Theory���(2006), no. 1, 222–239, DOI 10.1016/j.jnt.2005.07.004. MR2204744↑3, 4, 8
-
[6]
,Arithmetical properties of linear recurrent sequences, J. Number Theory���(2007), no. 1, 142–150, DOI 10.1016/j.jnt.2006.04.002. MR2287116↑1, 3, 4, 6
-
[7]
A. Dubickas and A. Novikas,Integer parts of powers of rational numbers, Math. Z.���(2005), no. 3, 635–648, DOI 10.1007/s00209-005-0827-4. MR2190349↑4
-
[8]
L. Flatto, J. C. Lagarias, and A. D. Pollington,On the range of fractional parts{�(���) � }, Acta Arith.��(1995), no. 2, 125–147, DOI 10.4064/aa-70-2-125-147. MR1322557↑1
-
[9]
Mend` es France,Les ensembles de B´ esineau, S´ eminaire Delange-Pisot-Poitou (15` eme ann´ ee: 1973/74), Th´ eorie des nombres, Fasc
M. Mend` es France,Les ensembles de B´ esineau, S´ eminaire Delange-Pisot-Poitou (15` eme ann´ ee: 1973/74), Th´ eorie des nombres, Fasc. 1, Secr´ etariat Math., Paris, 1975, pp. Exp. No. 7, 6 (French). MR0412139↑11
1973
-
[10]
Grekos, V
G. Grekos, V. Toma, and J. Tomanov´ a,A note on uniform or Banach density, Ann. Math. Blaise Pascal��(2010), no. 1, 153–163. MR2674656↑5
2010
-
[11]
G. H. Hardy,A problem of Diophantine approximation, J. Indian Math. Soc.��(1919), 162–166. ↑1, 3
1919
-
[12]
Hlawka,Zur formalen Theorie der Gleichverteilung in kompakten Gruppen, Rend
E. Hlawka,Zur formalen Theorie der Gleichverteilung in kompakten Gruppen, Rend. Circ. Mat. Palermo (2)�(1955), 33–47, DOI 10.1007/BF02846027 (German). MR0074489↑5
-
[13]
,Erbliche Eigenschaften in der Theorie der Gleichverteilung, Publ. Math. Debrecen�(1960), 181–186, DOI 10.5486/pmd.1960.7.1-4.16 (German). MR0125103↑5
-
[14]
Lawton,A note on well distributed sequences, Proc
B. Lawton,A note on well distributed sequences, Proc. Amer. Math. Soc.��(1959), 891–893, DOI 10.2307/2033616. MR0109818↑5
doi:10.2307/2033616 1959
-
[15]
Mahler,An unsolved problem on the powers of3�2, J
K. Mahler,An unsolved problem on the powers of3�2, J. Austral. Math. Soc.�(1968), 313–321. MR0227109↑1
1968
-
[16]
M. Morse and G. A. Hedlund,Symbolic Dynamics, Amer. J. Math.��(1938), no. 4, 815–866, DOI 10.2307/2371264. MR1507944↑8
doi:10.2307/2371264 1938
-
[17]
G. M. Petersen,‘Almost convergence’ and uniformly distributed sequences, Quart. J. Math. Oxford Ser. (2)�(1956), 188–191, DOI 10.1093/qmath/7.1.188. MR0095812↑5
-
[18]
Pisot,La r´ epartition modulo 1 et les nombres alg´ ebriques, Ann
C. Pisot,La r´ epartition modulo 1 et les nombres alg´ ebriques, Ann. Scuola Norm. Super. Pisa Cl. Sci. (2)�(1938), no. 3-4, 205–248 (French). MR1556807↑1, 3
1938
-
[19]
,R´ epartition(mod1)des puissances successives des nombres r´ eels, Comment. Math. Helv. ��(1946), 153–160, DOI 10.1007/BF02565954 (French). MR0017744↑3
-
[20]
Schinzel,On the reduced length of a polynomial with real coefficients, Funct
A. Schinzel,On the reduced length of a polynomial with real coefficients, Funct. Approx. Comment. Math.��(2006), 271–306, DOI 10.7169/facm/1229442629. MR2271619↑11
-
[21]
Vijayaraghavan,On the fractional parts of the powers of a number
T. Vijayaraghavan,On the fractional parts of the powers of a number. II, Proc. Cambridge Philos. Soc.��(1941), 349–357. MR0006217↑1, 3
1941
-
[22]
Weyl, ¨Uber die Gleichverteilung von Zahlen mod
H. Weyl, ¨Uber die Gleichverteilung von Zahlen mod. Eins, Math. Ann.��(1916), no. 3, 313–352, DOI 10.1007/BF01475864 (German). MR1511862↑1, 5 (Zhangchi Chen)School of Mathematical Sciences, Key Laboratory of MEA (Ministry of Education) and Shanghai Key Laboratory of PMMP, East China Normal University, Shanghai 200241, China Email address:�����������������...
This paper was first reviewed by grok-4.5 on July 12, 2026.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.