{"id":"d044e6db-d3a6-4f84-8bc2-095e32bbcd28","arxiv_id":"2501.18915","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"The paper argues that algebraic geometry offers a powerful dictionary for understanding deep learning models with polynomial or piecewise-polynomial activations.","lead":"This position paper proposes \"neuroalgebraic geometry\", a research program that studies neural network function spaces using algebraic geometry. It reviews how invariants like dimension, degree, and singularities of these spaces connect to sample complexity, expressivity, and training dynamics.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The Section 5.1 transfer from polynomial to continuous activations is sketched; Hausdorff closeness of neuromanifolds does not by itself transfer degree, singularities, or other algebro-geometric invariants central to the dictionary.","rationale":"The reader correctly identifies Section 5.1 as the weakest link, and I agree with the conditional verdict. My refinement is that the missing piece is not only the Hausdorff approximation lemma but also the stability of the dictionary's invariants under that approximation. Degree, singularities, and Euclidean distance degree are not continuous under Hausdorff convergence; only metric quantities like covering numbers pass directly. Therefore the abstract's broad claim that algebraic models 'can approximate arbitrary neuromanifolds' and the implied transfer of geometric control is not established. This does not undermine the core claim about algebraic models themselves: for those, the cited results on determinantal varieties, fiber-dimension theorem, and covering number bounds provide independent support. The paper is a position piece, so the appropriate condition is either a proof of the transfer lemma or a careful restriction of the transfer claim to metric invariants. Since this is exactly the condition the reader already imposes, the verdict remains unchanged.","tokens_in":15670,"tokens_out":16731,"duration_ms":171620,"concrete_test":"Two-part check. First, prove the missing estimate in Section 5.1: for a fixed MLP architecture, show that ||F_sigma(w) - F_p(w)||_infty <= C_arch ||sigma - p||_{C(K)} for all w in the compact parameter box, with an explicit C_arch depending on depth and weight bounds; if the bound cannot be established, the Hausdorff transfer fails. Second, take sigma = ReLU and p_n a sequence of polynomials converging uniformly on a fixed interval, compute the Zariski closure, degree, and singular locus of M_{p_n} for a one-hidden-unit network; if these invariants do not stabilize to a well-defined limit matching the ReLU geometry, restrict the claimed extension to metric invariants such as covering numbers.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central dictionary is well supported for algebraic models, but the paper's claim to extend beyond the algebraic domain rests on the one-paragraph argument in Section 5.1. There, uniform approximation of a continuous activation sigma by polynomials p on a compact interval is asserted to yield Hausdorff approximation of the neuromanifold M_sigma by M_p. This is plausible but not proved; it requires a uniform continuity estimate for the map sigma -> phi_sigma over the compact parameter space, with constants depending on depth and weight bounds. More importantly, even if d_H(M_sigma, M_p) -> 0, the algebro-geometric invariants featured in the dictionary — degree, singularities, Euclidean distance degree, data discriminants — are not continuous under Hausdorff convergence. Replacing ReLU by a polynomial approximant can change the singularity structure and the degree of the Zariski closure, so the approximation route cannot transfer those invariants to non-algebraic models. The paper explicitly transfers covering numbers, which are metric and stable, but the abstract and Section 1.1 promise more: 'approximate arbitrary neuromanifolds' and extension of 'results and techniques'. This overreach is the weakest load-bearing point.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This position paper proposes the term 'neuroalgebraic geometry' for the study of neuromanifolds of algebraic machine-learning models, i.e., models whose parameterization is polynomial in parameters and inputs, so that the resulting function space is a semi-algebraic variety. The paper sets out a dictionary connecting invariants such as dimension, degree, covering number, singularities, fibers, and critical points to expressivity, sample complexity, implicit bias, identifiability, and training dynamics. It reviews a substantial body of recent work, including tensor/rank models, linear and polynomial networks, CNNs, attention mechanisms, and determinantal varieties, and illustrates the framework with an explicit two-layer linear network in Appendix A and a sample-complexity argument in Appendix B. It also contains a short section on extending the framework beyond polynomial activations via polynomial approximation and via tropical geometry for ReLU networks.","tokens_in":15888,"tokens_out":7479,"duration_ms":71343,"significance":"The paper's core dictionary is plausible and well supported by the cited literature, and the concrete computations in Appendices A and B are useful illustrations for a position paper. The framing has real value for the community: it collects scattered results under one proposed research program and identifies open problems. The paper does not claim to prove new theorems about the dictionary itself, which is appropriate for an invitation. However, the paper's breadth claim, namely that algebraic models can approximate arbitrary neuromanifolds and hence that algebraic tools extend to continuous activations, rests on the one-paragraph argument in Section 5.1. That argument is not rigorous and, more seriously, cannot transfer the non-metric invariants such as degree, singularities, Euclidean distance degree, and data discriminants merely by Hausdorff approximation. This is the paper's weakest load-bearing point and needs to be fixed or explicitly re-scoped.","major_comments":[{"comment":"The central transfer claim is not established. The paragraph beginning 'More precisely, consider the example...' asserts that uniform approximation of a continuous activation sigma by polynomials on a compact interval yields an approximation of the neuromanifold M_sigma by the neuromanifold of an algebraic model in Hausdorff distance. This requires a uniform continuity estimate for the map sigma -> phi_sigma over the compact parameter space, with constants that do not blow up with depth or weight bounds; no such estimate or proof is given. More importantly, even if d_H(M_sigma, M_p) -> 0, the invariants featured in the dictionary, such as degree, singularities, Euclidean distance degree, and data discriminants, are not continuous under Hausdorff convergence: a small polynomial perturbation can smooth a cusp or change the degree of the Zariski closure. Thus the abstract's promise to 'approximate arbitrary neuromanifolds' and to extend 'results and techniques' is only justified for metric quantities such as covering numbers, as in the Zhang-Kileel example cited at the end of the section. The authors should either prove the transfer with explicit hypotheses or explicitly restrict the extension claim to metric and covering-number results.","section":"Section 5.1 (and Section 1.1)"},{"comment":"The sentence 'algebraic models are not only general, but can approximate arbitrary neuromanifolds' conflates two different statements. The Weierstrass theorem cited in Section 5.1 approximates individual continuous functions by polynomials; it does not, by itself, approximate the image of a parameterization map, which is a set of functions indexed by parameters. Since this sentence is used in Section 2.1 to contrast neuroalgebraic geometry with kernel methods ('which can approximate arbitrary neuromanifolds'), the overstatement is load-bearing for the paper's framing. If Section 5.1 is not made rigorous, this sentence and the corresponding abstract-level promise should be weakened to say that certain metric properties of neuromanifolds of continuous activations can be bounded using algebraic approximations.","section":"Section 1.1 (and Section 2.1)"}],"minor_comments":[{"comment":"In the first sentence of Section 2, 'relvance' should be 'relevance'.","section":"Section 2"},{"comment":"The metric in Theorem B.1 is denoted by d, while Section 4.1 uses d for the degree of a variety; this notational collision makes equations such as (13) harder to read. Consider denoting the metric by rho or another symbol.","section":"Appendix B"},{"comment":"Calling the degree 'an algebraic measure of how curved' the variety is imprecise; degree is an intersection-theoretic invariant and is not a curvature measure in any metric sense.","section":"Section 4.1"},{"comment":"The citation for the Weierstrass Approximation Theorem (de la Cerda, 2023) is an unusual source; a standard analysis textbook or a classical reference would be more helpful for the intended interdisciplinary audience.","section":"Section 5.1"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is essentially a survey and framing paper built on a series of recent works, several by the same authors. The editor may want to ask the authors to clarify in the introduction or in a related-work section which entries of the dictionary are new in this paper versus already established in Kileel et al. (2019), Trager et al. (2020), Kohn et al. (2022), Shahverdi et al. (2025a,b), and Henry et al. (2025). This is not a flaw per se for a position paper, but it affects the novelty claim. The main technical weakness is the approximation transfer in Section 5.1; if that remains a sketch, the paper should be re-scoped so that the central promise is about algebraic models, with only metric extensions claimed for continuous activations."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThis is a position paper and literature review proposing \"neuroalgebraic geometry\" as a research program: studying neuromanifolds of algebraic models (polynomial and ReLU networks, certain CNNs, linear attention) via algebraic geometry, with a dictionary linking dimension, degree, singularities, fibers, and critical points to expressivity, sample complexity, implicit bias, and training dynamics. It is exactly what it says on the tin: an invitation, not a new theorem paper. The dictionary (Table 1) is genuinely useful as an organizing device, and the survey of prior work—much of it the authors' own—is accurate and well-written. The worked example in Appendix A (linear two-layer network as a determinantal variety) is a nice pedagogical anchor, and Appendix B gives a clean geometric intuition for why dimension and degree bound sample complexity, with a standard covering-number argument.\n\nWhat is new is the synthesis and the terminology, not the underlying results. That is fine for a position paper, but it is worth saying plainly: if you are looking for a new theorem, you will not find it here.\n\nThe soft spot is Section 5.1, where the paper claims that continuous activations can be approximated by polynomials, giving Hausdorff approximation of the corresponding neuromanifolds, and that this \"allows for the extension of results and techniques\" to non-algebraic models. That is too strong. The one-paragraph argument gives no uniform continuity estimate for the map sigma -> phi_sigma over the parameter space, and even if Hausdorff distance goes to zero, algebro-geometric invariants—degree, singularities, Euclidean distance degree, data discriminants—are not continuous under Hausdorff convergence. Replacing ReLU with a polynomial can change the singularity structure and the Zariski closure entirely. The paper is careful when it uses the approximation only to transfer covering numbers, a metric quantity, citing Zhang & Kileel, but the abstract and Section 1.1 promise more. This overreach is the weakest load-bearing point and should be fixed by either proving the needed estimates or explicitly limiting the transfer claims to metric invariants.\n\nMinor: the self-citation density is high, but in a review of a field the authors largely created, that is not a flaw. The coverage of tropical geometry as an alternative for ReLU is a useful supplement.\n\nWho is this for? Researchers in mathematical deep learning who want a map of the terrain, and algebraic geometers looking for ML applications. It deserves a serious referee: the synthesis is valuable, the dictionary is likely to be cited, and the flaws are repairable. I would recommend accept with major revision, specifically to address Section 5.1's overreach.","headline":"A useful, well-written position paper proposing 'neuroalgebraic geometry' as a research program; the main weakness is an overreaching approximation claim in Section 5.1.","tokens_in":16442,"tokens_out":1556,"would_cite":true,"duration_ms":60441,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["14P10","68T07","62R01"],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that function spaces of polynomial-activation networks are semi-algebraic varieties, and that their algebro-geometric invariants control core aspects of learning.","keywords":["neuroalgebraic geometry","neuromanifold","semi-algebraic variety","algebraic geometry","sample complexity","implicit bias","singularities","metric algebraic geometry"],"falsifier":"Fix a small MLP architecture with a smooth non-polynomial activation, take compact parameter and input domains, and compute the Hausdorff distance between the true neuromanifold and the neuromanifold obtained by replacing the activation with its degree-$d$ polynomial approximation; if this distance does not go to zero as $d$ grows, the paper's bridge from algebraic models to continuous networks breaks.","tokens_in":15455,"feed_emoji":"🧮","tokens_out":8737,"duration_ms":77576,"temperature":0.7,"pith_summary":"This paper proposes that the function spaces parameterized by machine learning models—the neuromanifolds—should be studied with algebraic geometry rather than differential geometry alone. Its central claim is that for algebraic models, meaning networks with polynomial or piecewise-polynomial activation functions, the neuromanifold is a semi-algebraic variety. It then argues that algebro-geometric invariants of that variety, such as dimension, degree, singularities, fibers, and critical points, control sample complexity, expressivity, training dynamics, and implicit bias. If the claim is right, a precise dictionary becomes available: architecture choices translate into computable geometric invariants, and learning becomes a distance problem over these varieties. The paper offers this dictionary as an invitation to a research direction it calls neuroalgebraic geometry.","feed_headline":"Polynomial neural networks are algebraic varieties","feed_subtitle":"A proposed dictionary maps the geometry of these function spaces to sample complexity, expressivity, and training dynamics.","key_machinery":"The central object is the neuromanifold $\\mathcal{M}=\\{f_w : w\\in\\mathcal{W}\\}$, the image of the parameterization map $\\varphi:\\mathcal{W}\\to\\mathcal{V}$ inside a finite-dimensional ambient space of functions $\\mathcal{V}$. For algebraic models $\\varphi$ is polynomial, so $\\mathcal{M}$ is a semi-algebraic variety, meaning a set carved out by polynomial equalities and inequalities. The machinery is the pair $(\\mathcal{M},\\varphi)$: intrinsic invariants of the variety (dimension, degree, singularities) plus the map's fibers and critical points. Metric algebraic geometry then supplies the distance-based tools—covering-number bounds from dimension and degree, Voronoi cells around singularities, Euclidean distance degree, and data discriminants—that turn learning questions into concrete geometric computations.","core_discovery":"For a parametric model whose parameterization map $\\varphi: \\mathcal{W} \\to \\mathcal{V}$ is polynomial in both parameters and inputs, the image of $\\varphi$—the neuromanifold $\\mathcal{M}=\\{f_w : w\\in\\mathcal{W}\\}$—is a semi-algebraic variety by the Tarski–Seidenberg theorem. The paper argues that the invariants of this variety are machine-learning-relevant: dimension and degree bound covering numbers and hence sample complexity; singularities correspond to subnetworks and create implicit bias; fibers of the parameterization encode identifiability and symmetries; critical points and data discriminants structure the loss landscape; and Euclidean distance degree quantifies the complexity of distance minimization, and thus of fitting. Together these claims amount to a dictionary in which algebraic geometry is not an analogy but a working formalism for deep learning theory.","pith_inferences":["If the dictionary is correct, architecture design becomes prescriptive: one could choose depth, width, and activation degree to target a desired dimension, degree, or singularity locus of the neuromanifold.","The Hausdorff-approximation step in Section 5.1, once made fully rigorous, would transfer singularity and critical-point results to smooth practical activation functions; the paper sketches but does not prove this transfer.","A testable extension is to check, in a fixed small architecture, whether the number and type of spurious critical points predicted by Euclidean distance degree matches gradient descent's behavior on synthetic data with known ground truth.","The proposed dictionary also suggests a unification of singular learning theory and algebraic statistics under one geometric language, a connection the paper outlines but does not develop."],"forward_implications":["Sample complexity is governed by dimension and degree: the paper cites bounds of the form $\\log N_\\varepsilon(\\mathcal{M}) = O\\!\\left(m \\log\\frac{d}{\\varepsilon} + C\\right)$, converting covering-number bounds into sample-size guarantees.","Expressivity is quantified as a tubular volume: the set of functions within distance $\\varepsilon$ of the neuromanifold has volume bounded by the covering number, so a model's approximate expressive power is controlled by its dimension and degree.","Singular points of the neuromanifold act as implicit biases and, in many architectures, correspond exactly to subnetworks, offering a geometric explanation for automatic selection of simpler functions.","The loss landscape is organized by critical-point invariants: spurious critical points come from the parameterization's critical locus, and the number and type of real critical points change only across data discriminants.","The framework extends beyond polynomial activations: continuous activations can be approximated on compact sets by polynomials, and ReLU networks can be treated with tropical geometry, so the algebraic results carry over approximately."],"supporting_citations":[{"why":"Supplies the covering-number bound and the study of deep polynomial networks that anchors the dimension–degree–sample-complexity connection.","marker":"Kileel et al., 2019"},{"why":"Introduces spurious critical points and the determinantal variety description for linear networks, grounding the critical-point analysis.","marker":"Trager et al., 2020"},{"why":"Provides the geometry and spurious critical point analysis for linear convolutional networks, extending the framework beyond fully connected layers.","marker":"Kohn et al., 2022"},{"why":"Supplies the metric algebraic geometry toolbox, including Euclidean distance degree and data discriminants, for distance problems over varieties.","marker":"Breiding et al., 2024"},{"why":"Establishes singular learning theory, linking singularities of the model to learning dynamics and implicit bias.","marker":"Watanabe, 2009"},{"why":"Gives the sample-complexity results via covering numbers that connect geometric invariants to generalization guarantees.","marker":"Cucker & Smale, 2002"},{"why":"Extends covering-number bounds to ReLU networks by polynomial approximation, supporting the paper's bridge beyond the algebraic domain.","marker":"Zhang & Kileel, 2023"},{"why":"Provides the Tarski–Seidenberg semialgebraic-set theory that makes the neuromanifold of an algebraic model a semi-algebraic variety.","marker":"Bierstone & Milman, 1988"}],"fun_headline_variants":["Geometry of neural networks: an algebraic view","Semi-algebraic varieties in deep learning","The algebra behind neural network expressivity","A new framework: neuroalgebraic geometry","How geometry drives deep learning theory"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that approximating each activation function by polynomials on a compact set also approximates the entire function space closely enough in Hausdorff distance—a step the paper sketches but does not prove.","fun_headline_variants_meta":{"raw":{"variants":["Geometry of neural networks: an algebraic view","Semi-algebraic varieties in deep learning","The algebra behind neural network expressivity","A new framework: neuroalgebraic geometry","How geometry drives deep learning theory"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000533,"raw_usage":{"total_tokens":2505,"prompt_tokens":827,"completion_tokens":1678,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":443,"completion_tokens_details":{"reasoning_tokens":1614}},"tokens_in":443,"tokens_out":1678,"duration_ms":17528,"temperature":1.0,"reasoning_tokens":1614,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-09T21:57:21.999905+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Fix a small MLP architecture with a smooth non-polynomial activation, take compact parameter and input domains, and compute the Hausdorff distance between the true neuromanifold and the neuromanifold obtained by replacing the activation with its degree-$d$ polynomial approximation; if this distance does not go to zero as $d$ grows, the paper's bridge from algebraic models to continuous networks breaks.","supporting_citations":[{"cited_title":"On the expressive power of deep polynomial neural networks","cited_arxiv_id":null,"evidence_quote":"Supplies the covering-number bound and the study of deep polynomial networks that anchors the dimension–degree–sample-complexity connection."},{"cited_title":"Pure and spurious critical points: a geometric study of linear networks","cited_arxiv_id":null,"evidence_quote":"Introduces spurious critical points and the determinantal variety description for linear networks, grounding the critical-point analysis."},{"cited_title":"Geometry of linear convolutional networks","cited_arxiv_id":null,"evidence_quote":"Provides the geometry and spurious critical point analysis for linear convolutional networks, extending the framework beyond fully connected layers."},{"cited_title":"Metric Algebraic Geometry","cited_arxiv_id":null,"evidence_quote":"Supplies the metric algebraic geometry toolbox, including Euclidean distance degree and data discriminants, for distance problems over varieties."},{"cited_title":"Algebraic geometry and statistical learning theory, volume 25","cited_arxiv_id":null,"evidence_quote":"Establishes singular learning theory, linking singularities of the model to learning dynamics and implicit bias."},{"cited_title":"and Smale, S","cited_arxiv_id":null,"evidence_quote":"Gives the sample-complexity results via covering numbers that connect geometric invariants to generalization guarantees."},{"cited_title":"and Kileel, J","cited_arxiv_id":null,"evidence_quote":"Extends covering-number bounds to ReLU networks by polynomial approximation, supporting the paper's bridge beyond the algebraic domain."},{"cited_title":"and Milman, P","cited_arxiv_id":null,"evidence_quote":"Provides the Tarski–Seidenberg semialgebraic-set theory that makes the neuromanifold of an algebraic model a semi-algebraic variety."}],"review_version":1}