{"id":"dd77c724-928b-4c16-adf0-7f1b6c0c2445","arxiv_id":"2607.05546","paper_version":1,"verdict":"ACCEPT","confidence":"HIGH","novelty_score":7.5,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"Deep neural variation spaces remain small at any depth; univariate ReLU saturates after depth 2 up to a factor of 2, so norm-controlled deep ReLU nets cannot be highly oscillatory along any direction.","lead":"The paper builds a unified function-space norm for deep fully-connected nets that works for many practical activations, not just ReLU. Under this norm, depth barely enlarges the univariate ReLU class and forbids high-frequency behavior along lines, so some classic depth-expressivity claims vanish once rescaling is controlled.","discovery_kind":"unification","skeptic_critique":{"model":"grok-4.5","headline":"No significant objection identified","rationale":"The reader correctly isolates the geometry-specific construction of I as the most delicate step, yet that step is fully spelled out in Appendices A.16–A.18 and is self-contained. The remainder of the paper (representer theorem, Rademacher/entropy bounds, norm equivalences) is independent of the saturation result and is likewise carefully proved. Because the load-bearing claim holds under its stated hypotheses and no internal inconsistency appears, the ACCEPT verdict with high confidence stands.","tokens_in":73480,"tokens_out":452,"duration_ms":4454,"concrete_test":"Independently re-derive the bound I(f+) ≤ I(f) for a single interior negative interval [a,b] ⊂ (-1,1) using only the BV integration-by-parts identities (427)–(429) and the sign constraints f'(a-) ≤ 0 ≤ f'(b+); if the total-variation comparison (432)–(435) fails for any continuous piecewise-linear test function with one interior negative segment, the saturation constant is not secured.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The strongest claim (Theorem 4.3 and Corollary 4.8) is proved under the explicit geometry Ω = W = B = [-1,1] that the authors state. The functional I of Lemma 4.2 is constructed precisely for that interval, is shown non-increasing under positive part by a careful case analysis of negative components, and is bounded by 2 on B1 by direct evaluation of endpoint slopes and values. Remark 4.4 exhibits a family fα whose positive parts force the constant 2 to be sharp. The multivariate frequency-control corollary follows immediately by restriction to lines (Proposition 4.6) once the univariate saturation is in hand. No hidden assumption, circular step, or gap in the chain from I to the claimed containments is apparent; the geometry dependence is already flagged by the reader and does not undermine the stated theorems.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper introduces deep neural variation spaces VL whose unit balls BL are defined recursively as closed absolutely convex hulls of normalized activations applied to the previous layer, starting from a compact affine dictionary B1. This yields a Banach function-space norm compatible with both homogeneous and many non-homogeneous activations (ReLU, GELU, SiLU, etc.), recovers the path norm / generalized Barron spaces for ReLU, and supplies a representer theorem for norm-penalized data fitting. The authors prove Rademacher and Lp metric-entropy bounds on BL that grow at most mildly with depth, compare the norm to SOSW and several additive layerwise costs, and establish a univariate ReLU depth-saturation result B2 ⊂ BL ⊂ 2B2 (Theorem 4.3) on Ω = W = B = [−1, 1]. As a corollary, multivariate functions in BL cannot exhibit second-derivative total variation larger than 2 along any line through the origin, so high-frequency behavior along individual directions is forbidden under this norm.","tokens_in":73647,"tokens_out":943,"duration_ms":15668,"significance":"If the results hold, the paper supplies a clean, activation-general function-space language that separates genuine compositional nonlinearity from compounded layerwise rescaling, and shows that several widely cited depth-expressivity phenomena (Telgarsky sawteeth, high-frequency depth separations) rely on the latter once complexity is measured by a true function norm. The univariate saturation theorem and the line-restriction frequency-control corollary are sharp structural statements with complete proofs; the representer theorem and the mild depth dependence of the complexity bounds are useful for regularization and generalization analyses. Complete appendix proofs (Riesz–Markov, Banach–Alaoglu, Ledoux–Talagrand, BV theory) and explicit recovery of prior path-norm / Barron constructions are clear strengths.","major_comments":[],"minor_comments":[{"comment":"Abstract and §1: the phrase “some commonly cited expressivity benefits of depth disappear” is accurate for the VL norm, but a short clarifying sentence that the claim is relative to true function norms (not width or SOSW) would reduce possible misreading against the classical depth-separation literature.","section":null},{"comment":"Table 1 and Lemma 2.7: the depth-dependent equivalence constants for SELU, absolute value, and bent identity grow exponentially; a one-line remark in the main text that these bases remain modest for typical L would help practitioners.","section":null},{"comment":"Theorem 2.8: the width bound Kℓ ≤ N^{L−ℓ} is correctly stated as possibly improvable; citing the tighter N-per-layer bounds of Parhi–Nowak / Shenouda et al. more explicitly in the main text (rather than only in the comparison paragraph) would orient the reader.","section":null},{"comment":"Figure 2 / Remark 4.4: the sharpness example fα is clear; adding the numerical V2 values of (fα)+ for the three plotted α would make the approach to the constant 2 more immediate.","section":null},{"comment":"§3.2.1 and Appendix A.15: the ℓ1-pyramid example in B3 \\ B2 is useful; a brief pointer in the main text that this is currently the main explicit multivariate witness of strict depth increase would strengthen the open-problems discussion.","section":null},{"comment":"Notation: the dual use of BL for both the unit ball and (in places) the depth index is mostly clear from context, but a single sentence at the start of §2.2 fixing “BL always denotes the unit ball of VL” would avoid occasional ambiguity.","section":null}],"recommendation":"accept","confidential_remarks":"Strong theory paper with complete proofs and a genuinely new structural result (univariate saturation + frequency control under a true function norm). Fit for a top ML theory venue is high; the open characterization of BL for L ≥ 3 is honestly framed and does not undermine the published claims. No novelty or citation concerns."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The punchline is Theorem 4.3: under the natural geometry Ω = W = B = [-1,1], the univariate ReLU unit balls satisfy B2 ⊂ BL ⊂ 2B2 for every depth. Depth only rescales; it does not add shapes. The multivariate corollary (4.8) then forbids high second-derivative total variation along any line, so many classic depth-separation examples are ruled out once you control a genuine function norm rather than width or SOSW.\n\nWhat is new and solid: the recursive absolutely-convex construction with normalized activations σs works for a long list of non-homogeneous maps (GELU, SiLU, Mish, ELU, etc.) and recovers path-norm / generalized Barron spaces for ReLU. They give a representer theorem, Rademacher and Lp-entropy bounds that grow only mildly with depth, explicit norm-equivalence constants (Table 1), and a careful comparison showing SOSW and the various group-norm costs sit inside their balls after depth-dependent powers. The proofs are complete and use standard tools (Riesz–Markov, Banach–Alaoglu, BV theory, Ledoux–Talagrand) with careful bookkeeping; the functional I of Lemma 4.2 that is non-increasing under positive part is the technical heart of the saturation argument and looks correct. Remark 4.4 shows the constant 2 is sharp for that geometry.\n\nSoft spots, in proportion: the saturation geometry is special (they flag this), equivalence constants for non-ReLU activations can grow with L, and the multivariate deep balls BL \\ B2 remain almost completely uncharacterized beyond pyramid examples. Those are real open problems the authors state clearly; they do not break the theorems that are proved.\n\nThis is for people who care about function-space inductive bias and the interpretation of depth-separation results. Math and citations look solid; no circularity. I would send it to referees and I would cite the saturation and frequency-control statements.","headline":"Clean Banach-norm theory for deep nets that actually saturates depth for univariate ReLU and forces frequency control under true function-space cost.","tokens_in":74257,"tokens_out":499,"would_cite":true,"duration_ms":9543,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["68T07","41A46","46E35","62G05"],"pacs":[],"model":"grok-4.5","headline":"Under a true function-space norm, depth barely enlarges the class of representable ReLU functions and forbids high frequencies along every line.","keywords":["deep neural networks","variation spaces","path norm","depth saturation","metric entropy","Rademacher complexity","ReLU","function space norms"],"falsifier":"Exhibit a continuous univariate function whose second-derivative total variation exceeds 2 yet still lies in some deep unit ball BL for L greater than or equal to 3 under the paper’s exact choice of domain and dictionary, or prove that no such function exists.","tokens_in":74367,"feed_emoji":"📐","tokens_out":609,"duration_ms":5677,"temperature":0.7,"pith_summary":"The paper builds a single recursive variation space for deep fully connected networks that works for both homogeneous and non-homogeneous activations. Functions are measured by how large an absolutely convex combination of activated previous-layer functions they need, so the norm is a genuine Banach norm on the scalar output rather than an additive layer-wise parameter cost. With that control in place, the unit balls remain small at every depth: their Rademacher complexities and metric entropies grow only mildly with depth. In one dimension for ReLU the story is sharper still: every deep unit ball is trapped between the shallow unit ball and twice that ball. Consequently any multivariate function whose variation norm is bounded cannot oscillate rapidly along any straight line. The authors conclude that many famous “depth-separation” examples rely on hidden layer-wise rescaling that a true function norm forbids.","feed_headline":"Depth barely enlarges ReLU classes under true function norms","feed_subtitle":"High frequencies along any line are forbidden once layer-wise rescaling is controlled","key_machinery":"Deep neural variation spaces VL whose unit balls are the closed absolute convex hulls of normalized activations of the previous layer; the associated path-norm upper bound and the auxiliary functional I that is non-increasing under the positive-part map.","core_discovery":"Once network complexity is measured by a genuine function-space variation norm rather than width or sum-of-squared weights, depth alone does not create high-frequency or highly oscillatory functions; in the univariate ReLU case it merely multiplies the shallow unit ball by a depth-independent factor of two.","pith_inferences":[],"forward_implications":[],"fun_headline_variants":["Depth only rescales ReLU classes by constant under variation norms","No high frequencies from depth when complexity uses function norms","Univariate ReLU depth saturates to twice the shallow unit ball","Function-space norms strip depth of high-oscillation benefits","True variation norms keep deep ReLUs no richer than shallow up to scale"],"cache_read_input_tokens":65664,"weakest_assumption_plain":"The sharp depth-saturation constant of two is proved only for the compact interval [-1,1] with first-layer weights and biases also restricted to [-1,1]; the same geometry may not hold for other domains or dictionaries.","fun_headline_variants_meta":{"raw":{"variants":["Depth only rescales ReLU classes by constant under variation norms","No high frequencies from depth when complexity uses function norms","Univariate ReLU depth saturates to twice the shallow unit ball","Function-space norms strip depth of high-oscillation benefits","True variation norms keep deep ReLUs no richer than shallow up to scale"]},"model":"grok-4.5","effort":"low","cost_usd":0.00815,"raw_usage":{"total_tokens":1955,"prompt_tokens":799,"num_sources_used":0,"completion_tokens":90,"cost_in_usd_ticks":81500000,"prompt_tokens_details":{"text_tokens":799,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":1066,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":799,"tokens_out":90,"duration_ms":6707,"temperature":1.0,"reasoning_tokens":1066,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-11T06:00:29.701431+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Exhibit a continuous univariate function whose second-derivative total variation exceeds 2 yet still lies in some deep unit ball BL for L greater than or equal to 3 under the paper’s exact choice of domain and dictionary, or prove that no such function exists.","supporting_citations":[],"review_version":1}