Pith. sign in

REVIEW 2 cited by

Self-Expanding Neural Networks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2307.04526 v3 pith:WPU6UWN6 submitted 2023-07-10 cs.LG

classification cs.LG
keywords neuraltrainingarchitectureboundnetworknetworksonlyself-expanding
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The results of training a neural network are heavily dependent on the architecture chosen; and even a modification of only its size, however small, typically involves restarting the training process. In contrast to this, we begin training with a small architecture, only increase its capacity as necessary for the problem, and avoid interfering with previous optimization while doing so. We thereby introduce a natural gradient based approach which intuitively expands both the width and depth of a neural network when this is likely to substantially reduce the hypothetical converged training loss. We prove an upper bound on the ``rate'' at which neurons are added, and a computationally cheap lower bound on the expansion score. We illustrate the benefits of such Self-Expanding Neural Networks with full connectivity and convolutions in both classification and regression problems, including those where the appropriate architecture size is substantially uncertain a priori.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. The Cake that is Intelligence and Who Gets to Bake it: An AI Analogy and its Implications for Participation

    cs.AI 2025-02 accept novelty 5.0 of 10

    The authors expand LeCun's cake metaphor to the full AI lifecycle and argue that social outcomes are constrained by technical foundations such as the i.i.d. assumption, homogenization, catastrophic forgetting, and sur...

  2. Growing with Experience: Growing Neural Networks in Deep Reinforcement Learning

    cs.LG 2025-06 conditional novelty 4.0 of 10

    A simple PPO training scheme that adds hidden layers over time via function-preserving Net2Net morphisms outperforms static networks of the same final depth on MiniHack Room and MuJoCo Ant.

Pith tools