REVIEW 2 cited by
Self-Expanding Neural Networks
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
The results of training a neural network are heavily dependent on the architecture chosen; and even a modification of only its size, however small, typically involves restarting the training process. In contrast to this, we begin training with a small architecture, only increase its capacity as necessary for the problem, and avoid interfering with previous optimization while doing so. We thereby introduce a natural gradient based approach which intuitively expands both the width and depth of a neural network when this is likely to substantially reduce the hypothetical converged training loss. We prove an upper bound on the ``rate'' at which neurons are added, and a computationally cheap lower bound on the expansion score. We illustrate the benefits of such Self-Expanding Neural Networks with full connectivity and convolutions in both classification and regression problems, including those where the appropriate architecture size is substantially uncertain a priori.
Forward citations
Cited by 2 Pith papers
-
The Cake that is Intelligence and Who Gets to Bake it: An AI Analogy and its Implications for Participation
The authors expand LeCun's cake metaphor to the full AI lifecycle and argue that social outcomes are constrained by technical foundations such as the i.i.d. assumption, homogenization, catastrophic forgetting, and sur...
-
Growing with Experience: Growing Neural Networks in Deep Reinforcement Learning
A simple PPO training scheme that adds hidden layers over time via function-preserving Net2Net morphisms outperforms static networks of the same final depth on MiniHack Room and MuJoCo Ant.
Discussion (0). Continue with ORCID to comment.