Pith. sign in

REVIEW 4 cited by

Beyond Finite Layer Neural Networks: Bridging Deep Architectures and Numerical Differential Equations

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1710.10121 v3 pith:FWSNXOVD submitted 2017-10-27 cs.CV cs.LGstat.ML

classification cs.CVcs.LGstat.ML
keywords networksnumericalstochasticdeepdifferentialeffectiveequationslm-architecture
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
abstract

In our work, we bridge deep neural network design with numerical differential equations. We show that many effective networks, such as ResNet, PolyNet, FractalNet and RevNet, can be interpreted as different numerical discretizations of differential equations. This finding brings us a brand new perspective on the design of effective deep architectures. We can take advantage of the rich knowledge in numerical analysis to guide us in designing new and potentially more effective deep networks. As an example, we propose a linear multi-step architecture (LM-architecture) which is inspired by the linear multi-step method solving ordinary differential equations. The LM-architecture is an effective structure that can be used on any ResNet-like networks. In particular, we demonstrate that LM-ResNet and LM-ResNeXt (i.e. the networks obtained by applying the LM-architecture on ResNet and ResNeXt respectively) can achieve noticeably higher accuracy than ResNet and ResNeXt on both CIFAR and ImageNet with comparable numbers of trainable parameters. In particular, on both CIFAR and ImageNet, LM-ResNet/LM-ResNeXt can significantly compress ($>50$\%) the original networks while maintaining a similar performance. This can be explained mathematically using the concept of modified equation from numerical analysis. Last but not least, we also establish a connection between stochastic control and noise injection in the training process which helps to improve generalization of the networks. Furthermore, by relating stochastic training strategy with stochastic dynamic system, we can easily apply stochastic training to the networks with the LM-architecture. As an example, we introduced stochastic depth to LM-ResNet and achieve significant improvement over the original LM-ResNet on CIFAR10.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Learning stochastic differential equations using RNN with log signature features

    cs.LG 2019-08 conditional novelty 6.0 of 10

    A hybrid network that feeds coarse log-signature features into an RNN is universal for SDE solution maps and beats baseline RNNs on action and gesture recognition benchmarks.

  2. Neural Dynamics on Complex Networks

    cs.SI 2019-08 conditional novelty 6.0 of 10

    A graph neural network integrated over continuous time learns the differential equations governing networked systems and predicts their future states, outperforming several temporal-graph baselines on simulated dynamics.

  3. NeuPDE: Neural Network Based Ordinary and Partial Differential Equations for Modeling Time-Dependent Data

    cs.LG 2019-08 conditional novelty 6.0 of 10

    NeuPDE learns ODE/PDE models from data by parameterizing the differential equation's right-hand side with a neural network over monomial and derivative features.

  4. Deep Learning Theory Review: An Optimal Control and Dynamical Systems Perspective

    cs.LG 2019-08 conditional novelty 3.0 of 10

    A review that frames neural networks as dynamical systems, SGD as stochastic dynamics, and training as mean-field optimal control to unify deep learning theory.

Pith tools