Pith. sign in

REVIEW 1 cited by

A Scale Mixture Perspective of Multiplicative Noise in Neural Networks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1506.03208 v1 pith:ZVQLEORZ submitted 2015-06-10 stat.ML

classification stat.ML
keywords multiplicativenoiseweightsanalysisapproachdeepfunctiongaussian
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Corrupting the input and hidden layers of deep neural networks (DNNs) with multiplicative noise, often drawn from the Bernoulli distribution (or 'dropout'), provides regularization that has significantly contributed to deep learning's success. However, understanding how multiplicative corruptions prevent overfitting has been difficult due to the complexity of a DNN's functional form. In this paper, we show that when a Gaussian prior is placed on a DNN's weights, applying multiplicative noise induces a Gaussian scale mixture, which can be reparameterized to circumvent the problematic likelihood function. Analysis can then proceed by using a type-II maximum likelihood procedure to derive a closed-form expression revealing how regularization evolves as a function of the network's weights. Results show that multiplicative noise forces weights to become either sparse or invariant to rescaling. We find our analysis has implications for model compression as it naturally reveals a weight pruning rule that starkly contrasts with the commonly used signal-to-noise ratio (SNR). While the SNR prunes weights with large variances, seeing them as noisy, our approach recognizes their robustness and retains them. We empirically demonstrate our approach has a strong advantage over the SNR heuristic and is competitive to retraining with soft targets produced from a teacher model.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Complexity-Aware Training of Deep Neural Networks for Optimal Structure Discovery

    cs.LG 2024-11 conditional novelty 6.0 of 10

    Combined unit and layer pruning is achieved during a single training run via a complexity-aware variational objective, producing deterministic networks with pruning ratios around 25% to 77% FLOPS at modest accuracy cost.

Pith tools