REVIEW 2 cited by
Variational Learning is Effective for Large Deep Networks
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We give extensive empirical evidence against the common belief that variational learning is ineffective for large neural networks. We show that an optimizer called Improved Variational Online Newton (IVON) consistently matches or outperforms Adam for training large networks such as GPT-2 and ResNets from scratch. IVON's computational costs are nearly identical to Adam but its predictive uncertainty is better. We show several new use cases of IVON where we improve finetuning and model merging in Large Language Models, accurately predict generalization error, and faithfully estimate sensitivity to data. We find overwhelming evidence that variational learning is effective.
Forward citations
Cited by 2 Pith papers
-
Optimization Guarantees for Square-Root Natural-Gradient Variational Inference
For strongly concave log-likelihoods, square-root (Cholesky) parametrization of Gaussian variational inference yields exponential convergence guarantees for both the natural-gradient flow and a discrete-time natural-g...
-
Spectral-factorized Positive-definite Curvature Learning for NN Training
The paper derives a Riemannian update rule for the spectral factors of a positive-definite preconditioner, making arbitrary matrix roots fast and numerically stable for low-precision NN training.
Discussion (0). Continue with ORCID to comment.