REVIEW 13 cited by
Implicit Gradient Regularization
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Implicit Gradient Regularization
read the original abstract
Gradient descent can be surprisingly good at optimizing deep neural networks without overfitting and without explicit regularization. We find that the discrete steps of gradient descent implicitly regularize models by penalizing gradient descent trajectories that have large loss gradients. We call this Implicit Gradient Regularization (IGR) and we use backward error analysis to calculate the size of this regularization. We confirm empirically that implicit gradient regularization biases gradient descent toward flat minima, where test errors are small and solutions are robust to noisy parameter perturbations. Furthermore, we demonstrate that the implicit gradient regularization term can be used as an explicit regularizer, allowing us to control this gradient regularization directly. More broadly, our work indicates that backward error analysis is a useful theoretical approach to the perennial question of how learning rate, model size, and parameter regularization interact to determine the properties of overparameterized models optimized with gradient descent.
Forward citations
Cited by 13 Pith papers
-
Towards Universal Convergence of Backward Error in Linear System Solvers
Richardson and a new Krylov method MINBERR achieve universal (condition-free) backward-error rates 1/k and O(1/k^{2}) for PSD linear systems, with a near-universal O(log n / k) extension to general systems.
-
Towards Universal Convergence of Backward Error in Linear System Solvers
Richardson iteration achieves universal 1/k backward error on PSD systems, enabling O(n²/ε) solvers; MINBERR reaches O(1/k²) rate and O(n²/√ε) complexity, with empirical O(1/k) extension to general systems.
-
Characterizing and Correcting Effective Target Shift in Online Learning
Online kernel regression equals offline regression with shifted targets; correcting the targets lets online learning match offline performance and outperform true targets in continual image classification.
-
Estimating Implicit Regularization in Deep Learning
Gradient matching empirically recovers implicit regularization effects such as l2 penalties from early stopping and dropout in neural networks.
-
Avoiding unsafe sets when training with Langevin Dynamics
A Langevin training trajectory's chance of occupying a small failure region relaxes to about twice its tiny stationary value after a burn-in of order the dimension, unless the region's geometry gives a faster local re...
-
Avoiding unsafe sets when training with Langevin Dynamics
Langevin training trajectories on strongly convex losses avoid geometrically isolated failure regions with probability exponentially small in dimension after an O(d) burn-in, with a local spectral rate controlling tra...
-
Second-Order Path Kernel Interpolation Formulas in Machine Learning
Derives second-order path-kernel interpolation formulas for gradient descent, SGD, and momentum training, adding curvature terms and a concentration estimate around the expected prediction.
-
Thermodynamic Irreversibility of Training Algorithms
Four characterizations of irreversibility in training algorithms are equivalent to leading order in step size and produce an emergent force that breaks reparametrization symmetries while favoring minimum entropy produ...
-
Gradient Regularization Mitigates Reward Hacking in Reinforcement Learning from Human Feedback and Verifiable Rewards
Regularizing the policy-gradient norm during RLHF/RLVR training biases the policy toward flat optima where the proxy reward stays accurate, mitigating reward hacking better than a KL penalty.
-
Quantitative Understanding of PDF Fits and their Uncertainties
After an initial transient, a PDF-fitting neural network's output obeys f_t = U(t) f_0 + V(t) Y, a linear blend of the initial network and the data with explicit time-dependent operators.
-
How Optimality Structures Sparse Dictionaries: A Theory for Understanding SAE Representations
Derives optimality constraints for nonnegative joint dictionary learning that explain observed SAE behaviors such as feature splitting, absorption, and dense antipodal features.
-
Design Criteria for SGD Preconditioners: Local Conditioning, Noise Floors, and Basin Stability
The late-stage noise floor of preconditioned SGD is the product of the M-metric condition number and the preconditioned noise level, so the design goal is to improve conditioning while dampening noise.
-
Some Inverse Problems in Particle Physics
Lectures reviewing three established numerical methods for inverse problems in extracting PDFs and spectral functions from lattice QCD and experimental data.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.