Loss Surfaces, Mode Connectivity, and Fast Ensembling of DNNs

· 2018 · stat.ML · arXiv 1802.10026

3 Pith papers cite this work. Polarity classification is still indexing.

3 Pith papers citing it

open full Pith review browse 3 citing papers arXiv PDF

abstract

The loss functions of deep neural networks are complex and their geometric properties are not well understood. We show that the optima of these complex loss functions are in fact connected by simple curves over which training and test accuracy are nearly constant. We introduce a training procedure to discover these high-accuracy pathways between modes. Inspired by this new geometric insight, we also propose a new ensembling method entitled Fast Geometric Ensembling (FGE). Using FGE we can train high-performing ensembles in the time required to train a single model. We achieve improved performance compared to the recent state-of-the-art Snapshot Ensembles, on CIFAR-10, CIFAR-100, and ImageNet.

representative citing papers

Measuring Alignment-Induced Activation Shifts Correctly: A Template-Controlled Difference-in-Differences Protocol

cs.LG · 2026-05-23 · unverdicted · novelty 7.0

Introduces a template-controlled difference-in-differences protocol that corrects chat-template confounding when measuring alignment-induced activation shifts in LLMs and recovers the refusal direction with higher fidelity.

Recoverable but Not Stationary:Local Linear Structures in Weights and Activations

cs.LG · 2026-06-09 · unverdicted · novelty 6.0

Local low-rank task-gradient structures exist in weights and activations but are non-stationary, with initial recovery updates forming a basis capturing 77% of LoRA displacement and parameter steps aligning 0.58 cosine with CAA steering vectors.

Statistical Properties of Training & Generalization

stat.ML · 2026-06-18 · unverdicted · novelty 1.0

Review of neural scaling laws and their relation to constraints and inductive biases when applying machine learning to physics problems.

citing papers explorer

Showing 3 of 3 citing papers.

Measuring Alignment-Induced Activation Shifts Correctly: A Template-Controlled Difference-in-Differences Protocol cs.LG · 2026-05-23 · unverdicted · none · ref 4 · internal anchor
Introduces a template-controlled difference-in-differences protocol that corrects chat-template confounding when measuring alignment-induced activation shifts in LLMs and recovers the refusal direction with higher fidelity.
Recoverable but Not Stationary:Local Linear Structures in Weights and Activations cs.LG · 2026-06-09 · unverdicted · none · ref 6 · internal anchor
Local low-rank task-gradient structures exist in weights and activations but are non-stationary, with initial recovery updates forming a basis capturing 77% of LoRA displacement and parameter steps aligning 0.58 cosine with CAA steering vectors.
Statistical Properties of Training & Generalization stat.ML · 2026-06-18 · unverdicted · none · ref 62 · internal anchor
Review of neural scaling laws and their relation to constraints and inductive biases when applying machine learning to physics problems.

Loss Surfaces, Mode Connectivity, and Fast Ensembling of DNNs

fields

years

verdicts

representative citing papers

citing papers explorer