Online Batch Selection for Faster Training of Neural Networks

Loshchilov, Ilya, Hutter, Frank , title = · 2015 · cs.LG · arXiv 1511.06343

12 Pith papers cite this work. Polarity classification is still indexing.

12 Pith papers citing it

open full Pith review browse 12 citing papers arXiv PDF

abstract

Deep neural networks are commonly trained using stochastic non-convex optimization procedures, which are driven by gradient information estimated on fractions (batches) of the dataset. While it is commonly accepted that batch size is an important parameter for offline tuning, the benefits of online selection of batches remain poorly understood. We investigate online batch selection strategies for two state-of-the-art methods of stochastic gradient-based optimization, AdaDelta and Adam. As the loss function to be minimized for the whole dataset is an aggregation of loss functions of individual datapoints, intuitively, datapoints with the greatest loss should be considered (selected in a batch) more frequently. However, the limitations of this intuition and the proper control of the selection pressure over time are open questions. We propose a simple strategy where all datapoints are ranked w.r.t. their latest known loss value and the probability to be selected decays exponentially as a function of rank. Our experimental results on the MNIST dataset suggest that selecting batches speeds up both AdaDelta and Adam by a factor of about 5.

citation-role summary

background 1

citation-polarity summary

background 1

representative citing papers

Initialization is Half the Battle: Generating Diverse Images from a Guidance Potential Posterior

cs.CV · 2026-06-01 · unverdicted · novelty 7.0

DivIn samples initial noise from a guidance potential posterior via Langevin dynamics to improve diversity in class-to-image and text-to-image generation.

Disagreement-Regularized Importance Sampling for Adversarial Label Corruption

cs.LG · 2026-05-08 · unverdicted · novelty 7.0

DR-IS selects low-contamination subsets via bounded rank-disagreement in proxy ensembles under an ε-contamination model, with O(√(log(N/δ)/K)) concentration rates that certify separation when the expectation gap Δ' is positive.

Online Data Selection for Instruction Tuning via Gaussian Processes

cs.LG · 2026-06-29 · unverdicted · novelty 6.0

GAIA models continuous utility with Gaussian processes across semantic space and applies fixed-share Hedge updates to achieve dynamic regret guarantees while outperforming baselines on three datasets.

Selecting Samples on Graphs: A Unified Dataset Pruning Framework for Lossless Training Acceleration

cs.LG · 2026-06-11 · unverdicted · novelty 6.0

A graph-based unified dataset pruning framework that formulates pruning as MWCP, derives a greedy algorithm with formal approximation guarantees, and demonstrates substantial training acceleration on ImageNet.

Why SGD is not Brownian Motion: A New Perspective on Stochastic Dynamics

cs.LG · 2026-05-21 · unverdicted · novelty 6.0

SGD is reformulated via a master equation from discrete updates, producing a discrete Fokker-Planck equation that predicts non-stationary variance growth proportional to learning rate in flat Hessian directions.

Variance Matters: Improving Domain Adaptation via Stratified Sampling

cs.LG · 2025-12-04 · unverdicted · novelty 6.0

VaRDASS improves unsupervised domain adaptation by using stratified sampling to reduce variance in discrepancy estimation for measures like correlation alignment and MMD, with derived error bounds, an optimality proof for MMD under assumptions, and a k-means style algorithm.

Inverse Attention Guided Deep Crowd Counting Network

cs.CV · 2019-07-02 · unverdicted · novelty 6.0

IA-DCCN is a single-step VGG-16 network that infuses segmentation via inverse attention to improve crowd counting accuracy on three datasets with minimal overhead.

Dr. Post-Training: A Data Regularization Perspective on LLM Post-Training

cs.LG · 2026-05-08 · unverdicted · novelty 6.0

Dr. Post-Training reframes general data as a data-induced regularizer for LLM post-training updates, yielding a family of methods that outperform data-selection baselines on SFT, RLHF, and RLVR tasks.

Data Warmup: Complexity-Aware Curricula for Efficient Diffusion Training

cs.LG · 2026-04-08 · conditional · novelty 6.0

Data Warmup accelerates diffusion training on ImageNet by scheduling images from low to high complexity via a foreground-based metric and temperature-controlled sampler, improving FID and IS scores faster than uniform sampling.

RCAP: Robust, Class-Aware, Probabilistic Dynamic Dataset Pruning

cs.LG · 2026-06-10 · unverdicted · novelty 4.0

RCAP introduces class-aware probabilistic pruning that uses closed-form per-class fractions updated by loss and high-loss sampling to preserve worst-group accuracy at high pruning rates.

Learning to Reason at the Frontier of Learnability

cs.LG · 2025-02-17 · unverdicted · novelty 4.0

A curriculum sampling questions with high variance in success rate improves reinforcement learning performance for LLM reasoning tasks.

Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model

cs.CV · 2025-02-14 · unverdicted · novelty 4.0

Step-Video-T2V describes a 30B-parameter text-to-video model with custom Video-VAE, 3D DiT, flow matching, and Video-DPO that claims state-of-the-art results on a new internal benchmark.

citing papers explorer

Showing 12 of 12 citing papers.

Initialization is Half the Battle: Generating Diverse Images from a Guidance Potential Posterior cs.CV · 2026-06-01 · unverdicted · none · ref 61 · internal anchor
DivIn samples initial noise from a guidance potential posterior via Langevin dynamics to improve diversity in class-to-image and text-to-image generation.
Disagreement-Regularized Importance Sampling for Adversarial Label Corruption cs.LG · 2026-05-08 · unverdicted · none · ref 23
DR-IS selects low-contamination subsets via bounded rank-disagreement in proxy ensembles under an ε-contamination model, with O(√(log(N/δ)/K)) concentration rates that certify separation when the expectation gap Δ' is positive.
Online Data Selection for Instruction Tuning via Gaussian Processes cs.LG · 2026-06-29 · unverdicted · none · ref 43 · internal anchor
GAIA models continuous utility with Gaussian processes across semantic space and applies fixed-share Hedge updates to achieve dynamic regret guarantees while outperforming baselines on three datasets.
Selecting Samples on Graphs: A Unified Dataset Pruning Framework for Lossless Training Acceleration cs.LG · 2026-06-11 · unverdicted · none · ref 5 · internal anchor
A graph-based unified dataset pruning framework that formulates pruning as MWCP, derives a greedy algorithm with formal approximation guarantees, and demonstrates substantial training acceleration on ImageNet.
Why SGD is not Brownian Motion: A New Perspective on Stochastic Dynamics cs.LG · 2026-05-21 · unverdicted · none · ref 164 · internal anchor
SGD is reformulated via a master equation from discrete updates, producing a discrete Fokker-Planck equation that predicts non-stationary variance growth proportional to learning rate in flat Hessian directions.
Variance Matters: Improving Domain Adaptation via Stratified Sampling cs.LG · 2025-12-04 · unverdicted · none · ref 25 · internal anchor
VaRDASS improves unsupervised domain adaptation by using stratified sampling to reduce variance in discrepancy estimation for measures like correlation alignment and MMD, with derived error bounds, an optimality proof for MMD under assumptions, and a k-means style algorithm.
Inverse Attention Guided Deep Crowd Counting Network cs.CV · 2019-07-02 · unverdicted · none · ref 20 · internal anchor
IA-DCCN is a single-step VGG-16 network that infuses segmentation via inverse attention to improve crowd counting accuracy on three datasets with minimal overhead.
Dr. Post-Training: A Data Regularization Perspective on LLM Post-Training cs.LG · 2026-05-08 · unverdicted · none · ref 68
Dr. Post-Training reframes general data as a data-induced regularizer for LLM post-training updates, yielding a family of methods that outperform data-selection baselines on SFT, RLHF, and RLVR tasks.
Data Warmup: Complexity-Aware Curricula for Efficient Diffusion Training cs.LG · 2026-04-08 · conditional · none · ref 18
Data Warmup accelerates diffusion training on ImageNet by scheduling images from low to high complexity via a foreground-based metric and temperature-controlled sampler, improving FID and IS scores faster than uniform sampling.
RCAP: Robust, Class-Aware, Probabilistic Dynamic Dataset Pruning cs.LG · 2026-06-10 · unverdicted · none · ref 62 · internal anchor
RCAP introduces class-aware probabilistic pruning that uses closed-form per-class fractions updated by loss and high-loss sampling to preserve worst-group accuracy at high pruning rates.
Learning to Reason at the Frontier of Learnability cs.LG · 2025-02-17 · unverdicted · none · ref 47 · internal anchor
A curriculum sampling questions with high variance in success rate improves reinforcement learning performance for LLM reasoning tasks.
Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model cs.CV · 2025-02-14 · unverdicted · none · ref 268 · internal anchor
Step-Video-T2V describes a 30B-parameter text-to-video model with custom Video-VAE, 3D DiT, flow matching, and Video-DPO that claims state-of-the-art results on a new internal benchmark.

Online Batch Selection for Faster Training of Neural Networks

citation-role summary

citation-polarity summary

fields

years

verdicts

roles

polarities

representative citing papers

citing papers explorer