Pith. sign in

REVIEW 2 cited by

mL-BFGS: A Momentum-based L-BFGS for Distributed Large-Scale Neural Network Optimization

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2307.13744 v1 pith:IME5KWTL submitted 2023-07-25 cs.LG math.OC

classification cs.LGmath.OC
keywords ml-bfgsstochastictrainingl-bfgslarge-scaleconvergencehessianneural
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Quasi-Newton methods still face significant challenges in training large-scale neural networks due to additional compute costs in the Hessian related computations and instability issues in stochastic training. A well-known method, L-BFGS that efficiently approximates the Hessian using history parameter and gradient changes, suffers convergence instability in stochastic training. So far, attempts that adapt L-BFGS to large-scale stochastic training incur considerable extra overhead, which offsets its convergence benefits in wall-clock time. In this paper, we propose mL-BFGS, a lightweight momentum-based L-BFGS algorithm that paves the way for quasi-Newton (QN) methods in large-scale distributed deep neural network (DNN) optimization. mL-BFGS introduces a nearly cost-free momentum scheme into L-BFGS update and greatly reduces stochastic noise in the Hessian, therefore stabilizing convergence during stochastic optimization. For model training at a large scale, mL-BFGS approximates a block-wise Hessian, thus enabling distributing compute and memory costs across all computing nodes. We provide a supporting convergence analysis for mL-BFGS in stochastic settings. To investigate mL-BFGS potential in large-scale DNN training, we train benchmark neural models using mL-BFGS and compare performance with baselines (SGD, Adam, and other quasi-Newton methods). Results show that mL-BFGS achieves both noticeable iteration-wise and wall-clock speedup.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Domain-Decomposition Neural Surrogates for Scalable Decentralized Ensemble Kalman Filter Based Parameter Identification in High-Dimensional Stochastic PDEs

    cs.CE 2026-07 conditional novelty 5.5 of 10

    Domain-decomposed neural surrogates with augmented-Lagrange coupling, paired with a block-preconditioned decentralized EnKF, match FEM-EnKF and approach MCMC posteriors on 3D elastic parameter ID at reduced forecast cost.

  2. Can adversarial attacks by large language models be attributed?

    cs.AI 2024-11 conditional novelty 4.0 of 10

    Under a worst-case formal model, attributing LLM outputs to a specific model is provably impossible except in the narrow case of finitely many deterministic models.

Pith tools