Pith. sign in

REVIEW 2 cited by

Distributed Deep Learning Using Synchronous Stochastic Gradient Descent

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1602.06709 v1 pith:MOFG3A7Q submitted 2016-02-22 cs.DC cs.LG

classification cs.DCcs.LG
keywords nodesscalingtrainingalteringclusterdemonstratedesigndistributed
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We design and implement a distributed multinode synchronous SGD algorithm, without altering hyper parameters, or compressing data, or altering algorithmic behavior. We perform a detailed analysis of scaling, and identify optimal design points for different networks. We demonstrate scaling of CNNs on 100s of nodes, and present what we believe to be record training throughputs. A 512 minibatch VGG-A CNN training run is scaled 90X on 128 nodes. Also 256 minibatch VGG-A and OverFeat-FAST networks are scaled 53X and 42X respectively on a 64 node cluster. We also demonstrate the generality of our approach via best-in-class 6.5X scaling for a 7-layer DNN on 16 nodes. Thereafter we attempt to democratize deep-learning by training on an Ethernet based AWS cluster and show ~14X scaling on 16 nodes.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Shesha: Opportunistic In-network Acceleration of Asynchronous Distributed Reinforcement Learning

    cs.NI 2025-07 conditional novelty 6.0 of 10

    Olaf's opportunistic in-network aggregation and replacement of asynchronous DRL updates reduces model staleness and speeds up convergence under congestion.

  2. Model Fusion via Neuron Transplantation

    cs.LG 2025-02 conditional novelty 6.0 of 10

    A new fusion method, Neuron Transplantation, concatenates ensemble members and prunes back down to a single model's size, outperforming individual models after fine-tuning.

Pith tools