Fetching the paper…
Reading the bibliography…
Stochastic Gradient Descent (SGD) has become the de facto way to train deep neural networks in distributed clusters.
A. V. Gerbessiotis et al. , “Direct Bulk-Synchronous Parallel Algorithms,” J. Parallel Distrib. Comput. , 1994
1994
Earlier work this paper cites.
S. Hochreiter et al. , “Simplifying neural nets by discovering flat minima,” in NeurIPS , 1995
1995
Earlier work this paper cites.
S. Hochreiter et al. , “Flat minima,” Neural Computation , 1997
1997
Earlier work this paper cites.
A. Li et al. , “CloudCmp: comparing public cloud providers,” ACM IMC , 2010
2010
Earlier work this paper cites.
J. Xie et al. , “Improving mapreduce performance through data placement in heterogeneous hadoop clusters,” IPDPSW , 2010
2010
Earlier work this paper cites.
B. Recht et al. , “Hogwild: A Lock-Free Approach to Parallelizing Stochastic Gradient Descent,” NeurIPS , 2011
2011
Earlier work this paper cites.
J. Dean et al. , “Large scale distributed deep networks,” NeurIPS , 2012
2012
Earlier work this paper cites.
A. Senior et al. , “An empirical study of learning rates in deep neural networks for speech recognition,” ICASSP , 2013
2013
Earlier work this paper cites.
Q. Ho et al. , “More effective distributed ml via a stale synchronous parallel parameter server,” NeurIPS , 2013
2013
Earlier work this paper cites.
F. Yan et al. , “Performance Modeling and Scalability Optimization of Distributed Deep Learning Systems,” SIGKDD , 2015
2015
Earlier work this paper cites.
J. Chen et al. , “Revisiting distributed synchronous sgd,” arXiv preprint arXiv:1604.00981 , 2016
2016
Earlier work this paper cites.
K. He et al. , “Deep residual learning for image recognition,” CVPR , 2016
2016
Earlier work this paper cites.
K. Hsieh et al. , “Gaia: Geo-distributed machine learning approaching lan speeds,” NSDI , 2017
2017
Earlier work this paper cites.
G. Huang et al. , “Densely connected convolutional networks,” CVPR , 2017
2017
Earlier work this paper cites.
N. S. Keskar et al. , “On large-batch training for deep learning: Generalization gap and sharp minima,” ICLR , 2017
2017
Cited alongside, same era.
Q. Duan, “Cloud service performance evaluation: status, challenges, and opportunities–a survey from the system modeling perspective,” Digital Communications and Networks , 2017
2017
Cited alongside, same era.
2017
Cited alongside, same era.
2017
Cited alongside, same era.
A. Krizhevsky et al. , “Cifar-10,” http://www.cs.toronto.edu/~kriz/cifar.html , 2017
2017
B. Kleinberg et al. , “An alternative view: When does sgd escape local minima?” ICML , 2018
2018
Later among the works it cites.
A. Vaswani et al. , “Tensor2tensor for neural machine translation,” CoRR , 2018
2018
Later among the works it cites.
S. L. Smith et al. , “A bayesian perspective on generalization and stochastic gradient descent,” ICLR , 2018
2018
Later among the works it cites.
S. Li et al. , “Speeding up Deep Learning with Transient Servers,” ICAC , 2019
2019
Later among the works it cites.
T. Ben-Nun et al. , “Demystifying Parallel and Distributed Deep Learning: An In-depth Concurrency Analysis,” ACM Comput. Surv. , 2019
2019
Later among the works it cites.
X. Zhao et al. , “Dynamic stale synchronous parallel distributed training for deep learning,” ICDCS , 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
X. Gastaldi, “Shake-shake regularization,” arXiv preprint arXiv:1705.07485 , 2017
2017
Cited alongside, same era.
F. Chollet, “Xception: Deep learning with depthwise separable convolutions,” CVPR , 2017
2017
Cited alongside, same era.
2017
Cited alongside, same era.
W. Wen et al. , “TernGrad: Ternary Gradients to Reduce Communication in Distributed Deep Learning,” NeurIPS , 2017
2017
Cited alongside, same era.
D. Alistarh et al. , “QSGD: Communication-Efficient SGD via Gradient Quantization and Encoding,” NeurIPS , 2017
2017
Cited alongside, same era.
L. N. Smith, “Cyclical learning rates for training neural networks,” WACV , 2017
2017
Cited alongside, same era.
S. Shi et al. , “Performance modeling and evaluation of distributed deep learning frameworks on gpus,” DASC/PiCom/DataCom/CyberSciTech , 2018
2018
Cited alongside, same era.
2019
Later among the works it cites.
W. Jiang et al. , “A novel stochastic gradient descent algorithm based on grouping over heterogeneous cluster systems for distributed deep learning,” CCGRID , 2019
2019
Later among the works it cites.
2019
Later among the works it cites.
C. Coleman et al. , “Analysis of dawnbench, a time-to-accuracy machine learning performance benchmark,” SIGOPS Operating Systems Review , 2019
2019
Later among the works it cites.
P. Chaudhari et al. , “Entropy-sgd: Biasing gradient descent into wide valleys,” Journal of Statistical Mechanics: Theory and Experiment , 2019
2019
Later among the works it cites.
S. Li et al. , “Characterizing and Modeling Distributed Training with Transient Cloud GPU Servers,” ICDCS , 2020
2020
Later among the works it cites.
2020
Later among the works it cites.
A. Or et al. , “Resource Elasticity in Distributed Deep Learning,” Proceedings of Machine Learning and Systems , 2020
2020
Later among the works it cites.