Fetching the paper…
Reading the bibliography…
Deep learning models are trained on servers with many GPUs, and training must scale with the number of GPUs.
Deep Learning with Elastic Averaging SGD
S. Zhang, A. Choromanska, and Y. LeCun · 1901
Earlier work this paper cites.
A stochastic approximation method
H. Robbins and S. Monro · 1951
Earlier work this paper cites.
Some methods of speeding up the convergence of iteration methods
B. Polyak · 1964
Earlier work this paper cites.
A method of solving a convex programming problem with convergence rate O(1/sqr(k))
Y. Nesterov · 1983
Earlier work this paper cites.
Learning internal representations by error propagation
D. E. Rumelhart, G. E. Hinton, and R. J. Williams · 1986
Earlier work this paper cites.
Efficient estimators from a slowly convergent Robbins-Monro process
D. Ruppert · 1988
Earlier work this paper cites.
New stochastic approximation type procedures
B. Polyak · 1990
Earlier work this paper cites.
Acceleration of stochastic approximation by averaging
B. Polyak and A. Juditsky · 1992
Earlier work this paper cites.
Flat minima
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
On-line learning and stochastic approximations
L. Bottou · 1998
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner · 1998
Earlier work this paper cites.
Distributed optimization and statistical learning via the alternating direction method of multipliers
S. Boyd, N. Parikh, E. Chu, B. Peleato, and J. Eckstein · 2011
Earlier work this paper cites.
Hogwild: A Lock-Free Approach to Parallelizing Stochastic Gradient Descent
B. Recht, C. Re, S. Wright, and F. Niu · 2011
Earlier work this paper cites.
Towards optimal one pass large scale learning with averaged stochastic gradient descent
W. Xu · 2011
Earlier work this paper cites.
Large Scale Distributed Deep Networks
J. Dean, G. Corrado, R. Monga, K. Chen, M. Devin, M. Mao, M. aurelio Ranzato, A. Senior, P. Tucker, K. Yang, Q. V. Le, and A. Y. Ng · 2012
Earlier work this paper cites.
Deep Neural Networks for Acoustic Modeling in Speech Recognition: The Shared Views of Four Research Groups
G. Hinton, L. Deng, D. Yu, G. E. Dahl, A. Mohamed, N. Jaitly, A. Senior, V. Vanhoucke, P. Nguyen, T. N. Sainath, and B. Kingsbury · 2012
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
A. Krizhevsky, I. Sutskever, and G. E. Hinton · 2012
Earlier work this paper cites.
Efficient BackProp
Y. A. LeCun, L. Bottou, G. B. Orr, and K.-R. Müller · 2012
Earlier work this paper cites.
More Effective Distributed ML via a Stale Synchronous Parallel Parameter Server
Q. Ho, J. Cipar, H. Cui, S. Lee, J. K. Kim, P. B. Gibbons, G. A. Gibson, G. Ganger, and E. P. Xing · 2013
Earlier work this paper cites.
On the importance of initialization and momentum in deep learning
I. Sutskever, J. Martens, G. Dahl, and G. Hinton · 2013
Earlier work this paper cites.
Exploiting bounded staleness to speed up big data analytics
H. Cui, J. Cipar, Q. Ho, J. K. Kim, S. Lee, A. Kumar, J. Wei, W. Dai, G. R. Ganger, P. B. Gibbons, G. A. Gibson, and E. P. Xing · 2014
Earlier work this paper cites.
Caffe: Convolutional architecture for fast feature embedding
Y. Jia, E. Shelhamer, J. Donahue, S. Karayev, J. Long, R. B. Girshick, S. Guadarrama, and T. Darrell · 2014
Earlier work this paper cites.
One weird trick for parallelizing convolutional neural networks
A. Krizhevsky · 2014
Earlier work this paper cites.
Scaling Distributed Machine Learning with the Parameter Server
M. Li, D. G. Andersen, J. W. Park, A. J. Smola, A. Ahmed, V. Josifovski, J. Long, E. J. Shekita, and B.-Y. Su · 2014
Cited alongside, same era.
Efficient mini-batch training for stochastic optimization
M. Li, T. Zhang, Y. Chen, and A. J. Smola · 2014
Cited alongside, same era.
Dogwild! – Distributed Hogwild for CPU and GPU
C. Noel and S. Osindero · 2014
Cited alongside, same era.
DimmWitted: A Study of Main-memory Statistical Analytics
C. Zhang and C. Ré · 2014
Cited alongside, same era.
Asynchronous stochastic convex optimization: the noise is in the noise and sgd don't care
S. Chaturapruek, J. C. Duchi, and C. Ré · 2015
Cited alongside, same era.
MXNet: A flexible and efficient machine learning library for heterogeneous distributed systems
The microsoft 2016 conversational speech recognition system
W. Xiong, J. Droppo, X. Huang, F. Seide, M. Seltzer, A. Stolcke, D. Yu, and G. Zweig · 2017
Later among the works it cites.
Scaling SGD batch size to 32K for ImageNet training
J. Zhang and I. Mitliagkas · 2017
Later among the works it cites.
YellowFin and the art of momentum tuning
J. Zhang and I. Mitliagkas · 2017
Later among the works it cites.
Optimization methods for large-scale machine learning
L. Bottou, F. Curtis, and J. Nocedal · 2018
Later among the works it cites.
A new golden age in computer architecture: Empowering the machine-learning revolution
J. Dean, D. Patterson, and C. Young · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
T. Chen, M. Li, Y. Li, M. Lin, N. Wang, M. Wang, T. Xiao, B. Xu, C. Zhang, and Z. Zhang · 2015
Cited alongside, same era.
The loss surfaces of multilayer networks
A. Choromanska, M. Henaff, M. Mathieu, G. Ben Arous, and Y. LeCun · 2015
Cited alongside, same era.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2015
Cited alongside, same era.
Tensorflow: A system for large-scale machine learning
M. Abadi, P. Barham, J. Chen, Z. Chen, A. Davis, J. Dean, M. Devin, S. Ghemawat, G. Irving, M. Isard, M. Kudlur, J. Levenberg, R. Monga, S. Moore, D. G. Murray, B. Steiner, P. Tucker, V. Vasudevan, P. Warden, M. Wicke, Y. Yu, and X. Zheng · 2016
Cited alongside, same era.
Revisiting distributed synchronous SGD
J. Chen, R. Monga, S. Bengio, and R. Józefowicz · 2016
Cited alongside, same era.
GeePS: Scalable deep learning on distributed GPUs with a GPU-specialized parameter server
H. Cui, H. Zhang, G. R. Ganger, P. B. Gibbons, and E. P. Xing · 2016
Cited alongside, same era.
Addressing the straggler problem for iterative convergent parallel ML
A. Harlap, H. Cui, W. Dai, J. Wei, G. R. Ganger, P. B. Gibbons, G. A. Gibson, and E. P. Xing · 2016
Cited alongside, same era.
FlexPS: Flexible parallelism control in parameter server architecture
Y. Huang, T. Jin, Y. Wu, Z. Cai, X. Yan, F. Yang, J. Li, Y. Guo, and J. Cheng · 2018
Later among the works it cites.
X. Jia, S. Song, W. He, Y. Wang, H. Rong, F. Zhou, L. Xie, Z. Guo, Y. Yang, L. Yu, T. Chen, G. Hu, S. Shi, and X. Chu · 2018
Later among the works it cites.
Exploring hidden dimensions in accelerating convolutional neural networks
Z. Jia, S. Lin, C. R. Qi, and A. Aiken · 2018
Later among the works it cites.
Asynchronous decentralized parallel stochastic gradient descent
X. Lian, W. Zhang, C. Zhang, and J. Liu · 2018
Later among the works it cites.
Revisiting small batch training for deep neural networks
D. Masters and C. Luschi · 2018
Later among the works it cites.
Ray: A distributed framework for emerging AI applications
P. Moritz, R. Nishihara, S. Wang, A. Tumanov, R. Liaw, E. Liang, M. Elibol, Z. Yang, W. Paul, M. I. Jordan, and I. Stoica · 2018
Later among the works it cites.
Accelerating model search with model batching
D. Narayanan, K. Santhanam, and M. Zaharia · 2018
Later among the works it cites.
https://developer.nvidia.com/nccl
NVIDIA Collective Communications Library (NCCL), 2018 · 2018
Later among the works it cites.
https://www.nvidia.com/en-us/data-center/nvlink/
NVLINK FABRIC, 2018 · 2018
Later among the works it cites.
https://www.microway.com/product/ octoputer-4u-10-gpu-server-single-root-complex/
Octoputer 4U 10-GPU Server with Single Root Complex for GPU-Direct, 2018 · 2018
Later among the works it cites.
https://pytorch.org
PyTorch, 2018 · 2018
Later among the works it cites.
Litz: Elastic framework for high-performance distributed machine learning
A. Qiao, A. Aghayev, W. Yu, H. Chen, Q. Ho, G. A. Gibson, and E. P. Xing · 2018
Later among the works it cites.
Horovod: fast and easy distributed deep learning in TensorFlow
A. Sergeev and M. D. Balso · 2018
Later among the works it cites.
Measuring the effects of data parallelism on neural network training
C. J. Shallue, J. Lee, J. Antognini, J. Sohl-Dickstein, R. Frostig, and G. E. Dahl · 2018
Later among the works it cites.
https://github.com/tensorflow/benchmarks
TensorFlow Benchmarks, 2018 · 2018
Later among the works it cites.
https://github.com/geifmany/cifar-vgg
VGG16 models for CIFAR-10 and CIFAR-100 using Keras, 2018 · 2018
Later among the works it cites.
Superneurons: dynamic GPU memory management for training deep neural networks
L. Wang, J. Ye, Y. Zhao, W. Wu, A. Li, S. L. Song, Z. Xu, and T. Kraska · 2018
Later among the works it cites.