Fetching the paper…
Reading the bibliography…
Distributed training of deep learning models on large-scale training data is typically conducted with asynchronous stochastic optimization to maximize the rate of updates, at the cost of additional noise introduced from asynchrony.
Learning multiple layers of features from tiny images
Alex Krizhevsky and Geoffrey Hinton · 2009
Earlier work this paper cites.
Parallelized stochastic gradient descent
Martin Zinkevich, Markus Weimer, Lihong Li, and Alex J Smola · 2010
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
John Duchi, Elad Hazan, and Yoram Singer · 2011
Earlier work this paper cites.
Hogwild: A lock-free approach to parallelizing stochastic gradient descent
Benjamin Recht, Christopher Re, Stephen Wright, and Feng Niu · 2011
Earlier work this paper cites.
Large scale distributed deep networks
J. Dean, G. S. Corrado, R. Monga, K. Chen, M. Devin, Q. V. Le, M. Z. Mao, M. A. Ranzato, A. Senior, P. Tucker, K. Yang, and A. Y. Ng · 2012
Earlier work this paper cites.
Deep neural networks for acoustic modeling in speech recognition
G. Hinton, L. Deng, D. Yu, G. Dahl, A. Mohamed, N. Jaitly, A. Senior, V. Vanhoucke, P. Nguyen, T. N. Sainath, and B. Kingsbury · 2012
Earlier work this paper cites.
Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude
Tijmen Tieleman and Geoffrey Hinton · 2012
Earlier work this paper cites.
The tail at scale
Jeffrey Dean and Luiz André Barroso · 2013
Earlier work this paper cites.
Estimation, optimization, and parallelism when data is sparse
John Duchi, Michael I Jordan, and Brendan McMahan · 2013
Earlier work this paper cites.
Project adam: Building an efficient and scalable deep learning training system
T. Chilimbi, Y. Suzue, J. Apacible, and K. Kalyanaraman · 2014
Cited alongside, same era.
Adam: A method for stochastic optimization
Diederik Kingma and Jimmy Ba · 2014
Cited alongside, same era.
Scaling distributed machine learning with the parameter server
Mu Li, David G Andersen, Jun Woo Park, Alexander J Smola, Amr Ahmed, Vanja Josifovski, James Long, Eugene J Shekita, and Bor-Yiing Su · 2014
Cited alongside, same era.
TensorFlow: Large-scale machine learning on heterogeneous systems, 2015
Martín Abadi, Ashish Agarwal, Paul Barham, Eugene Brevdo, Zhifeng Chen, Craig Citro, Greg S. Corrado, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Ian Goodfellow, Andrew Harp, Geoffrey Irving, Michael Isard, Yangqing Jia, Rafal Jozefowicz, Lukasz Kaiser, Manjunath Kudlur, Josh Levenberg, Dan Mané, Rajat Monga, Sherry Moore, Derek Murray, Chris Olah, Mike Schuster, Jonathon Shlens, Benoit Steiner, Ilya Sutskever, Kunal Talwar, Paul Tucker, Vincent Vanhoucke, Vijay Vasudevan, Fernanda Viégas, Oriol Vinyals, Pete Warden, Martin Wattenberg, Martin Wicke, Yuan Yu, and Xiaoqiang Zheng · 2015
Cited alongside, same era.
Petuum: A new platform for distributed machine learning on big data
Eric P Xing, Qirong Ho, Wei Dai, Jin Kyu Kim, Jinliang Wei, Seunghak Lee, Xun Zheng, Pengtao Xie, Abhimanu Kumar, and Yaoliang Yu · 2015
Later among the works it cites.
Splash: User-friendly programming interface for parallelizing stochastic algorithms
Yuchen Zhang and Michael I Jordan · 2015
Later among the works it cites.
Revisiting distributed synchronous sgd
Jianmin Chen, Rajat Monga, Samy Bengio, and Rafal Jozefowicz · 2016
Closest in time.
Distributed deep learning using synchronous stochastic gradient descent
Dipankar Das, Sasikanth Avancha, Dheevatsa Mudigere, Karthikeyan Vaidynathan, Srinivas Sridharan, Dhiraj Kalamkar, Bharat Kaul, and Pradeep Dubey · 2016
Closest in time.
On large-batch training for deep learning: Generalization gap and sharp minima
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Christopher M De Sa, Ce Zhang, Kunle Olukotun, and Christopher Ré · 2015
Cited alongside, same era.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
S. Ioffe and C. Szegedy · 2015
Cited alongside, same era.
Perturbed iterate analysis for asynchronous stochastic optimization
Horia Mania, Xinghao Pan, Dimitris Papailiopoulos, Benjamin Recht, Kannan Ramchandran, and Michael I Jordan · 2015
Cited alongside, same era.
On variance reduction in stochastic gradient descent and its asynchronous variants
Sashank J Reddi, Ahmed Hefny, Suvrit Sra, Barnabas Poczos, and Alex J Smola · 2015
Cited alongside, same era.
Imagenet large scale visual recognition challenge
Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, Alexander C. Berg, and Li Fei-Fei · 2015
Cited alongside, same era.
Deep learning with elastic averaging sgd
Sixin Zhang, Anna E Choromanska, and Yann LeCun
Cited in the paper.
Staleness-aware async-sgd for distributed deep learning
Wei Zhang, Suyog Gupta, Xiangru Lian, and Ji Liu
Cited in the paper.
Nitish Shirish Keskar, Dheevatsa Mudigere, Jorge Nocedal, Mikhail Smelyanskiy, and Ping Tak Peter Tang · 2016
Closest in time.
Asaga: Asynchronous parallel saga
Rémi Leblond, Fabian Pedregosa, and Simon Lacoste-Julien · 2016
Closest in time.
Conditional image generation with pixelcnn decoders
Aaron van den Oord, Nal Kalchbrenner, Oriol Vinyals, Lasse Espeholt, Alex Graves, and Koray Kavukcuoglu · 2016
Closest in time.
Rethinking the inception architecture for computer vision
C. Szegedy, V. Vanhoucke, S. Ioffe, J. Shlens, and Z. Wojna · 2016
Closest in time.
Ako: Decentralised deep learning with partial gradient exchange
Pijika Watcharapichat, Victoria Lopez Morales, Raul Castro Fernandez, and Peter Pietzuch · 2016
Closest in time.