Fetching the paper…
Reading the bibliography…
Many distributed machine learning (ML) systems adopt the non-synchronous execution in order to alleviate the network communication bottleneck, resulting in stale parameters that do not reflect the latest updates.
The mnist database of handwritten digits
Yann LeCun · 1998
Earlier work this paper cites.
On the momentum term in gradient descent learning algorithms
Ning Qian · 1999
Earlier work this paper cites.
Finding scientific topics
Thomas L. Griffiths and Mark Steyvers · 2004
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky and Geoffrey Hinton · 2009
Earlier work this paper cites.
Slow learners are fast
John Langford, Alexander J Smola, and Martin Zinkevich · 2009
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
Xavier Glorot and Yoshua Bengio · 2010
Earlier work this paper cites.
Rectified linear units improve restricted boltzmann machines
Vinod Nair and Geoffrey E Hinton · 2010
Earlier work this paper cites.
An architecture for parallel topic models
Alexander Smola and Shravan Narayanamurthy · 2010
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
John Duchi, Elad Hazan, and Yoram Singer · 2011
Earlier work this paper cites.
Hogwild: A lock-free approach to parallelizing stochastic gradient descent
Benjamin Recht, Christopher Re, Stephen J. Wright, and Feng Niu · 2011
Earlier work this paper cites.
Scalable inference in latent variable models
Amr Ahmed, Mohamed Aly, Joseph Gonzalez, Shravan Narayanamurthy, and Alexander J. Smola · 2012
Earlier work this paper cites.
Large scale distributed deep networks
Jeffrey Dean, Greg Corrado, Rajat Monga, Kai Chen, Matthieu Devin, Mark Mao, Andrew Senior, Paul Tucker, Ke Yang, Quoc V Le, et al · 2012
Earlier work this paper cites.
Neural networks for machine learning
Geoffrey Hinton · 2012
Earlier work this paper cites.
Distributed graphlab: A framework for machine learning and data mining in the cloud
Yucheng Low, Joseph Gonzalez, Aapo Kyrola, Danny Bickson, Carlos Guestrin, and Joseph M. Hellerstein · 2012
Earlier work this paper cites.
Deep learning with cots hpc systems
Adam Coates, Brody Huval, Tao Wang, David Wu, Bryan Catanzaro, and Ng Andrew · 2013
Cited alongside, same era.
More effective distributed ml via a stale synchronous parallel parameter server
Qirong Ho, James Cipar, Henggang Cui, Seunghak Lee, Jin Kyu Kim, Phillip B. Gibbons, Garth A. Gibson, Greg Ganger, and Eric Xing · 2013
Cited alongside, same era.
On statistics, computation and scalability
Michael I Jordan et al · 2013
Cited alongside, same era.
Asymptotically exact, embarrassingly parallel mcmc
Willie Neiswanger, Chong Wang, and Eric Xing · 2013
Cited alongside, same era.
NOMAD: Non-locking, stochastic multi-machine algorithm for asynchronous and decentralized matrix completion
Hyokun Yun, Hsiang-Fu Yu, Cho-Jui Hsieh, SVN Vishwanathan, and Inderjit Dhillon · 2013
Cited alongside, same era.
Managed communication and consistency for fast data-parallel iterative analytics
Jinliang Wei, Wei Dai, Aurick Qiao, Qirong Ho, Henggang Cui, Gregory R Ganger, Phillip B Gibbons, Garth A Gibson, and Eric P Xing · 2015
Later among the works it cites.
Lightlda: Big topic models on modest computer clusters
Jinhui Yuan, Fei Gao, Qirong Ho, Wei Dai, Jinliang Wei, Xun Zheng, Eric Po Xing, Tie-Yan Liu, and Wei-Ying Ma · 2015
Later among the works it cites.
Revisiting distributed synchronous SGD
Jianmin Chen, Rajat Monga, Samy Bengio, and Rafal Jozefowicz · 2016
Later among the works it cites.
Geeps: Scalable deep learning on distributed gpus with a gpu-specialized parameter server
Henggang Cui, Hao Zhang, Gregory R Ganger, Phillip B Gibbons, and Eric P Xing · 2016
Later among the works it cites.
Omnivore: An optimizer for multi-device deep learning on cpus and gpus
Stefan Hadjis, Ce Zhang, Ioannis Mitliagkas, Dan Iter, and Christopher Ré · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Project adam: Building an efficient and scalable deep learning training system
Trishul Chilimbi, Yutaka Suzue, Johnson Apacible, and Karthik Kalyanaraman · 2014
Cited alongside, same era.
Exploiting bounded staleness to speed up big data analytics
Henggang Cui, James Cipar, Qirong Ho, Jin Kyu Kim, Seunghak Lee, Abhimanu Kumar, Jinliang Wei, Wei Dai, Gregory R. Ganger, Phillip B. Gibbons, Garth A. Gibson, and Eric P. Xing · 2014
Cited alongside, same era.
Adam: A method for stochastic optimization
Diederik Kingma and Jimmy Ba · 2014
Cited alongside, same era.
Fugue: Slow-worker-agnostic distributed learning for big models on big data
Abhimanu Kumar, Alex Beutel, Qirong Ho, and Eric P Xing · 2014
Cited alongside, same era.
Delay-tolerant algorithms for asynchronous distributed online learning
Brendan McMahan and Matthew Streeter · 2014
Cited alongside, same era.
Asynchronous distributed admm for consensus optimization
Ruiliang Zhang and James T. Kwok · 2014
Cited alongside, same era.
Analysis of high-performance distributed ml at scale through parameter server consistency models
Wei Dai, Abhimanu Kumar, Jinliang Wei, Qirong Ho, Garth Gibson, and Eric P. Xing · 2015
Cited alongside, same era.
The movielens datasets: History and context
F Maxwell Harper and Joseph A Konstan · 2016
Later among the works it cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Later among the works it cites.
On large-batch training for deep learning: Generalization gap and sharp minima
Nitish Shirish Keskar, Dheevatsa Mudigere, Jorge Nocedal, Mikhail Smelyanskiy, and Ping Tak Peter Tang · 2016
Later among the works it cites.
Strads: a distributed framework for scheduled model parallel machine learning
Jin Kyu Kim, Qirong Ho, Seunghak Lee, Xun Zheng, Wei Dai, Garth A Gibson, and Eric P Xing · 2016
Later among the works it cites.
Asynchrony begets momentum, with an application to deep learning
Ioannis Mitliagkas, Ce Zhang, Stefan Hadjis, and Christopher Ré · 2016
Later among the works it cites.
Bayes and big data: The consensus monte carlo algorithm
Steven L Scott, Alexander W Blocker, Fernando V Bonassi, Hugh A Chipman, Edward I George, and Robert E McCulloch · 2016
Later among the works it cites.
Accurate, large minibatch SGD: training imagenet in 1 hour
Priya Goyal, Piotr Dollár, Ross Girshick, Pieter Noordhuis, Lukasz Wesolowski, Aapo Kyrola, Andrew Tulloch, Yangqing Jia, and Kaiming He · 2017
Later among the works it cites.
Revisiting small batch training for deep neural networks
Dominic Masters and Carlo Luschi · 2018
Closest in time.