Fetching the paper…
Reading the bibliography…
Parallelism is a ubiquitous method for accelerating machine learning algorithms.
Distributed systems , volume 12
Sape Mullender et al · 1993
Earlier work this paper cites.
Replication management using the state-machine approach, distributed systems
Fred B Schneider · 1993
Earlier work this paper cites.
Asynchronous iterations with flexible communication for linear systems
Andreas Frommer and Daniel B Szyld · 1998
Earlier work this paper cites.
Using MPI-2: Advanced features of the message passing interface
William Gropp, Rajeev Thakur, and Ewing Lusk · 1999
Earlier work this paper cites.
Fastest mixing markov chain on a graph
Stephen Boyd, Persi Diaconis, and Lin Xiao · 2004
Earlier work this paper cites.
Solving large scale linear prediction problems using stochastic gradient descent algorithms
Tong Zhang · 2004
Earlier work this paper cites.
Asynchronous iterations with flexible communication: contracting operators
Didier El Baz, Andreas Frommer, and Pierre Spiteri · 2005
Earlier work this paper cites.
Large-scale machine learning with stochastic gradient descent
Léon Bottou · 2010
Earlier work this paper cites.
Hogwild: A lock-free approach to parallelizing stochastic gradient descent
Benjamin Recht, Christopher Re, Stephen Wright, and Feng Niu · 2011
Earlier work this paper cites.
Distributed delayed stochastic optimization
Alekh Agarwal and John C Duchi · 2011
Earlier work this paper cites.
On the importance of initialization and momentum in deep learning
Ilya Sutskever, James Martens, George Dahl, and Geoffrey Hinton · 2013
Earlier work this paper cites.
More effective distributed ml via a stale synchronous parallel parameter server
Qirong Ho, James Cipar, Henggang Cui, Seunghak Lee, Jin Kyu Kim, Phillip B Gibbons, Garth A Gibson, Greg Ganger, and Eric P Xing · 2013
Earlier work this paper cites.
Stochastic first-and zeroth-order methods for nonconvex stochastic programming
Saeed Ghadimi and Guanghui Lan · 2013
Earlier work this paper cites.
Accelerating stochastic gradient descent using predictive variance reduction
Rie Johnson and Tong Zhang · 2013
Earlier work this paper cites.
Dogwild!-distributed hogwild for cpu & gpu
Cyprien Noel and Simon Osindero · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Asynchronous parallel stochastic gradient for nonconvex optimization
Xiangru Lian, Yijun Huang, Yuncheng Li, and Ji Liu · 2015
Earlier work this paper cites.
Staleness-aware async-sgd for distributed deep learning
Wei Zhang, Suyog Gupta, Xiangru Lian, and Ji Liu · 2015
Cited alongside, same era.
Taming the wild: A unified analysis of hogwild-style algorithms
Christopher M De Sa, Ce Zhang, Kunle Olukotun, and Christopher Ré · 2015
Cited alongside, same era.
An overview of gradient descent optimization algorithms
Sebastian Ruder · 2016
Cited alongside, same era.
Revisiting distributed synchronous sgd
Jianmin Chen, Xinghao Pan, Rajat Monga, Samy Bengio, and Rafal Jozefowicz · 2016
Cited alongside, same era.
Cyclades: Conflict-free asynchronous machine learning
Xinghao Pan, Maximilian Lam, Stephen Tu, Dimitris Papailiopoulos, Ce Zhang, Michael I Jordan, Kannan Ramchandran, and Christopher Ré · 2016
Cited alongside, same era.
Distributed learning over unreliable networks
Chen Yu, Hanlin Tang, Cedric Renggli, Simon Kassing, Ankit Singla, Dan Alistarh, Ce Zhang, and Ji Liu · 2018
Later among the works it cites.
On the convergence of a class of adam-type algorithms for non-convex optimization
Xiangyi Chen, Sijia Liu, Ruoyu Sun, and Mingyi Hong · 2018
Later among the works it cites.
On the convergence of adaptive gradient methods for nonconvex optimization
Dongruo Zhou, Yiqi Tang, Ziyan Yang, Yuan Cao, and Quanquan Gu · 2018
Later among the works it cites.
Don’t use large mini-batches, use local sgd
Tao Lin, Sebastian U Stich, Kumar Kshitij Patel, and Martin Jaggi · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Unified convergence analysis of stochastic momentum methods for convex and non-convex optimization
Tianbao Yang, Qihang Lin, and Zhe Li · 2016
Cited alongside, same era.
Markov chains and mixing times , volume 107
David A Levin and Yuval Peres · 2017
Cited alongside, same era.
Understanding and optimizing asynchronous low-precision stochastic gradient descent
Christopher De Sa, Matthew Feldman, Christopher Ré, and Kunle Olukotun · 2017
Cited alongside, same era.
Pipe-sgd: A decentralized pipelined sgd framework for distributed deep net training
Youjie Li, Mingchao Yu, Songze Li, Salman Avestimehr, Nam Sung Kim, and Alexander Schwing · 2018
Cited alongside, same era.
The convergence of stochastic gradient descent in asynchronous shared memory
Dan Alistarh, Christopher De Sa, and Nikola Konstantinov · 2018
Cited alongside, same era.
D2: Decentralized training over decentralized data
Hanlin Tang, Xiangru Lian, Ming Yan, Ce Zhang, and Ji Liu · 2018
Cited alongside, same era.
Cola: Decentralized linear learning
Lie He, An Bian, and Martin Jaggi · 2018
Cited alongside, same era.
Sebastian U Stich · 2018
Later among the works it cites.
Asynchronous decentralized optimization in directed networks
Jiaqi Zhang and Keyou You · 2019
Later among the works it cites.
Dadam: A consensus-based distributed adaptive gradient method for online optimization
Parvin Nazari, Davoud Ataee Tarzanagh, and George Michailidis · 2019
Later among the works it cites.
Distributed learning over unreliable networks
Chen Yu, Hanlin Tang, Cedric Renggli, Simon Kassing, Ankit Singla, Dan Alistarh, Ce Zhang, and Ji Liu · 2019
Later among the works it cites.
Fundamental Results on Asynchronous Parallel Optimization Algorithms
Robert Rafaeil Hannah · 2019
Later among the works it cites.
Distributed learning with sublinear communication
Jayadev Acharya, Christopher De Sa, Dylan J Foster, and Karthik Sridharan · 2019
Later among the works it cites.
Error feedback fixes signsgd and other gradient compression schemes
Sai Praneeth Karimireddy, Quentin Rebjock, Sebastian U Stich, and Martin Jaggi · 2019
Later among the works it cites.
On the convergence of adam and beyond
Sashank J Reddi, Satyen Kale, and Sanjiv Kumar · 2019
Later among the works it cites.
Lower bounds for finding stationary points i
Yair Carmon, John C Duchi, Oliver Hinder, and Aaron Sidford · 2019
Later among the works it cites.
Lower bounds for non-convex stochastic optimization
Yossi Arjevani, Yair Carmon, John C Duchi, Dylan J Foster, Nathan Srebro, and Blake Woodworth · 2019
Later among the works it cites.
Elastic consistency: A general consistency model for distributed stochastic gradient descent
Dan Alistarh, Bapi Chatterjee, and Vyacheslav Kungurtsev · 2020
Closest in time.
Moniqua: Modulo quantized communication in decentralized sgd
Yucheng Lu and Christopher De Sa · 2020
Closest in time.