Fetching the paper…
Reading the bibliography…
We study the asynchronous stochastic gradient descent algorithm for distributed training over $n$ workers which have varying computation and communication frequency over time.
A Stochastic Approximation Method
Herbert Robbins and Sutton Monro · 1951
Earlier work this paper cites.
Parallel and Distributed Computation: Numerical Methods
D.P. Bertsekas and J.N. Tsitsiklis · 1989
Earlier work this paper cites.
Backpropagation convergence via deterministic nonmonotone perturbed minimization
Olvi L. Mangasarian and Mikhail V. Solodov · 1994
Earlier work this paper cites.
Distributed training strategies for the structured perceptron
Ryan McDonald, Keith Hall, and Gideon Mann · 2010
Earlier work this paper cites.
Parallelized stochastic gradient descent
Martin Zinkevich, Markus Weimer, Lihong Li, and Alex J Smola · 2010
Earlier work this paper cites.
Distributed delayed stochastic optimization
Alekh Agarwal and John C Duchi · 2011
Earlier work this paper cites.
HOGWILD!: A lock-free approach to parallelizing stochastic gradient descent
Feng Niu, Benjamin Recht, Christopher Re, and Stephen J. Wright · 2011
Earlier work this paper cites.
Hogwild: A lock-free approach to parallelizing stochastic gradient descent
Benjamin Recht, Christopher Re, Stephen Wright, and Feng Niu · 2011
Earlier work this paper cites.
Large scale distributed deep networks
Jeffrey Dean, Greg Corrado, Rajat Monga, Kai Chen, Matthieu Devin, Mark Mao, Marc’aurelio Ranzato, Andrew Senior, Paul Tucker, Ke Yang, Quoc V. Le, and Andrew Y. Ng · 2012
Earlier work this paper cites.
Optimal distributed online prediction using mini-batches
Ofer Dekel, Ran Gilad-Bachrach, Ohad Shamir, and Lin Xiao · 2012
Earlier work this paper cites.
Stochastic first- and zeroth-order methods for nonconvex stochastic programming
Saeed Ghadimi and Guanghui Lan · 2013
Earlier work this paper cites.
Delay-tolerant algorithms for asynchronous distributed online learning
Brendan McMahan and Matthew Streeter · 2014
Earlier work this paper cites.
Asynchronous stochastic convex optimization: the noise is in the noise and SGD don’t care
Sorathan Chaturapruek, John C Duchi, and Christopher Ré · 2015
Earlier work this paper cites.
Asynchronous parallel stochastic gradient for nonconvex optimization
Xiangru Lian, Yijun Huang, Yuncheng Li, and Ji Liu · 2015
Earlier work this paper cites.
Arda Aytekin, Hamid Reza Feyzmahdavian, and Mikael Johansson · 2016
Earlier work this paper cites.
An asynchronous mini-batch algorithm for regularized stochastic optimization
H. R. Feyzmahdavian, A. Aytekin, and M. Johansson · 2016
Earlier work this paper cites.
Federated learning of deep networks using model averaging
H. Brendan McMahan, Eider Moore, Daniel Ramage, and Blaise Agüera y Arcas · 2016
Earlier work this paper cites.
Adadelay: Delay adaptive distributed stochastic optimization
Suvrit Sra, Adams Wei Yu, Mu Li, and Alex Smola · 2016
Earlier work this paper cites.
Staleness-aware async-sgd for distributed deep learning
Wei Zhang, Suyog Gupta, Xiangru Lian, and Ji Liu · 2016
Cited alongside, same era.
QSGD: Communication-efficient SGD via gradient quantization and encoding
Dan Alistarh, Demjan Grubic, Jerry Li, Ryota Tomioka, and Milan Vojnovic · 2017
Cited alongside, same era.
Can decentralized algorithms outperform centralized algorithms? a case study for decentralized parallel stochastic gradient descent
Xiangru Lian, Ce Zhang, Huan Zhang, Cho-Jui Hsieh, Wei Zhang, and Ji Liu · 2017
Cited alongside, same era.
Perturbed iterate analysis for asynchronous stochastic optimization
Horia Mania, Xinghao Pan, Dimitris Papailiopoulos, Benjamin Recht, Kannan Ramchandran, and Michael I. Jordan · 2017
Cited alongside, same era.
Asynchronous stochastic gradient descent with delay compensation
Shuxin Zheng, Qi Meng, Taifeng Wang, Wei Chen, Nenghai Yu, Zhi-Ming Ma, and Tie-Yan Liu · 2017
Cited alongside, same era.
A unified theory of decentralized sgd with changing topology and local updates
Anastasia Koloskova, Nicolas Loizou, Sadra Boreiri, Martin Jaggi, and Sebastian U. Stich · 2020
Later among the works it cites.
Distributed gradient methods for convex machine learning problems in networks: Distributed optimization
Angelia Nedić · 2020
Later among the works it cites.
The error-feedback framework: SGD with delayed gradients
Sebastian U. Stich and Sai Praneeth Karimireddy · 2020
Later among the works it cites.
A survey on large-scale machine learning
Meng Wang, Weijie Fu, Xiangnan He, Shijie Hao, and Xindong Wu · 2020
Later among the works it cites.
Yikai Yan, Chaoyue Niu, Yucheng Ding, Zhenzhe Zheng, Fan Wu, Guihai Chen, Shaojie Tang, and Zhihua Wu · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
The convergence of sparsified gradient methods
Dan Alistarh, Torsten Hoefler, Mikael Johansson, Nikola Konstantinov, Sarit Khirirat, and Cedric Renggli · 2018
Cited alongside, same era.
Optimization methods for large-scale machine learning
L. Bottou, F. Curtis, and J. Nocedal · 2018
Cited alongside, same era.
Slow and stale gradients can win the race: Error-runtime trade-offs in distributed sgd
Sanghamitra Dutta, Gauri Joshi, Soumyadip Ghosh, Parijat Dube, and Priya Nagpurkar · 2018
Cited alongside, same era.
Improved asynchronous parallel optimization analysis for stochastic incremental methods
Remi Leblond, Fabian Pedregosa, and Simon Lacoste-Julien · 2018
Cited alongside, same era.
Lower bounds for non-convex stochastic optimization
Yossi Arjevani, Yair Carmon, John C. Duchi, Dylan J. Foster, Nathan Srebro, and Blake E. Woodworth · 2019
Cited alongside, same era.
Stochastic gradient push for distributed deep learning
Mahmoud Assran, Nicolas Loizou, Nicolas Ballas, and Michael Rabbat · 2019
Cited alongside, same era.
Towards federated learning at scale: System design
Keith Bonawitz, Hubert Eichner, Wolfgang Grieskamp, Dzmitry Huba, Alex Ingerman, Vladimir Ivanov, Chloé Kiddon, Jakub Konečný, Stefano Mazzocchi, Brendan McMahan, Timon Van Overveldt, David Petrou, Daniel Ramage, and Jason Roselander · 2019
Cited alongside, same era.
Federated learning under arbitrary communication patterns
Dmitrii Avdiukhin and Shiva Kasiviswanathan · 2021
Later among the works it cites.
Learning under delayed feedback: Implicitly adapting to gradient delays
Rotem Zamir Aviv, Ido Hakimi, Assaf Schuster, and Kfir Yehuda Levy · 2021
Later among the works it cites.
Asynchronous stochastic optimization robust to arbitrary delays
Alon Cohen, Amit Daniely, Yoel Drori, Tomer Koren, and Mariano Schain · 2021
Later among the works it cites.
Decentralized optimization with heterogeneous delays: a continuous-time approach , 2021
Mathieu Even, Hadrien Hendrikx, and Laurent Massoulie · 2021
Later among the works it cites.
Fast federated learning in the presence of arbitrary device unavailability
Xinran Gu, Kaixuan Huang, Jingzhao Zhang, and Longbo Huang · 2021
Later among the works it cites.
Zero-shot text-to-image generation
Aditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray, Chelsea Voss, Alec Radford, Mark Chen, and Ilya Sutskever · 2021
Later among the works it cites.
Critical parameters for scalable distributed learning with large batches and asynchronous updates
Sebastian Stich, Amirkeivan Mohtashami, and Martin Jaggi · 2021
Later among the works it cites.
Anarchic federated learning , 2021
Haibo Yang, Xin Zhang, Prashant Khanduri, and Jia Liu · 2021
Later among the works it cites.
Asynchronous SGD beats minibatch SGD under arbitrary delays , 2022
Konstantin Mishchenko, Francis Bach, Mathieu Even, and Blake Woodworth · 2022
Closest in time.
Federated learning with buffered asynchronous aggregation
John Nguyen, Kshitiz Malik, Hongyua Zhan, Ashka Yousefpour, Mike Rabbat, Mani Malek, and Dzmitry Huba · 2022
Closest in time.
Hierarchical text-conditional image generation with clip latents
Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen · 2022
Closest in time.
Delay-adaptive step-sizes for asynchronous learning , 2022
Xuyang Wu, Sindri Magnusson, Hamid Reza Feyzmahdavian, and Mikael Johansson · 2022
Closest in time.