Fetching the paper…
Reading the bibliography…
The existing analysis of asynchronous stochastic gradient descent (SGD) degrades dramatically when any delay is large, giving the impression that performance depends primarily on the delay.
Asynchronous iterative methods for multiprocessors
Gerard M. Baudet · 1978
Earlier work this paper cites.
Problem complexity and method efficiency in optimization
Arkadii Semenovich Nemirovsky and David Borisovich Yudin · 1983
Earlier work this paper cites.
Distributed asynchronous deterministic and stochastic gradient optimization algorithms
John Tsitsiklis, Dimitri Bertsekas, and Michael Athans · 1986
Earlier work this paper cites.
Distributed asynchronous incremental subgradient methods
Angelia Nedić, Dimitri P. Bertsekas, and Vivek S. Borkar · 2001
Earlier work this paper cites.
Distributed subgradient methods for multi-agent optimization
Angelia Nedić and Asuman Ozdaglar · 2009
Earlier work this paper cites.
Parallelized stochastic gradient descent
Martin Zinkevich, Markus Weimer, Lihong Li, and Alex J. Smola · 2010
Earlier work this paper cites.
Distributed delayed stochastic optimization
Alekh Agarwal and John C. Duchi · 2011
Earlier work this paper cites.
Better mini-batch algorithms via accelerated gradient methods
Andrew Cotter, Ohad Shamir, Nati Srebro, and Karthik Sridharan · 2011
Earlier work this paper cites.
Hogwild!: A lock-free approach to parallelizing stochastic gradient descent
Benjamin Recht, Christopher Re, Stephen Wright, and Feng Niu · 2011
Earlier work this paper cites.
Large scale distributed deep networks
Jeffrey Dean, Greg Corrado, Rajat Monga, Kai Chen, Matthieu Devin, Mark Mao, Marc’aurelio Ranzato, Andrew Senior, Paul Tucker, Ke Yang, et al · 2012
Earlier work this paper cites.
Optimal distributed online prediction using mini-batches
Ofer Dekel, Ran Gilad-Bachrach, Ohad Shamir, and Lin Xiao · 2012
Earlier work this paper cites.
Stochastic first-and zeroth-order methods for nonconvex stochastic programming
Saeed Ghadimi and Guanghui Lan · 2013
Earlier work this paper cites.
Delay-tolerant algorithms for asynchronous distributed online learning
Brendan McMahan and Matthew J. Streeter · 2014
Earlier work this paper cites.
Convex optimization: Algorithms and complexity
Sébastien Bubeck · 2015
Earlier work this paper cites.
Asynchronous parallel stochastic gradient for nonconvex optimization
Xiangru Lian, Yijun Huang, Yuncheng Li, and Ji Liu · 2015
Earlier work this paper cites.
Massively parallel methods for deep reinforcement learning
Arun Nair, Praveen Srinivasan, Sam Blackwell, Cagdas Alcicek, Rory Fearon, Alessandro De Maria, Vedavyas Panneershelvam, Mustafa Suleyman, Charles Beattie, Stig Petersen, et al · 2015
Earlier work this paper cites.
Revisiting distributed synchronous SGD
Jianmin Chen, Xinghao Pan, Rajat Monga, Samy Bengio, and Rafal Jozefowicz · 2016
Earlier work this paper cites.
Delay-tolerant online convex optimization: Unified analysis and adaptive-gradient algorithms
Pooria Joulani, András György, and Csaba Szepesvári · 2016
Earlier work this paper cites.
Asynchronous methods for deep reinforcement learning
Volodymyr Mnih, Adria Puigdomenech Badia, Mehdi Mirza, Alex Graves, Timothy Lillicrap, Tim Harley, David Silver, and Koray Kavukcuoglu · 2016
Earlier work this paper cites.
Adadelay: Delay adaptive distributed stochastic optimization
Suvrit Sra, Adams Wei Yu, Mu Li, and Alexander J. Smola · 2016
Cited alongside, same era.
Qsgd: Communication-efficient sgd via gradient quantization and encoding
Dan Alistarh, Demjan Grubic, Jerry Li, Ryota Tomioka, and Milan Vojnovic · 2017
Cited alongside, same era.
Lower bounds for finding stationary points I
Yair Carmon, John C. Duchi, Oliver Hinder, and Aaron Sidford · 2017
Cited alongside, same era.
Accurate, large minibatch SGD: Training imagenet in 1 hour
Priya Goyal, Piotr Dollár, Ross Girshick, Pieter Noordhuis, Lukasz Wesolowski, Aapo Kyrola, Andrew Tulloch, Yangqing Jia, and Kaiming He · 2017
Cited alongside, same era.
Can decentralized algorithms outperform centralized algorithms? a case study for decentralized parallel stochastic gradient descent
Xiangru Lian, Ce Zhang, Huan Zhang, Cho-Jui Hsieh, Wei Zhang, and Ji Liu · 2017
Advances and open problems in federated learning, 2019
Peter Kairouz, H. Brendan McMahan, Brendan Avent, Aurélien Bellet, Mehdi Bennis, Arjun Nitin Bhagoji, Keith Bonawitz, Zachary Charles, Graham Cormode, Rachel Cummings, Rafael G. L. D’Oliveira, Salim El Rouayheb, David Evans, Josh Gardner, Zachary Garrett, Adrià Gascón, Badih Ghazi, Phillip B. Gibbons, Marco Gruteser, Zaid Harchaoui, Chaoyang He, Lie He, Zhouyuan Huo, Ben Hutchinson, Justin Hsu, Martin Jaggi, Tara Javidi, Gauri Joshi, Mikhail Khodak, Jakub Konečný, Aleksandra Korolova, Farinaz Koushanfar, Sanmi Koyejo, Tancrède Lepoint, Yang Liu, Prateek Mittal, Mehryar Mohri, Richard Nock, Ayfer Özgür, Rasmus Pagh, Mariana Raykova, Hang Qi, Daniel Ramage, Ramesh Raskar, Dawn Song, Weikang Song, Sebastian U. Stich, Ziteng Sun, Ananda Theertha Suresh, Florian Tramèr, Praneeth Vepakomma, Jianyu Wang, Li Xiong, Zheng Xu, Qiang Yang, Felix X. Yu, Han Yu, and Sen Zhao · 2019
Later among the works it cites.
A tight convergence analysis for stochastic gradient descent with delayed updates
Yossi Arjevani, Ohad Shamir, and Nathan Srebro · 2020
Later among the works it cites.
Advances in asynchronous parallel and distributed optimization
Mahmoud Assran, Arda Aytekin, Hamid Reza Feyzmahdavian, Mikael Johansson, and Michael G. Rabbat · 2020
Later among the works it cites.
Better theory for SGD in the nonconvex world
Ahmed Khaled and Peter Richtárik · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Perturbed iterate analysis for asynchronous stochastic optimization
Horia Mania, Xinghao Pan, Dimitris Papailiopoulos, Benjamin Recht, Kannan Ramchandran, and Michael I. Jordan · 2017
Cited alongside, same era.
Communication-efficient learning of deep networks from decentralized data
Brendan McMahan, Eider Moore, Daniel Ramage, Seth Hampson, and Blaise Aguera y Arcas · 2017
Cited alongside, same era.
Asynchronous stochastic gradient descent with delay compensation
Shuxin Zheng, Qi Meng, Taifeng Wang, Wei Chen, Nenghai Yu, Zhi-Ming Ma, and Tie-Yan Liu · 2017
Cited alongside, same era.
signSGD: Compressed optimisation for non-convex problems
Jeremy Bernstein, Yu-Xiang Wang, Kamyar Azizzadenesheli, and Animashree Anandkumar · 2018
Cited alongside, same era.
Improved asynchronous parallel optimization analysis for stochastic incremental methods
Rémi Leblond, Fabian Pedregosa, and Simon Lacoste-Julien · 2018
Cited alongside, same era.
A delay-tolerant proximal-gradient algorithm for distributed learning
Konstantin Mishchenko, Franck Iutzeler, Jérôme Malick, and Massih-Reza Amini · 2018
Cited alongside, same era.
Ray: A distributed framework for emerging AI applications
Philipp Moritz, Robert Nishihara, Stephanie Wang, Alexey Tumanov, Richard Liaw, Eric Liang, Melih Elibol, Zongheng Yang, William Paul, Michael I. Jordan, and Ion Stoica · 2018
Cited alongside, same era.
Tighter theory for local SGD on identical and heterogeneous data
Ahmed Khaled, Konstantin Mishchenko, and Peter Richtárik · 2020
Later among the works it cites.
The error-feedback framework: Better rates for SGD with delayed gradients and compressed updates
Sebastian U. Stich and Sai Praneeth Karimireddy · 2020
Later among the works it cites.
Is local SGD better than minibatch SGD?
Blake Woodworth, Kumar Kshitij Patel, Sebastian Stich, Zhen Dai, Brian Bullins, Brendan Mcmahan, Ohad Shamir, and Nathan Srebro · 2020
Later among the works it cites.
Asynchronous distributed learning: Adapting to gradient delays without prior knowledge
Rotem Zamir Aviv, Ido Hakimi, Assaf Schuster, and Kfir Yehuda Levy · 2021
Later among the works it cites.
FedAT: a high-performance and communication-efficient federated learning system with asynchronous tiers
Zheng Chai, Yujing Chen, Ali Anwar, Liang Zhao, Yue Cheng, and Huzefa Rangwala · 2021
Later among the works it cites.
Asynchronous stochastic optimization robust to arbitrary delays
Alon Cohen, Amit Daniely, Yoel Drori, Tomer Koren, and Mariano Schain · 2021
Later among the works it cites.
Decentralized optimization with heterogeneous delays: a continuous-time approach
Mathieu Even, Hadrien Hendrikx, and Laurent Massoulié · 2021
Later among the works it cites.
FjORD: Fair and accurate federated learning under heterogeneous targets with ordered dropout
Samuel Horváth, Stefanos Laskaridis, Mario Almeida, Ilias Leontiadis, Stylianos I. Venieris, and Nicholas D. Lane · 2021
Later among the works it cites.
Asynchronous decentralized SGD with quantized and local updates
Giorgi Nadiradze, Amirmojtaba Sabour, Peter Davies, Shigang Li, and Dan Alistarh · 2021
Later among the works it cites.
Moshpit SGD: Communication-efficient decentralized training on heterogeneous unreliable devices
Max Ryabinin, Eduard Gorbunov, Vsevolod Plokhotnyuk, and Gennady Pekhimenko · 2021
Later among the works it cites.
Lower bounds for non-convex stochastic optimization
Yossi Arjevani, Yair Carmon, John C. Duchi, Dylan J. Foster, Nathan Srebro, and Blake Woodworth · 2022
Closest in time.
Sharper convergence guarantees for asynchronous SGD for distributed and federated learning
Anastasia Koloskova, Sebastian U. Stich, and Martin Jaggi · 2022
Closest in time.
Federated learning with buffered asynchronous aggregation
John Nguyen, Kshitiz Malik, Hongyuan Zhan, Ashkan Yousefpour, Mike Rabbat, Mani Malek, and Dzmitry Huba · 2022
Closest in time.
Delay-adaptive step-sizes for asynchronous learning
Xuyang Wu, Sindri Magnusson, Hamid Reza Feyzmahdavian, and Mikael Johansson · 2022
Closest in time.