Fetching the paper…
Reading the bibliography…
In scalable machine learning systems, model training is often parallelized over multiple nodes that run without tight synchronization.
Introduction to optimization. optimization software
Boris T Polyak · 1987
Earlier work this paper cites.
Convergence rate and termination of asynchronous iterative algorithms
Dimitri P Bertsekas and John N Tsitsiklis · 1989
Earlier work this paper cites.
Rcv1: A new benchmark collection for text categorization research
David D Lewis, Yiming Yang, Tony Russell-Rose, and Fan Li · 2004
Earlier work this paper cites.
A convergent incremental gradient method with a constant step size
Doron Blatt, Alfred O Hero, and Hillel Gauchman · 2007
Earlier work this paper cites.
Hogwild!: A lock-free approach to parallelizing stochastic gradient descent
Benjamin Recht, Christopher Re, Stephen Wright, and Feng Niu · 2011
Earlier work this paper cites.
A stochastic gradient method with an exponential convergence rate for finite training sets
Nicolas Le Roux, Mark Schmidt, and Francis Bach · 2012
Earlier work this paper cites.
Large scale distributed deep networks
Jeffrey Dean, Greg S Corrado, Rajat Monga, Kai Chen, Matthieu Devin, Quoc V Le, Mark Z Mao, Marc’Aurelio Ranzato, Andrew Senior, Paul Tucker, et al · 2012
Earlier work this paper cites.
The mnist database of handwritten digit images for machine learning research [best of the web]
Li Deng · 2012
Earlier work this paper cites.
Parameter server for distributed machine learning
Mu Li, Li Zhou, Zichao Yang, Aaron Li, Fei Xia, David G Andersen, and Alexander Smola · 2013
Earlier work this paper cites.
An asynchronous parallel stochastic coordinate descent algorithm
Ji Liu, Steve Wright, Christopher Ré, Victor Bittorf, and Srikrishna Sridhar · 2014
Earlier work this paper cites.
Asynchronous stochastic coordinate descent: Parallelism and convergence properties
Ji Liu and Stephen J Wright · 2015
Cited alongside, same era.
ARock: an algorithmic framework for asynchronous parallel coordinate updates
Zhimin Peng, Yangyang Xu, Ming Yan, and Wotao Yin · 2016
Cited alongside, same era.
Analysis and implementation of an asynchronous optimization algorithm for the parameter server
Arda Aytekin, Hamid Reza Feyzmahdavian, and Mikael Johansson · 2016
Cited alongside, same era.
Asynchronous parallel algorithms for nonconvex big-data optimization: Model and convergence
Loris Cannelli, Francisco Facchinei, Vyacheslav Kungurtsev, and Gesualdo Scutari · 2016
Cited alongside, same era.
The asynchronous palm algorithm for nonsmooth nonconvex problems
Damek Davis · 2016
Global convergence rate of proximal incremental aggregated gradient methods
N Denizcan Vanli, Mert Gurbuzbalaban, and Asuman Ozdaglar · 2018
Later among the works it cites.
A delay-tolerant proximal-gradient algorithm for distributed learning
Konstantin Mishchenko, Franck Iutzeler, Jérôme Malick, and Massih-Reza Amini · 2018
Later among the works it cites.
On unbounded delays in asynchronous parallel fixed-point algorithms
Robert Hannah and Wotao Yin · 2018
Later among the works it cites.
Improved asynchronous parallel optimization analysis for stochastic incremental methods
Rémi Leblond, Fabian Pedregosa, and Simon Lacoste-Julien · 2018
Later among the works it cites.
General proximal incremental aggregated gradient algorithms: Better and novel results under general scheme
Tao Sun, Yuejiao Sun, Dongsheng Li, and Qing Liao · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Linear convergence of gradient and proximal-gradient methods under the polyak-łojasiewicz condition
Hamed Karimi, Julie Nutini, and Mark Schmidt · 2016
Cited alongside, same era.
On the convergence rate of incremental aggregated gradient algorithms
Mert Gurbuzbalaban, Asuman Ozdaglar, and Pablo A Parrilo · 2017
Cited alongside, same era.
Asynchronous coordinate descent under more realistic assumption
Tao Sun, Robert Hannah, and Wotao Yin · 2017
Cited alongside, same era.
Iteration complexity analysis of block coordinate descent methods
Mingyi Hong, Xiangfeng Wang, Meisam Razaviyayn, and Zhi-Quan Luo · 2017
Cited alongside, same era.
Xiaoge Deng, Tao Sun, Feng Liu, and Feng Huang · 2020
Later among the works it cites.
Accelerating incremental gradient optimization with curvature information
Hoi-To Wai, Wei Shi, César A Uribe, Angelia Nedić, and Anna Scaglione · 2020
Later among the works it cites.
Taming convergence for asynchronous stochastic gradient descent with unbounded delay in non-convex learning
Xin Zhang, Jia Liu, and Zhengyuan Zhu · 2020
Later among the works it cites.
Asynchronous iterations in optimization: New sequence results and sharper algorithmic guarantees
Hamid Reza Feyzmahdavian and Mikael Johansson · 2021
Later among the works it cites.