Fetching the paper…
Reading the bibliography…
We provide tight finite-time convergence bounds for gradient descent and stochastic gradient descent on quadratic functions, when the gradients are delayed and reflect iterates from $\tau$ rounds ago.
Problem complexity and method efficiency in optimization. 1983
AS Nemirovsky and DB Yudin · 1983
Earlier work this paper cites.
Enumerative combinatorics. vol. i, the wadsworth & brooks/cole mathematics series, wadsworth & brooks, 1986
Richard P Stanley · 1986
Earlier work this paper cites.
Parallel and distributed computation: numerical methods
Dimitri P Bertsekas and John N Tsitsiklis · 1989
Earlier work this paper cites.
Distributed asynchronous incremental subgradient methods
A Nedić, Dimitri P Bertsekas, and Vivek S Borkar · 2001
Earlier work this paper cites.
Introductory lectures on convex optimization
Yurii Nesterov · 2004
Earlier work this paper cites.
generatingfunctionology
Herbert S Wilf · 2005
Earlier work this paper cites.
Analytic combinatorics
Philippe Flajolet and Robert Sedgewick · 2009
Earlier work this paper cites.
Slow learners are fast
John Langford, Alex J Smola, and Martin Zinkevich · 2009
Earlier work this paper cites.
Distributed delayed stochastic optimization
Alekh Agarwal and John C Duchi · 2011
Earlier work this paper cites.
Better mini-batch algorithms via accelerated gradient methods
Andrew Cotter, Ohad Shamir, Nati Srebro, and Karthik Sridharan · 2011
Cited alongside, same era.
Hogwild: A lock-free approach to parallelizing stochastic gradient descent
Benjamin Recht, Christopher Re, Stephen Wright, and Feng Niu · 2011
Cited alongside, same era.
Optimal distributed online prediction using mini-batches
Ofer Dekel, Ran Gilad-Bachrach, Ohad Shamir, and Lin Xiao · 2012
Cited alongside, same era.
Online learning under delayed feedback
Pooria Joulani, Andras Gyorgy, and Csaba Szepesvári · 2013
Cited alongside, same era.
Stochastic gradient descent for non-smooth optimization: Convergence results and optimal averaging schemes
Ohad Shamir and Tong Zhang · 2013
Cited alongside, same era.
Mini-batch primal and dual methods for svms
Martin Takác, Avleen Singh Bijral, Peter Richtárik, and Nati Srebro · 2013
Asynchronous stochastic convex optimization: the noise is in the noise and sgd don’t care
Sorathan Chaturapruek, John C Duchi, and Christopher Ré · 2015
Later among the works it cites.
An optimal randomized incremental gradient method
Guanghui Lan and Yi Zhou · 2015
Later among the works it cites.
Perturbed iterate analysis for asynchronous stochastic optimization
Horia Mania, Xinghao Pan, Dimitris Papailiopoulos, Benjamin Recht, Kannan Ramchandran, and Michael I Jordan · 2015
Later among the works it cites.
Adadelay: Delay adaptive distributed stochastic convex optimization
Suvrit Sra, Adams Wei Yu, Mu Li, and Alexander J Smola · 2015
Later among the works it cites.
An asynchronous mini-batch algorithm for regularized stochastic optimization
Hamid Reza Feyzmahdavian, Arda Aytekin, and Mikael Johansson · 2016
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
A delayed proximal gradient method with linear convergence rate
Hamid Reza Feyzmahdavian, Arda Aytekin, and Mikael Johansson · 2014
Cited alongside, same era.
Distributed stochastic optimization and learning
Ohad Shamir and Nathan Srebro · 2014
Cited alongside, same era.
Convex optimization: Algorithms and complexity
Sébastien Bubeck et al · 2015
Cited alongside, same era.
Later among the works it cites.
Decentralized consensus algorithm with delayed and stochastic gradients
Benjamin Sirb and Xiaojing Ye · 2016
Later among the works it cites.
Tight complexity bounds for optimizing composite objectives
Blake E Woodworth and Nati Srebro · 2016
Later among the works it cites.
Improved asynchronous parallel optimization analysis for stochastic incremental methods
Rémi Leblond, Fabian Pederegosa, and Simon Lacoste-Julien · 2018
Closest in time.