Fetching the paper…
Reading the bibliography…
For SGD based distributed stochastic optimization, computation complexity, measured by the convergence rate in terms of the number of stochastic gradient calls, and communication complexity, measured by the number of inter-node communication rounds, are two most important performance metrics.
Gradient methods for minimizing functionals
Polyak, B. T · 1963
Earlier work this paper cites.
Problem complexity and method efficiency in optimization
Nemirovsky, A. S. and Yudin, D. B · 1983
Earlier work this paper cites.
New proximal point algorithms for convex minimization
Güler, O · 1992
Earlier work this paper cites.
Nonlinear Programming
Bertsekas, D. P · 1999
Earlier work this paper cites.
Introductory Lectures on Convex Optimization: A Basic Course
Nesterov, Y · 2004
Earlier work this paper cites.
The tradeoffs of large scale learning
Bottou, L. and Bousquet, O · 2008
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A. and Hinton, G · 2009
Earlier work this paper cites.
Robust stochastic approximation approach to stochastic programming
Nemirovski, A., Juditsky, A., Lan, G., and Shapiro, A · 2009
Earlier work this paper cites.
Optimal distributed online prediction using mini-batches
Dekel, O., Gilad-Bachrach, R., Shamir, O., and Xiao, L · 2012
Earlier work this paper cites.
Hybrid deterministic-stochastic methods for data fitting
Friedlander, M. P. and Schmidt, M · 2012
Earlier work this paper cites.
An accelerated inexact proximal point algorithm for convex minimization
He, B. and Yuan, X · 2012
Earlier work this paper cites.
Lacoste-Julien, S., Schmidt, M., and Bach, F · 2012
Earlier work this paper cites.
Making gradient descent optimal for strongly convex stochastic optimization
Rakhlin, A., Shamir, O., and Sridharan, K · 2012
Cited alongside, same era.
Inexact and accelerated proximal point algorithms
Salzo, S. and Villa, S · 2012
Cited alongside, same era.
Stochastic first-and zeroth-order methods for nonconvex stochastic programming
Ghadimi, S. and Lan, G · 2013
Cited alongside, same era.
Zhang, L., Yang, T., Jin, R., and He, X · 2013
Cited alongside, same era.
Beyond the regret minimization barrier: an optimal algorithm for stochastic strongly-convex optimization
Hazan, E. and Kale, S · 2014
Cited alongside, same era.
Asynchronous parallel stochastic gradient for nonconvex optimization
Adabatch: Adaptive batch sizes for training deep neural networks
Devarakonda, A., Naumov, M., and Garland, M · 2017
Later among the works it cites.
A simple parallel algorithm with an O ( 1 / t ) {O}(1/t) convergence rate for general convex programs
Yu, H. and Neely, M. J · 2017
Later among the works it cites.
Optimization methods for large-scale machine learning
Bottou, L., Curtis, F. E., and Nocedal, J · 2018
Later among the works it cites.
A linear speedup analysis of distributed deep learning with sparse and quantized communication
Jiang, P. and Agrawal, G · 2018
Later among the works it cites.
Don’t use large mini-batches, use local SGD
Lin, T., Stich, S. U., and Jaggi, M · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Lian, X., Huang, Y., Li, Y., and Liu, J · 2015
Cited alongside, same era.
A universal catalyst for first-order optimization
Lin, H., Mairal, J., and Harchaoui, Z · 2015
Cited alongside, same era.
Mini-batch stochastic approximation methods for nonconvex stochastic composite optimization
Ghadimi, S., Lan, G., and Zhang, H · 2016
Cited alongside, same era.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Cited alongside, same era.
Linear convergence of gradient and proximal-gradient methods under the Polyak-Lojasiewicz condition
Karimi, H., Nutini, J., and Schmidt, M · 2016
Cited alongside, same era.
Proximally guided stochastic subgradient method for nonsmooth, nonconvex problems
Davis, D. and Grimmer, B · 2017
Cited alongside, same era.
Automated inference with adaptive batches
De, S., Yadav, A., Jacobs, D., and Goldstein, T · 2017
Cited alongside, same era.
Catalyst for gradient-based nonconvex optimization
Paquette, C., Lin, H., Drusvyatskiy, D., Mairal, J., and Harchaoui, Z · 2018
Later among the works it cites.
Understanding generalization and stochastic gradient descent
Smith, S. L. and Le, Q. V · 2018
Later among the works it cites.
Don’t decay the learning rate, increase the batch size
Smith, S. L., Kindermans, P.-J., Ying, C., and Le, Q. V · 2018
Later among the works it cites.
Local SGD converges fast and communicates little
Stich, S. U · 2018
Later among the works it cites.
Wang, J. and Joshi, G · 2018
Later among the works it cites.
Yu, H., Yang, S., and Zhu, S · 2018
Later among the works it cites.