Fetching the paper…
Reading the bibliography…
Synchronous mini-batch SGD is state-of-the-art for large-scale distributed machine learning.
Distributed optimization and statistical learning via the alternating direction method of multipliers
Stephen Boyd, Neal Parikh, Eric Chu, Borja Peleato, and Jonathan Eckstein · 1935
Earlier work this paper cites.
A stochastic approxiation method
H. Robbins and S. Monro · 1951
Earlier work this paper cites.
On asymptotic normality in stochastic approximation
Vaclav Fabian · 1968
Earlier work this paper cites.
Efficient estimations from a slowly convergent Robbins-Monro process
D. Ruppert · 1988
Earlier work this paper cites.
Acceleration of stochastic approximation by averaging
B. T. Polyak and A. B. Juditsky · 1992
Earlier work this paper cites.
Co-coercivity and its role in the convergence of iterative schemes for solving variational inequalities
Dao Li Zhu and Patrice Marcotte · 1996
Earlier work this paper cites.
Introductory Lectures on Convex Optimization: A Basic Course
Y. Nesterov · 2004
Earlier work this paper cites.
Solving large scale linear prediction problems using stochastic gradient descent algorithms
T. Zhang · 2004
Earlier work this paper cites.
Parallel stochastic gradient descent
Olivier Delalleau and Yoshua Bengio · 2007
Earlier work this paper cites.
Confidence Level Solutions for Stochastic Programming
Yu. Nesterov and J. Ph. Vial · 2008
Earlier work this paper cites.
Slow Learners are Fast
J. Langford, A. Smola, and M. Zinkevich · 2009
Earlier work this paper cites.
Efficient large-scale distributed training of conditional maximum entropy models
Ryan Mcdonald, Mehryar Mohri, Nathan Silberman, Dan Walker, and Gideon S Mann · 2009
Earlier work this paper cites.
Robust Stochastic Approximation Approach to Stochastic Programming
A. Nemirovski, A. Juditsky, G. Lan, and A. Shapiro · 2009
Earlier work this paper cites.
Stochastic convex optimization
S. Shalev-Shwartz, O. Shamir, N. Srebro, and K. Sridharan · 2009
Earlier work this paper cites.
Distributed training strategies for the structured perceptron
Ryan McDonald, Keith Hall, and Gideon Mann · 2010
Earlier work this paper cites.
Parallelized stochastic gradient descent
Martin Zinkevich, Markus Weimer, Lihong Li, and Alex J Smola · 2010
Earlier work this paper cites.
Distributed Delayed Stochastic Optimization
A. Agarwal and J. C. Duchi · 2011
Earlier work this paper cites.
Non-asymptotic Analysis of Stochastic Approximation Algorithms for Machine Learning
Francis Bach and Eric Moulines · 2011
Earlier work this paper cites.
Non-asymptotic analysis of stochastic approximation algorithms for machine learning
Eric Moulines and Francis R Bach · 2011
Earlier work this paper cites.
HOGWILD!: A Lock-Free Approach to Parallelizing Stochastic Gradient Descent
F. Niu, B. Recht, C. Re, and S. J. Wright · 2011
Earlier work this paper cites.
Making Gradient Descent Optimal for Strongly Convex Stochastic Optimization
A. Rakhlin, O. Shamir, and K. Sridharan · 2011
Earlier work this paper cites.
Hogwild: A lock-free approach to parallelizing stochastic gradient descent
Benjamin Recht, Christopher Re, Stephen Wright, and Feng Niu · 2011
Earlier work this paper cites.
A simpler approach to obtaining an O(1/t) rate for the stochastic projected subgradient method
S. Lacoste-Julien, M. Schmidt, and F. Bach · 2012
Earlier work this paper cites.
Distributed principal component analysis on networks via directed graphical models, March 2012
Z. Meng, A. Wiesel, and A. O. Hero · 2012
Earlier work this paper cites.
Making gradient descent optimal for strongly convex stochastic optimization
Alexander Rakhlin, Ohad Shamir, Karthik Sridharan, et al · 2012
Earlier work this paper cites.
Communication-efficient algorithms for statistical optimization
Yuchen Zhang, Martin J Wainwright, and John C Duchi · 2012
Earlier work this paper cites.
Non-strongly-convex smooth stochastic approximation with convergence rate O(1/n)
F. Bach and E. Moulines · 2013
Earlier work this paper cites.
GPU Asynchronous Stochastic Gradient Descent to Speed Up Neural Network Training
T. Paine, H. Jin, J. Yang, Z. Lin, and T. Huang · 2013
Earlier work this paper cites.
Stochastic Gradient Descent for Non-smooth Optimization: Convergence Results and Optimal Averaging Schemes
O. Shamir and T. Zhang · 2013
Earlier work this paper cites.
Mini-batch primal and dual methods for svms
Martin Takáč, Avleen Bijral, Peter Richtárik, and Nathan Srebro · 2013
Earlier work this paper cites.
Information-theoretic lower bounds for distributed statistical estimation with communication constraints
Yuchen Zhang, John Duchi, Michael I Jordan, and Martin J Wainwright · 2013
Earlier work this paper cites.
A fast parallel sgd for matrix factorization in shared memory systems
Yong Zhuang, Wei-Sheng Chin, Yu-Chin Juan, and Chih-Jen Lin · 2013
Earlier work this paper cites.
Adaptivity of averaged stochastic gradient descent to local strong convexity for logistic regression
F. Bach · 2014
Cited alongside, same era.
Optimality guarantees for distributed statistical estimation
J. C. Duchi, M. I. Jordan, M. J. Wainwright, and Y. Zhang · 2014
Cited alongside, same era.
On model parallelization and scheduling strategies for distributed machine learning
Seunghak Lee, Jin Kyu Kim, Xun Zheng, Qirong Ho, Garth A Gibson, and Eric P Xing · 2014
Cited alongside, same era.
Stochastic Gradient Descent, Weighted Sampling, and the Randomized Kaczmarz algorithm
Deanna Needell, Rachel Ward, and Nati Srebro · 2014
Cited alongside, same era.
Communication-efficient distributed optimization using an approximate newton-type method
Ohad Shamir, Nati Srebro, and Tong Zhang · 2014
Cited alongside, same era.
Deep learning with Elastic Averaging SGD
Federated learning of deep networks using model averaging
H. Brendan McMahan, Eider Moore, Daniel Ramage, and Blaise Agüera y Arcas · 2016
Later among the works it cites.
AIDE: Fast and Communication Efficient Distributed Optimization
S. J. Reddi, J. Konečný, P. Richtárik, B. Póczós, and A. Smola · 2016
Later among the works it cites.
On the optimality of averaging in distributed statistical learning
Jonathan D. Rosenblatt and Boaz Nadler · 2016
Later among the works it cites.
On Large-Batch Training for Deep Learning: Generalization Gap and Sharp Minima
N. Shirish Keskar, D. Mudigere, J. Nocedal, M. Smelyanskiy, and P. T. P. Tang · 2016
Later among the works it cites.
CoCoA: A General Framework for Communication-Efficient Distributed Optimization
V. Smith, S. Forte, C. Ma, M. Takac, M. I. Jordan, and M. Jaggi · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
S. Zhang, A. Choromanska, and Y. LeCun · 2014
Cited alongside, same era.
Communication complexity of distributed convex learning and optimization
Yossi Arjevani and Ohad Shamir · 2015
Cited alongside, same era.
Mark Braverman, Ankit Garg, Tengyu Ma, Huy L. Nguyen, and David P. Woodruff · 2015
Cited alongside, same era.
A fast parallel stochastic gradient method for matrix factorization in shared memory systems
Wei-Sheng Chin, Yong Zhuang, Yu-Chin Juan, and Chih-Jen Lin · 2015
Cited alongside, same era.
Efficient Distributed SGD with Variance Reduction
S. De and T. Goldstein · 2015
Cited alongside, same era.
Averaged least-mean-squares: bias-variance trade-offs and optimal sampling distributions
A. Défossez and F. Bach · 2015
Cited alongside, same era.
Asynchronous stochastic convex optimization
J. C. Duchi, S. Chaturapruek, and C. Ré · 2015
Cited alongside, same era.
Hantian Zhang, Jerry Li, Kaan Kara, Dan Alistarh, Ji Liu, and Ce Zhang · 2016
Later among the works it cites.
Parallel SGD: When does averaging help?
J. Zhang, C. De Sa, I. Mitliagkas, and C. Ré · 2016
Later among the works it cites.
Fast asynchronous parallel stochastic gradient descent: A lock-free approach with convergence guarantee
Shen-Yi Zhao and Wu-Jun Li · 2016
Later among the works it cites.
Bridging the gap between constant step size stochastic gradient descent and markov chains
Aymeric Dieuleveut, Alain Durmus, and Francis Bach · 2017
Later among the works it cites.
Optimal non-asymptotic bound of the Ruppert-Polyak averaging without strong convexity
S. Gadat and F. Panloup · 2017
Later among the works it cites.
On the rates of convergence of Parallelized Averaged Stochastic Gradient Algorithms
A. Baggioni Godichon and S. Saadane · 2017
Later among the works it cites.
On the rates of convergence of parallelized averaged stochastic gradient algorithms
Baggioni Antoine Godichon and Sofiane Saadane · 2017
Later among the works it cites.
Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour
P. Goyal, P. Dollár, R. Girshick, P. Noordhuis, L. Wesolowski, A. Kyrola, A. Tulloch, Y. Jia, and K. He · 2017
Later among the works it cites.
Accurate, large minibatch sgd: training imagenet in 1 hour
Priya Goyal, Piotr Dollár, Ross Girshick, Pieter Noordhuis, Lukasz Wesolowski, Aapo Kyrola, Andrew Tulloch, Yangqing Jia, and Kaiming He · 2017
Later among the works it cites.
Accelerating Stochastic Gradient Descent
P. Jain, S. M. Kakade, R. Kidambi, P. Netrapalli, and A. Sidford · 2017
Later among the works it cites.
Distributed optimization with arbitrary local solvers
Chenxin Ma, Jakub Konečnỳ, Martin Jaggi, Virginia Smith, Michael I Jordan, Peter Richtárik, and Martin Takáč · 2017
Later among the works it cites.
On-chip training of recurrent neural networks with limited numerical precision
Taesik Na, Jong Hwan Ko, Jaeha Kung, and Saibal Mukhopadhyay · 2017
Later among the works it cites.
Large-scale distributed l-bfgs
Maryam M Najafabadi, Taghi M Khoshgoftaar, Flavio Villanustre, and John Holt · 2017
Later among the works it cites.
Breaking the Nonsmooth Barrier: A Scalable Parallel Method for Composite Optimization
F. Pedregosa, R. Leblond, and S. Lacoste-Julien · 2017
Later among the works it cites.
Optimal algorithms for smooth and strongly convex distributed optimization in networks
K. Scaman, F. Bach, S. Bubeck, Y. Tat Lee, and L. Massoulié · 2017
Later among the works it cites.
Gradient sparsification for communication-efficient distributed optimization
Jianqiao Wangni, Jialei Wang, Ji Liu, and Tong Zhang · 2017
Later among the works it cites.
Terngrad: Ternary gradients to reduce communication in distributed deep learning
Wei Wen, Cong Xu, Feng Yan, Chunpeng Wu, Yandan Wang, Yiran Chen, and Hai Li · 2017
Later among the works it cites.
Large Batch Training of Convolutional Networks
Y. You, I. Gitman, and B. Ginsburg · 2017
Later among the works it cites.
ZipML: Training linear models with end-to-end low precision, and a little bit of deep learning, 06–11 Aug 2017
Hantian Zhang, Jerry Li, Kaan Kara, Dan Alistarh, Ji Liu, and Ce Zhang · 2017
Later among the works it cites.
The Convergence of Stochastic Gradient Descent in Asynchronous Shared Memory
D. Alistarh, C. De Sa, and N. Konstantinov · 2018
Later among the works it cites.
Asynchronous Stochastic Quasi-Newton MCMC for Non-Convex Optimization
U. Şimşekli, Ç. Yıldız, T. H. Nguyen, G. Richard, and A. Taylan Cemgil · 2018
Later among the works it cites.
A Stochastic Large-scale Machine Learning Algorithm for Distributed Features and Observations
B. Fang and D. Klabjan · 2018
Later among the works it cites.
Distributed learning with compressed gradients
S. Khirirat, H. R. Feyzmahdavian, and M. Johansson · 2018
Later among the works it cites.
Improved asynchronous parallel optimization analysis for stochastic incremental methods
R. Leblond, F. Pedregosa, and S. Lacoste-Julien · 2018
Later among the works it cites.
Don’t Use Large Mini-Batches, Use Local SGD
T. Lin, S. U. Stich, and M. Jaggi · 2018
Later among the works it cites.
Local SGD Converges Fast and Communicates Little
S. U. Stich · 2018
Later among the works it cites.
Parallel Restarted SGD for Non-Convex Optimization with Faster Convergence and Less Communication
H. Yu, S. Yang, and S. Zhu · 2018
Later among the works it cites.