Fetching the paper…
Reading the bibliography…
Local SGD is a communication-efficient variant of SGD for large-scale training, where multiple GPUs perform SGD independently and average the model parameters periodically.
Differentiation of the limit mapping in a dynamical system
KJ Falconer · 1983
Earlier work this paper cites.
Solutions of a stochastic differential equation forced onto a manifold by a large drift
G. S. Katzenberger · 1991
Earlier work this paper cites.
Flat minima
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Reinforcement learning - an introduction
Richard S. Sutton and Andrew G. Barto · 1998
Earlier work this paper cites.
Invariant manifolds for weak solutions to stochastic equations
Damir Filipović · 2000
Earlier work this paper cites.
Invariant manifold reduction for stochastic dynamical systems
Aijun Du and JinQiao Duan · 2007
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky et al · 2009
Earlier work this paper cites.
Efficient large-scale distributed training of conditional maximum entropy models
Gideon Mann, Ryan T. McDonald, Mehryar Mohri, Nathan Silberman, and Dan Walker · 2009
Earlier work this paper cites.
Parallelized stochastic gradient descent
Martin Zinkevich, Markus Weimer, Lihong Li, and Alex Smola · 2010
Earlier work this paper cites.
Hogwild: A lock-free approach to parallelizing stochastic gradient descent
Benjamin Recht, Christopher Ré, Stephen J. Wright, and Feng Niu · 2011
Earlier work this paper cites.
Practical Recommendations for Gradient-Based Training of Deep Architectures , pp. 437–478
Yoshua Bengio · 2012
Earlier work this paper cites.
Efficient BackProp , pp. 9–48
Yann A. LeCun, Léon Bottou, Genevieve B. Orr, and Klaus-Robert Müller · 2012
Earlier work this paper cites.
Stochastic differential equations: an introduction with applications
Bernt Øksendal · 2013
Earlier work this paper cites.
One weird trick for parallelizing convolutional neural networks
Alex Krizhevsky · 2014
Earlier work this paper cites.
Parallel training of dnns with natural gradient and parameter averaging
Daniel Povey, Xiaohui Zhang, and Sanjeev Khudanpur · 2014
Earlier work this paper cites.
1-bit stochastic gradient descent and its application to data-parallel distributed training of speech dnns
Frank Seide, Hao Fu, Jasha Droppo, Gang Li, and Dong Yu · 2014
Earlier work this paper cites.
Improving deep neural network acoustic models using generalized maxout networks
Xiaohui Zhang, Jan Trmal, Daniel Povey, and Sanjeev Khudanpur · 2014
Earlier work this paper cites.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Earlier work this paper cites.
ImageNet Large Scale Visual Recognition Challenge
Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, Alexander C. Berg, and Li Fei-Fei · 2015
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
K. Simonyan and A. Zisserman · 2015
Earlier work this paper cites.
Scalable distributed DNN training using commodity GPU cloud computing
Nikko Strom · 2015
Earlier work this paper cites.
Experiments on parallel training of deep neural network using model averaging
Hang Su and Haoyu Chen · 2015
Earlier work this paper cites.
Revisiting distributed synchronous SGD
Jianmin Chen, Xinghao Pan, Rajat Monga, Samy Bengio, and Rafal Jozefowicz · 2016
Earlier work this paper cites.
Scalable training of deep learning machines by incremental block training with intra-block parallel optimization and blockwise model-update filtering
Kai Chen and Qiang Huo · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Cited alongside, same era.
Gaussian error linear units (gelus)
Dan Hendrycks and Kevin Gimpel · 2016
Cited alongside, same era.
Sharp minima can generalize for deep nets
Laurent Dinh, Razvan Pascanu, Samy Bengio, and Yoshua Bengio · 2017
Cited alongside, same era.
Accurate, large minibatch SGD: Training imagenet in 1 hour
Priya Goyal, Piotr Dollár, Ross Girshick, Pieter Noordhuis, Lukasz Wesolowski, Aapo Kyrola, Andrew Tulloch, Yangqing Jia, and Kaiming He · 2017
Cited alongside, same era.
Train longer, generalize better: closing the generalization gap in large batch training of neural networks
Elad Hoffer, Itay Hubara, and Daniel Soudry · 2017
Cited alongside, same era.
Geometric insights into the convergence of nonlinear TD learning
David Brandfonbrener and Joan Bruna · 2020
Later among the works it cites.
Convergence rates for the stochastic gradient descent method for non-convex objective functions
Benjamin Fehrman, Benjamin Gess, and Arnulf Jentzen · 2020
Later among the works it cites.
Fantastic generalization measures and where to find them
Yiding Jiang, Behnam Neyshabur, Hossein Mobahi, Dilip Krishnan, and Samy Bengio · 2020
Later among the works it cites.
Scaffold: Stochastic controlled averaging for federated learning
Sai Praneeth Karimireddy, Satyen Kale, Mehryar Mohri, Sashank Reddi, Sebastian Stich, and Ananda Theertha Suresh · 2020
Later among the works it cites.
Tighter theory for local SGD on identical and heterogeneous data
Ahmed Khaled, Konstantin Mishchenko, and Peter Richtárik · 2020
Later among the works it cites.
Reconciling modern deep learning with traditional optimization analyses: The intrinsic learning rate
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Wenqing Hu, Chris Junchi Li, Lei Li, and Jian-Guo Liu · 2017
Cited alongside, same era.
Three factors influencing minima in SGD
Stanisław Jastrzębski, Zachary Kenton, Devansh Arpit, Nicolas Ballas, Asja Fischer, Yoshua Bengio, and Amos Storkey · 2017
Cited alongside, same era.
On large-batch training for deep learning: Generalization gap and sharp minima
Nitish Shirish Keskar, Dheevatsa Mudigere, Jorge Nocedal, Mikhail Smelyanskiy, and Ping Tak Peter Tang · 2017
Cited alongside, same era.
Communication-efficient learning of deep networks from decentralized data
Brendan McMahan, Eider Moore, Daniel Ramage, Seth Hampson, and Blaise Aguera y Arcas · 2017
Cited alongside, same era.
Exploring generalization in deep learning
Behnam Neyshabur, Srinadh Bhojanapalli, David Mcallester, and Nati Srebro · 2017
Cited alongside, same era.
Searching for activation functions
Prajit Ramachandran, Barret Zoph, and Quoc V Le · 2017
Cited alongside, same era.
Highly scalable deep learning training system with mixed-precision: Training imagenet in four minutes
Xianyan Jia, Shutao Song, Wei He, Yangzihao Wang, Haidong Rong, Feihu Zhou, Liqiang Xie, Zhenyu Guo, Yuanzhou Yang, Liwei Yu, et al · 2018
Cited alongside, same era.
Zhiyuan Li, Kaifeng Lyu, and Sanjeev Arora · 2020
Later among the works it cites.
On the generalization benefit of noise in stochastic gradient descent
Samuel Smith, Erich Elsen, and Soham De · 2020
Later among the works it cites.
Large batch optimization for deep learning: Training BERT in 76 minutes
Yang You, Jing Li, Sashank Reddi, Jonathan Hseu, Sanjiv Kumar, Srinadh Bhojanapalli, Xiaodan Song, James Demmel, Kurt Keutzer, and Cho-Jui Hsieh · 2020
Later among the works it cites.
Why are adaptive methods good for attention models?
Jingzhao Zhang, Sai Praneeth Karimireddy, Andreas Veit, Seungyeon Kim, Sashank Reddi, Sanjiv Kumar, and Suvrit Sra · 2020
Later among the works it cites.
Label noise SGD provably prefers flat global minimizers
Alex Damian, Tengyu Ma, and Jason D. Lee · 2021
Later among the works it cites.
Sharpness-aware minimization for efficiently improving generalization
Pierre Foret, Ariel Kleiner, Hossein Mobahi, and Behnam Neyshabur · 2021
Later among the works it cites.
Exponential escape efficiency of SGD from sharp minima in non-stationary regime
Hikaru Ibayashi and Masaaki Imaizumi · 2021
Later among the works it cites.
Advances and open problems in federated learning
Peter Kairouz, H Brendan McMahan, Brendan Avent, Aurélien Bellet, Mehdi Bennis, Arjun Nitin Bhagoji, Kallista Bonawitz, Zachary Charles, Graham Cormode, Rachel Cummings, et al · 2021
Later among the works it cites.
On linear stability of SGD and input-smoothness of neural networks
Chao Ma and Lexing Ying · 2021
Later among the works it cites.
Trade-offs of Local SGD at scale: An empirical study
Jose Javier Gonzalez Ortiz, Jonathan Frankle, Mike Rabbat, Ari Morcos, and Nicolas Ballas · 2021
Later among the works it cites.
On the origin of implicit regularization in stochastic gradient descent
Samuel L Smith, Benoit Dherin, David Barrett, and Soham De · 2021
Later among the works it cites.
Cooperative SGD: A unified framework for the design and analysis of local-update SGD algorithms
Jianyu Wang and Gauri Joshi · 2021
Later among the works it cites.
A diffusion theory for deep learning dynamics: Stochastic gradient descent exponentially favors flat minima
Zeke Xie, Issei Sato, and Masashi Sugiyama · 2021
Later among the works it cites.
Understanding the unstable convergence of gradient descent
Kwangjun Ahn, Jingzhao Zhang, and Suvrit Sra · 2022
Later among the works it cites.
Sharp bounds for federated averaging (Local SGD) and continuous perspective
Margalit R Glasgow, Honglin Yuan, and Tengyu Ma · 2022
Later among the works it cites.
Fast mixing of stochastic gradient descent with normalization and weight decay
Zhiyuan Li, Tianhao Wang, and Dingli Yu · 2022
Later among the works it cites.
Understanding the generalization benefit of normalization layers: Sharpness reduction, 2022
Kaifeng Lyu, Zhiyuan Li, and Sanjeev Arora · 2022
Later among the works it cites.
On the SDEs and scaling rules for adaptive gradient algorithms
Sadhika Malladi, Kaifeng Lyu, Abhishek Panigrahi, and Sanjeev Arora · 2022
Later among the works it cites.
On the unreasonable effectiveness of federated averaging with heterogeneous data
Jianyu Wang, Rudrajit Das, Gauri Joshi, Satyen Kale, Zheng Xu, and Tong Zhang · 2022
Later among the works it cites.