Fetching the paper…
Reading the bibliography…
Stochastic Gradient Descent (SGD) is a central tool in machine learning.
A Stochastic Approximation Method
Herbert Robbins and Sutton Monro · 1951
Earlier work this paper cites.
Nonlinear Programming
D. Bertsekas · 1999
Earlier work this paper cites.
Incremental subgradient methods for nondifferentiable optimization
A. Geary and D.P. Bertsekas · 2001
Earlier work this paper cites.
Non-Asymptotic Analysis of Stochastic Approximation Algorithms for Machine Learning
Francis Bach and Eric Moulines · 2011
Earlier work this paper cites.
Incremental proximal methods for large scale convex optimization
Dimitri P. Bertsekas · 2011
Earlier work this paper cites.
Mini-batch Stochastic Approximation Methods for Nonconvex Stochastic Composite Optimization
Saeed Ghadimi, Guanghui Lan, and Hongchao Zhang · 2013
Earlier work this paper cites.
Understanding Machine Learning: From Theory to Algorithms
Shai Ben-David and Shai Shalev-Shwartz · 2014
Earlier work this paper cites.
Convex optimization: Algorithms and complexity
Sébastien Bubeck · 2015
Cited alongside, same era.
Optimization Methods for Large-Scale Machine Learning
Léon Bottou, Frank E. Curtis, and Jorge Nocedal · 2016
Cited alongside, same era.
Without-Replacement Sampling for Stochastic Gradient Methods: Convergence Results and Application to Distributed Optimization
Ohad Shamir · 2016
Cited alongside, same era.
Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour
Priya Goyal, Piotr Dollár, Ross Girshick, Pieter Noordhuis, Lukasz Wesolowski, Aapo Kyrola, Andrew Tulloch, Yangqing Jia Kaiming, and He Facebook · 2017
Cited alongside, same era.
Train longer, generalize better: closing the generalization gap in large batch training of neural networks
Elad Hoffer, Itay Hubara, and Daniel Soudry · 2017
Cited alongside, same era.
Three Factors Influencing Minima in SGD
The Power of Interpolation: Understanding the Effectiveness of SGD in Modern Over-parametrized Learning
Siyuan Ma, Raef Bassily, and Mikhail Belkin · 2017
Later among the works it cites.
On exponential convergence of SGD in non-convex over-parametrized learning
Raef Bassily, Mikhail Belkin, and Siyuan Ma · 2018
Closest in time.
Risk and parameter convergence of logistic regression
Ziwei Ji and Matus Telgarsky · 2018
Closest in time.
Don’t Decay the Learning Rate, Increase the Batch Size
Samuel L. Smith, Pieter-Jan Kindermans, Chris Ying, and Quoc V. Le · 2018
Closest in time.
When Will Gradient Methods Converge to Max-margin Classifier under ReLU Models?
Tengyu Xu, Yi Zhou, Kaiyi Ji, and Yingbin Liang · 2018
Closest in time.
Convergence of Gradient Descent on Separable Data
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Stanislaw Jastrzebski, Zachary Kenton, Devansh Arpit, Nicolas Ballas, Asja Fischer, Yoshua Bengio, and Amos Storkey · 2017
Cited alongside, same era.
Implicit Bias of Gradient Descent on Linear Convolutional Networks
Suriya Gunasekar, Jason Lee, Daniel Soudry, and Nathan Srebro
Cited in the paper.
Characterizing implicit bias in terms of optimization geometry
Suriya Gunasekar, Jason D. Lee, Daniel Soudry, and Nathan Srebro
Cited in the paper.
The implicit bias of gradient descent on separable data
Daniel Soudry, Elad Hoffer, Mor Shpigel Nacson, Suriya Gunasekar, and Nathan Srebro
Cited in the paper.
The implicit bias of gradient descent on separable data
Daniel Soudry, Elad Hoffer, and Nathan Srebro
Cited in the paper.
Mor Shpigel Nacson, Jason Lee, Suriya Gunasekar, Nathan Srebro, and Daniel Soudry · 2019
Closest in time.