Fetching the paper…
Reading the bibliography…
Stochastic gradient descent (SGD) is widely used in machine learning.
Probability inequalities for sums of bounded random variables
Wassily Hoeffding · 1963
Earlier work this paper cites.
Simplifying neural nets by discovering flat minima
Sepp Hochreiter and Jürgen Schmidhuber · 1995
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
John C. Duchi, Elad Hazan, and Yoram Singer · 2011
Earlier work this paper cites.
Accelerating stochastic gradient descent using predictive variance reduction
Rie Johnson and Tong Zhang · 2013
Earlier work this paper cites.
On the importance of initialization and momentum in deep learning
Ilya Sutskever, James Martens, George E. Dahl, and Geoffrey E. Hinton · 2013
Earlier work this paper cites.
SAGA: A fast incremental gradient method with support for non-strongly convex composite objectives
Aaron Defazio, Francis R. Bach, and Simon Lacoste-Julien · 2014
Earlier work this paper cites.
Qualitatively characterizing neural network optimization problems
Ian J. Goodfellow and Oriol Vinyals · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Introductory Lectures on Convex Optimization: A Basic Course
Yurii Nesterov · 2014
Earlier work this paper cites.
Escaping from saddle points - online stochastic gradient for tensor decomposition
Rong Ge, Furong Huang, Chi Jin, and Yang Yuan · 2015
Earlier work this paper cites.
Train faster, generalize better: Stability of stochastic gradient descent
M. Hardt, B. Recht, and Y. Singer · 2015
Cited alongside, same era.
Improved SVRG for non-strongly-convex or sum-of-non-convex objectives
Zeyuan Allen-Zhu and Yang Yuan · 2016
Cited alongside, same era.
Entropy-SGD: Biasing Gradient Descent Into Wide Valleys
P. Chaudhari, A. Choromanska, S. Soatto, Y. LeCun, C. Baldassi, C. Borgs, J. Chayes, L. Sagun, and R. Zecchina · 2016
Cited alongside, same era.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Cited alongside, same era.
Densely Connected Convolutional Networks
G. Huang, Z. Liu, K. Q. Weinberger, and L. van der Maaten · 2016
Cited alongside, same era.
Minimizing finite sums with the stochastic average gradient
Mark Schmidt, Nicolas Le Roux, and Francis Bach · 2016
Cited alongside, same era.
How to escape saddle points efficiently
Chi Jin, Rong Ge, Praneeth Netrapalli, Sham M. Kakade, and Michael I. Jordan · 2017
Later among the works it cites.
On large-batch training for deep learning: Generalization gap and sharp minima
Nitish Shirish Keskar, Dheevatsa Mudigere, Jorge Nocedal, Mikhail Smelyanskiy, and Ping Tak Peter Tang · 2017
Later among the works it cites.
Convergence analysis of two-layer neural networks with relu activation
Yuanzhi Li and Yang Yuan · 2017
Later among the works it cites.
SGDR: stochastic gradient descent with restarts
Ilya Loshchilov and Frank Hutter · 2017
Later among the works it cites.
Stochastic Gradient Descent as Approximate Bayesian Inference
S. Mandt, M. D. Hoffman, and D. M. Blei · 2017
Later among the works it cites.
Generalization Bounds of SGLD for Non-convex Learning: Two Theoretical Viewpoints
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Stochastic gradient descent performs variational inference, converges to limit cycles for deep networks
P. Chaudhari and S. Soatto · 2017
Cited alongside, same era.
Sharp Minima Can Generalize For Deep Nets
L. Dinh, R. Pascanu, S. Bengio, and Y. Bengio · 2017
Cited alongside, same era.
Train longer, generalize better: closing the generalization gap in large batch training of neural networks
Elad Hoffer, Itay Hubara, and Daniel Soudry · 2017
Cited alongside, same era.
Snapshot ensembles: Train 1, get m for free
Gao Huang, Yixuan Li, Geoff Pleiss, Zhuang Liu, John E. Hopcroft, and Kilian Q. Weinberger · 2017
Cited alongside, same era.
W. Mou, L. Wang, X. Zhai, and K. Zheng · 2017
Later among the works it cites.
The Impact of Local Geometry and Batch Size on the Convergence and Divergence of Stochastic Gradient Descent
V. Patel · 2017
Later among the works it cites.
A Bayesian Perspective on Generalization and Stochastic Gradient Descent
S. L. Smith and Q. V. Le · 2017
Later among the works it cites.
A hitting time analysis of stochastic gradient langevin dynamics
Yuchen Zhang, Percy Liang, and Moses Charikar · 2017
Later among the works it cites.