Fetching the paper…
Reading the bibliography…
We propose a class of very simple modifications of gradient descent and stochastic gradient descent.
A stochastic approximation method
H. Robinds and S. Monro · 1951
Earlier work this paper cites.
A method for solving the convex programming problem with convergence rate o ( 1 / k 2 ) o(1/k^{2})
Y. Nesterov · 1983
Earlier work this paper cites.
Two problems with backpropagation and other steepest-descent learning procedures for networks
R. Sutton · 1986
Earlier work this paper cites.
Convergence analysis of gradient descent stochastic algorithms
A. Shapiro and Y. Wardi · 1996
Earlier work this paper cites.
Matrix Analysis
R. Bhatia · 1997
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner · 1998
Earlier work this paper cites.
Introductory lectures on convex programming volume i: Basic course
Y. Nesterov · 1998
Earlier work this paper cites.
Nonlinear programming
D. P. Bertsekas · 1999
Earlier work this paper cites.
On the momentum term in gradient descent learning algorithms
N. Qian · 1999
Earlier work this paper cites.
Sobolev gradients and joint variational image segmentation, denoising, and deblurring
M. Jung, G. Chung, G. Sundaramoorthi, L. Vese, and A. Yuille · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
A. Krizhevsky · 2009
Earlier work this paper cites.
https://terrytao.wordpress.com/2010/01/03/254a-notes-1-concentration-of-measure/
254a, notes 1: Concentration of measure · 2010
Earlier work this paper cites.
Partial differential equations
L.C. Evans · 2010
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
J. Duchi, E. Hazan, and Y. Singer · 2011
Earlier work this paper cites.
Bayesian learning via stochastic gradient langevin dynamics
M. Welling and Y. Teh · 2011
Earlier work this paper cites.
Stochastic gradient descent tricks
L. Bottou · 2012
Earlier work this paper cites.
Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude
T. Tieleman and G. Hinton · 2012
Cited alongside, same era.
Adadelta: An adaptive learning rate method
M. Zeiler · 2012
Cited alongside, same era.
Accelerating stochastic gradient descent using predictive variance reduction
R. Johoson and T. Zhang · 2013
Cited alongside, same era.
Playing Atari with deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. Graves, I. Antonoglou, D. Wierstra, and M. Riedmiller · 2013
Cited alongside, same era.
An empirical study of learning rates in deep neural networks for speech recognition
A. Senior, G. Heigold, M. Ranzato, and K. Yang · 2013
Cited alongside, same era.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Later among the works it cites.
Mastering the game of go with deep neural networks and tree search
D. Silver and et al · 2016
Later among the works it cites.
Towards principled methods for training generative adversarial networks
M. Arjovsky and L. Bottou · 2017
Later among the works it cites.
Deep relaxation: partial differential equations for optimizing deep neural networks
P. Chaudhari, A. Oberman, S. Osher, S. Soatto, and C. Guillame · 2017
Later among the works it cites.
Nonconvex finite-sum optimization via scsg methods
L. Lei, C. Ju, J. Chen, and M. Jordan · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Stochastic gradient descent for non-smooth optimization: Convergence results and optimal averaging schemes
O. Shamir and T. Zhang · 2013
Cited alongside, same era.
Saga: A fast incremental gradient method with support for non-strongly convex composite objectives
A. Defazio and F. Bach · 2014
Cited alongside, same era.
Generative adversarial nets
I. J. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. C. Courville, and Y. Bengio · 2014
Cited alongside, same era.
Adam: A method for stochastic optimization
D. Kingma and J. Ba · 2014
Cited alongside, same era.
Deep learning in neural networks: An overview
J Schmidhuber · 2014
Cited alongside, same era.
Human-level control through deep reinforcement learning
Mnih and et al · 2015
Cited alongside, same era.
Unsupervised representation learning with deep convolutional generative adversarial networks
A. Radford, L. Metz, and S. Chintala · 2015
Cited alongside, same era.
H. Li, Z. Xu, G. Taylor, and T. Goldstein · 2017
Later among the works it cites.
S. Chintala M. Arjovsky and L. Bottou · 2017
Later among the works it cites.
Stochastic gradient descent as approximate bayesian inference
S. Mandt, M. Hoffman, and D. Blei · 2017
Later among the works it cites.
Automatic differentiation in pytorch
A. Paszke, S. Gross, S. Chintala, G. Chanan, E. Yang, Z. DeVito, Z. Lin, A. Desmaison, L. Antiga, and A. Lerer · 2017
Later among the works it cites.
Katyusha: The first direct acceleration of stochastic gradient methods
Z. Allen-Zhu · 2018
Closest in time.
Optimization methods for large-scale machine learning
L. Bottou, E. F. Curtis, and J. Nocedal · 2018
Closest in time.
Dnn’s sharpest directions along the sgd trajectory
S. Jastrzebski, Z. Kenton, N. Ballas, A. Fischer, Y. Bengio, and A. Storkey · 2018
Closest in time.
Cs231n: Convolutional neural networks for visual recognition
F. Li and et al · 2018
Closest in time.
On the convergence of adam and beyond
S. Reddi, S. Kale, and S. Kumar · 2018
Closest in time.
Group normalization
Y. Wu and K. He · 2018
Closest in time.
Privacy-preserving erm by laplacian smoothing stochastic gradient descent
B. Wang, Q. Gu, M. Boedihardjo, F. Barekat, and S. Osher · 2019
Closest in time.