Fetching the paper…
Reading the bibliography…
In this paper, we consider a general stochastic optimization problem which is often at the core of supervised learning, such as deep learning and linear classification.
A stochastic approximation method
H. Robbins and S. Monro · 1951
Earlier work this paper cites.
On the limited memory BFGS method for large scale optimization
D. C. Liu and J. Nocedal · 1989
Earlier work this paper cites.
Introductory lectures on convex optimization: a basic course
Y. Nesterov · 2004
Earlier work this paper cites.
The tradeoffs of large scale learning
L. Bottou and O. Bousquet · 2007
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
J. Duchi, E. Hazan, and Y. Singer · 2011
Cited alongside, same era.
Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude
T. Tieleman and G. Hinton · 2012
Cited alongside, same era.
Accelerating stochastic gradient descent using predictive variance reduction
R. Johnson and T. Zhang · 2013
Cited alongside, same era.
SAGA: A fast incremental gradient method with support for non-strongly convex composite objectives
A. Defazio, F. Bach, and S. Lacoste-Julien · 2014
Cited alongside, same era.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2014
Later among the works it cites.
Optimization methods for large-scale machine learning
L. Bottou, F. E. Curtis, and J. Nocedal · 2016
Later among the works it cites.
Minimizing finite sums with the stochastic average gradient
M. Schmidt, N. Le Roux, and F. Bach · 2016
Later among the works it cites.
SARAH: A novel method for machine learning problems using stochastic recursive gradient
L. Nguyen, J. Liu, K. Scheinberg, and M. Takáč · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…