Fetching the paper…
Reading the bibliography…
Stochastic Gradient Descent (SGD) is a workhorse in machine learning, yet its slow convergence can be a computational bottleneck.
A stochastic approximation method
H. Robbins and S. Monro · 1951
Earlier work this paper cites.
Near-optimal hashing algorithms for approximate nearest neighbor in high dimensions
A. Andoni and P. Indyk · 2008
Earlier work this paper cites.
Large-scale machine learning with stochastic gradient descent
L. Bottou · 2010
Earlier work this paper cites.
Pegasos: Primal estimated sub-gradient solver for SVM
S. Shalev-Shwartz, Y. Singer, N. Srebro, and A. Cotter · 2011
Earlier work this paper cites.
Accelerating stochastic gradient descent using predictive variance reduction
R. Johnson and T. Zhang · 2013
Cited alongside, same era.
Semi-stochastic gradient descent methods
J. Konečnỳ and P. Richtárik · 2013
Cited alongside, same era.
Minimizing finite sums with the stochastic average gradient
M. Schmidt, N. L. Roux, and F. Bach · 2013
Cited alongside, same era.
Stochastic dual coordinate ascent methods for regularized loss
S. Shalev-Shwartz and T. Zhang · 2013
Cited alongside, same era.
SAGA: A fast incremental gradient method with support for non-strongly convex composite objectives
A. Defazio, F. Bach, and S. Lacoste-Julien · 2014
Later among the works it cites.
Convergence rate of stochastic gradient with constant step size
M. Schmidt · 2014
Later among the works it cites.
Randomized partition trees for nearest neighbor search
S. Dasgupta and K. Sinha · 2015
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…