Fetching the paper…
Reading the bibliography…
Many stochastic optimization algorithms work by estimating the gradient of the cost function on the fly by sampling datapoints uniformly at random from a training set.
The nonstochastic multiarmed bandit problem
Peter Auer, Nicolo Cesa-Bianchi, Yoav Freund, and Robert E Schapire · 2002
Earlier work this paper cites.
Large-scale machine learning with stochastic gradient descent
Léon Bottou · 2010
Earlier work this paper cites.
Libsvm: a library for support vector machines
Chih-Chung Chang and Chih-Jen Lin · 2011
Earlier work this paper cites.
Semi-stochastic gradient descent methods
Jakub Konecnỳ and Peter Richtárik · 2013
Earlier work this paper cites.
Stochastic multi-armed-bandit problem with non-stationary rewards
Omar Besbes, Yonatan Gur, and Assaf Zeevi · 2014
Earlier work this paper cites.
Saga: A fast incremental gradient method with support for non-strongly convex composite objectives
Aaron Defazio, Francis Bach, and Simon Lacoste-Julien · 2014
Earlier work this paper cites.
A proximal stochastic gradient method with progressive variance reduction
Lin Xiao and Tong Zhang · 2014
Cited alongside, same era.
Accelerating minibatch stochastic gradient descent using stratified sampling
Peilin Zhao and Tong Zhang · 2014
Cited alongside, same era.
Stochastic optimization with importance sampling
Peilin Zhao and Tong Zhang · 2014
Cited alongside, same era.
Non-uniform stochastic average gradient method for training conditional random fields
Mark Schmidt, Reza Babanezhad, Mohamed Ahmed, Aaron Defazio, Ann Clifton, and Anoop Sarkar · 2015
Cited alongside, same era.
Convex optimization
Nisheeth K Vishnoi · 2015
Cited alongside, same era.
Svrg++ with non-uniform sampling
Stochastic optimization with importance sampling for regularized loss minimization
Peilin Zhao and Tong Zhang · 2015
Later among the works it cites.
Improved svrg for non-strongly-convex or sum-of-non-convex objectives
Zeyuan Allen-Zhu and Yang Yuan · 2016
Later among the works it cites.
A stochastic quasi-newton method for large-scale optimization
Richard H Byrd, Samantha L Hansen, Jorge Nocedal, and Yoram Singer · 2016
Later among the works it cites.
Importance sampling for minibatches
Dominik Csiba and Peter Richtárik · 2016
Later among the works it cites.
Stochastic learning on imbalanced data: Determinantal point processes for mini-batch diversification
Cheng Zhang, Hedvig Kjellstrom, and Stephan Mandt · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Tamás Kern and András György
Cited in the paper.
Closest in time.