Fetching the paper…
Reading the bibliography…
Mini-batch stochastic gradient descent and variants thereof have become standard for large-scale empirical risk minimization like the training of neural networks.
A stochastic approximation method
Robbins, Herbert and Monro, Sutton · 1951
Earlier work this paper cites.
Learning representations by back-propagating errors
Rumelhart, David E., Hinton, Geoffrey E., and Williams, Ronald J · 1986
Earlier work this paper cites.
Gradient-based learning applied to document recognition
LeCun, Yann, Bottou, Léon, Bengio, Yoshua, and Haffner, Patrick · 1998
Earlier work this paper cites.
Information Theory, Inference and Learning Algorithms
MacKay, David JC · 2003
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, Alex · 2009
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
Duchi, John, Hazan, Elad, and Singer, Yoram · 2011
Earlier work this paper cites.
Reading digits in natural images with unsupervised feature learning
Netzer, Yuval, Wang, Tao, Coates, Adam, Bissacco, Alessandro, Wu, Bo, and Ng, Andrew Y · 2011
Earlier work this paper cites.
Sample size selection in optimization methods for machine learning
Byrd, Richard H, Chin, Gillian M, Nocedal, Jorge, and Wu, Yuchen · 2012
Earlier work this paper cites.
Hybrid deterministic-stochastic methods for data fitting
Friedlander, Michael P and Schmidt, Mark · 2012
Cited alongside, same era.
ADADELTA: An adaptive learning rate method
Zeiler, Matthew D · 2012
Cited alongside, same era.
Accelerating stochastic gradient descent using predictive variance reduction
Johnson, Rie and Zhang, Tong · 2013
Cited alongside, same era.
No more pesky learning rates
Schaul, Tom, Zhang, Sixin, and LeCun, Yann · 2013
Cited alongside, same era.
SAGA: A fast incremental gradient method with support for non-strongly convex composite objectives
Defazio, Aaron, Bach, Francis, and Lacoste-Julien, Simon · 2014
Cited alongside, same era.
Stochastic gradient descent, weighted sampling, and the randomized Kaczmarz algorithm
Probabilistic line searches for stochastic optimization
Mahsereci, Maren and Hennig, Philipp · 2015
Later among the works it cites.
Non-uniform stochastic average gradient method for training conditional random fields
Schmidt, Mark, Babanezhad, Reza, Ahmed, Mohamed Osama, Defazio, Aaron, Clifton, Ann, and Sarkar, Anoop · 2015
Later among the works it cites.
Stochastic optimization with importance sampling for regularized loss minimization
Zhao, Peilin and Zhang, Tong · 2015
Later among the works it cites.
Optimization methods for large-scale machine learning
Bottou, Léon, Curtis, Frank E, and Nocedal, Jorge · 2016
Closest in time.
Importance sampling for minibatches
Csiba, Dominik and Richtárik, Peter · 2016
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Needell, Deanna, Ward, Rachel, and Srebro, Nati · 2014
Cited alongside, same era.
StopWasting my gradients: Practical SVRG
Harikandeh, Reza, Ahmed, Mohamed Osama, Virani, Alim, Schmidt, Mark, Konečný, Jakub, and Sallinen, Scott · 2015
Cited alongside, same era.
Adam: A method for stochastic optimization
Kingma, Diederik and Ba, Jimmy · 2015
Cited alongside, same era.
Starting small—learning with adaptive sample sizes
Daneshmand, Hadi, Lucchi, Aurelien, and Hofmann, Thomas · 2016
Closest in time.
Cost-sensitive approach for batch size optimization
Pirotta, Matteo and Restelli, Marcello · 2016
Closest in time.
Automated inference with adaptive batches
De, Soham, Yadav, Abhay, Jacobs, David, and Goldstein, Tom · 2017
Closest in time.