Fetching the paper…
Reading the bibliography…
Classical stochastic gradient methods for optimization rely on noisy gradient approximations that become progressively less accurate as iterates approach a solution.
A stochastic approximation method
Herbert Robbins and Sutton Monro · 1951
Earlier work this paper cites.
Gradient methods for minimizing functionals
Boris Teodorovich Polyak · 1963
Earlier work this paper cites.
Two-point step size gradient methods
Jonathan Barzilai and Jonathan M Borwein · 1988
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Yann LeCun, Léon Bottou, Yoshua Bengio, and Patrick Haffner · 1998
Earlier work this paper cites.
Comparative accuracies of artificial neural networks and discriminant analysis in predicting forest cover types from cartographic variables
Jock A Blackard and Denis J Dean · 1999
Earlier work this paper cites.
Ijcnn 2001 neural network competition
Danil Prokhorov · 2001
Earlier work this paper cites.
Unsupervised learning of invariant feature hierarchies with applications to object recognition
Marc Aurelio Ranzato, Fu Jie Huang, Y-Lan Boureau, and Yann LeCun · 2007
Earlier work this paper cites.
Learning multiple layers of features from tiny images, 2009
Alex Krizhevsky and Geoffrey Hinton · 2009
Earlier work this paper cites.
High-order methods for basis pursuit
Tom Goldstein and Simon Setzer · 2010
Earlier work this paper cites.
The million song dataset
Thierry Bertin-Mahieux, Daniel P.W. Ellis, Brian Whitman, and Paul Lamere · 2011
Earlier work this paper cites.
Reading digits in natural images with unsupervised feature learning
Yuval Netzer, Tao Wang, Adam Coates, Alessandro Bissacco, Bo Wu, and Andrew Y Ng · 2011
Earlier work this paper cites.
Making gradient descent optimal for strongly convex stochastic optimization
Alexander Rakhlin, Ohad Shamir, and Karthik Sridharan · 2011
Earlier work this paper cites.
Sample size selection in optimization methods for machine learning
Richard H Byrd, Gillian M Chin, Jorge Nocedal, and Yuchen Wu · 2012
Cited alongside, same era.
Hybrid deterministic-stochastic methods for data fitting
Michael P Friedlander and Mark Schmidt · 2012
Cited alongside, same era.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton · 2012
Cited alongside, same era.
Adadelta: an adaptive learning rate method
Matthew D Zeiler · 2012
Cited alongside, same era.
Accelerating stochastic gradient descent using predictive variance reduction
Rie Johnson and Tong Zhang · 2013
Cited alongside, same era.
Introductory lectures on convex optimization: A basic course , volume 87
Accelerating stochastic gradient descent via online learning to sample
Guillaume Bouchard, Théo Trouillon, Julien Perez, and Adrien Gaidon · 2015
Later among the works it cites.
Stopwasting my gradients: Practical svrg
Reza Harikandeh, Mohamed Osama Ahmed, Alim Virani, Mark Schmidt, Jakub Konečnỳ, and Scott Sallinen · 2015
Later among the works it cites.
Probabilistic line searches for stochastic optimization
Maren Mahsereci and Philipp Hennig · 2015
Later among the works it cites.
On variance reduction in stochastic gradient descent and its asynchronous variants
Sashank J Reddi, Ahmed Hefny, Suvrit Sra, Barnabas Poczos, and Alex J Smola · 2015
Later among the works it cites.
Importance sampling for minibatches
Dominik Csiba and Peter Richtárik · 2016
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yurii Nesterov · 2013
Cited alongside, same era.
No more pesky learning rates
Tom Schaul, Sixin Zhang, and Yann LeCun · 2013
Cited alongside, same era.
Minimizing finite sums with the stochastic average gradient
Mark Schmidt, Nicolas Le Roux, and Francis Bach · 2013
Cited alongside, same era.
A field guide to forward-backward splitting with a FASTA implementation
Tom Goldstein, Christoph Studer, and Richard Baraniuk · 2014
Cited alongside, same era.
Adam: A method for stochastic optimization
Diederik Kingma and Jimmy Ba · 2014
Cited alongside, same era.
Stochastic gradient descent, weighted sampling, and the randomized kaczmarz algorithm
Deanna Needell, Rachel Ward, and Nati Srebro · 2014
Cited alongside, same era.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman · 2014
Cited alongside, same era.
Efficient distributed SGD with variance reduction
Soham De and Tom Goldstein · 2016
Closest in time.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Closest in time.
Linear convergence of gradient and proximal-gradient methods under the polyak-łojasiewicz condition
Hamed Karimi, Julie Nutini, and Mark Schmidt · 2016
Closest in time.
On large-batch training for deep learning: Generalization gap and sharp minima
Nitish Shirish Keskar, Dheevatsa Mudigere, Jorge Nocedal, Mikhail Smelyanskiy, and Ping Tak Peter Tang · 2016
Closest in time.
Barzilai-borwein step size for stochastic gradient descent
Conghui Tan, Shiqian Ma, Yu-Hong Dai, and Yuqiu Qian · 2016
Closest in time.
Automated inference with adaptive batches
Soham De, Abhay Yadav, David Jacobs, and Tom Goldstein · 2017
Closest in time.