A stochastic approximation method
H. Robbins and S. Monro · 1951
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner · 1998
Earlier work this paper cites.
Learning multiple layers of features from tiny images
A. Krizhevsky · 2009
Earlier work this paper cites.
Online learning for matrix factorization and sparse coding
J. Mairal, F. Bach, J. Ponce, and G. Sapiro · 2010
Earlier work this paper cites.
LIBSVM: a library for support vector machines
C.-C. Chang and C.-J. Lin · 2011
Earlier work this paper cites.
Stochastic first and zeroth-order methods for nonconvex stochastic programming
S. Ghadimi and G. Lan · 2013
Earlier work this paper cites.
Memory limited, streaming PCA
I. Mitliagkas, C. Caramanis, and P. Jain · 2013
Earlier work this paper cites.
Sparse bilinear logistic regression
Original
J. V. Shi, Y. Xu, and R. G. Baraniuk · 2014
Earlier work this paper cites.
Striving for simplicity: The all convolutional net
Original
J. T. Springenberg, A. Dosovitskiy, T. Brox, and M. Riedmiller · 2014
Earlier work this paper cites.
Sparse convolutional neural networks
B. Liu, M. Wang, H. Foroosh, M. Tappen, and M. Pensky · 2015
Earlier work this paper cites.
Block stochastic gradient iteration for convex and nonconvex optimization
Y. Xu and W. Yin · 2015
Earlier work this paper cites.
Variance reduction for faster non-convex optimization
Z. Allen-Zhu and E. Hazan · 2016
Earlier work this paper cites.
Accelerated gradient methods for nonconvex nonlinear and stochastic programming
S. Ghadimi and G. Lan · 2016
Earlier work this paper cites.
Mini-batch stochastic approximation methods for nonconvex stochastic composite optimization
S. Ghadimi, G. Lan, and H. Zhang · 2016
Earlier work this paper cites.