B. T. Polyak, “Some methods of speeding up the convergence of iteration methods,” USSR Computational Mathematics and Mathematical Physics , vol. 4, no. 5, pp. 1–17, 1964
1964
Earlier work this paper cites.
Y. E. Nesterov, “A method for solving the convex programming problem with convergence rate o ( 1 / k 2 ) o(1/k^{2}) ,” in Dokl. Akad. Nauk SSSR , vol. 269, 1983, pp. 543–547
1983
Earlier work this paper cites.
H. Robbins and S. Monro, “A stochastic approximation method,” in Herbert Robbins Selected Papers . Springer, 1985, pp. 102–109
1985
Earlier work this paper cites.
L. Bottou, “Online learning and stochastic approximations,” On-line learning in neural networks , vol. 17, no. 9, p. 142, 1998
1998
Earlier work this paper cites.
Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner, “Gradient-based learning applied to document recognition,” Proceedings of the IEEE , vol. 86, no. 11, pp. 2278–2324, 1998
1998
Earlier work this paper cites.
A. Krizhevsky, “Learning multiple layers of features from tiny images,” Master’s thesis, University of Tront , 2009
2009
Earlier work this paper cites.
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, “Imagenet: A large-scale hierarchical image database,” in Computer Vision and Pattern Recognition, 2009. CVPR 2009. IEEE Conference on . Ieee, 2009, pp. 248–255
2009
Earlier work this paper cites.
H. B. McMahan and M. Streeter, “Adaptive bound optimization for online convex optimization,” COLT 2010 , p. 244, 2010
2010
Earlier work this paper cites.
J. Duchi, E. Hazan, and Y. Singer, “Adaptive subgradient methods for online learning and stochastic optimization,” Journal of Machine Learning Research , vol. 12, no. Jul, pp. 2121–2159, 2011
2011
Earlier work this paper cites.
A. Krizhevsky, I. Sutskever, and G. E. Hinton, “Imagenet classification with deep convolutional neural networks,” in Advances in neural information processing systems , 2012, pp. 1097–1105
2012
Earlier work this paper cites.
G. Hinton, N. Srivastava, and K. Swersky, “Neural networks for machine learning lecture 6a overview of mini-batch gradient descent,” Cited on , p. 14, 2012
2012
Earlier work this paper cites.
S. Ghadimi and G. Lan, “Stochastic first-and zeroth-order methods for nonconvex stochastic programming,” SIAM Journal on Optimization , vol. 23, no. 4, pp. 2341–2368, 2013
2013
Earlier work this paper cites.
R. Johnson and T. Zhang, “Accelerating stochastic gradient descent using predictive variance reduction,” in Advances in neural information processing systems , 2013, pp. 315–323
2013
Earlier work this paper cites.
Y. Nesterov, Introductory lectures on convex optimization: A basic course . Springer Science & Business Media, 2013, vol. 87
2013
Earlier work this paper cites.
I. Sutskever, J. Martens, G. Dahl, and G. Hinton, “On the importance of initialization and momentum in deep learning,” in International conference on machine learning , 2013, pp. 1139–1147
2013
Earlier work this paper cites.