Some methods of speeding up the convergence of iteration methods
B. T. Polyak · 1964
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky · 2009
Earlier work this paper cites.
Introductory lectures on convex optimization: A basic course , volume 87
Yurii Nesterov · 2013
Earlier work this paper cites.
On the importance of initialization and momentum in deep learning
Ilya Sutskever, James Martens, George Dahl, and Geoffrey Hinton · 2013
Earlier work this paper cites.
Stochastic gradient descent, weighted sampling, and the randomized kaczmarz algorithm
Deanna Needell, Rachel Ward, and Nati Srebro · 2014
Earlier work this paper cites.
From averaging to acceleration, there is only a step-size
Nicolas Flammarion and Francis Bach · 2015
Earlier work this paper cites.
Global convergence of the heavy-ball method for convex optimization
Euhanna Ghadimi, Hamid Feyzmahdavian, and Mikael Johansson · 2015
Earlier work this paper cites.
ImageNet Large Scale Visual Recognition Challenge
Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, Alexander C. Berg, and Li Fei-Fei · 2015
Earlier work this paper cites.
Adaptive averaging in accelerated descent dynamics
Walid Krichene, Alexandre Bayen, and Peter L Bartlett · 2016
Earlier work this paper cites.