Fetching the paper…
Reading the bibliography…
Second-order methods for neural network optimization have several advantages over methods based on first-order gradient descent, including better scaling to large mini-batch sizes and fewer updates needed for convergence.
Acceleration of stochastic approximation by averaging
B. T. Polyak and A. B. Juditsky · 1992
Earlier work this paper cites.
Fast exact multiplication by the Hessian
B. A. Pearlmutter · 1994
Earlier work this paper cites.
Long short-term memory
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
Natural gradient works efficiently in learning
S.-I. Amari · 1998
Earlier work this paper cites.
Efficient backprop
Y. A. LeCun, L. Bottou, G. B. Orr, and K.-R. Müller · 1998
Earlier work this paper cites.
On the momentum term in gradient descent learning algorithms
N. Qian · 1999
Earlier work this paper cites.
On the truncated conjugate gradient method
Y. Yuan · 2000
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner · 2001
Earlier work this paper cites.
Learning precise timing with LSTM recurrent networks
F. A. Gers, N. N. Schraudolph, and J. Schmidhuber · 2002
Earlier work this paper cites.
Fast curvature matrix-vector products for second-order gradient descent
N. N. Schraudolph · 2002
Earlier work this paper cites.
Large Scale Machine Learning
R. Collobert · 2004
Earlier work this paper cites.
A fast learning algorithm for deep belief nets
G. E. Hinton, S. Osindero, and Y. W. Teh · 2006
Earlier work this paper cites.
Methods of information geometry , volume 191
S.-i. Amari and H. Nagaoka · 2007
Earlier work this paper cites.
Topmoumoute online natural gradient algorithm
N. Le Roux, P.-A. Manzagol, and Y. Bengio · 2008
Earlier work this paper cites.
Deep learning via Hessian-free optimization
J. Martens · 2010
Cited alongside, same era.
On the use of stochastic Hessian information in optimization methods for machine learning
R. H. Byrd, G. M. Chin, W. Neveitt, and J. Nocedal · 2011
Cited alongside, same era.
Adaptive subgradient methods for online learning and stochastic optimization
J. Duchi, E. Hazan, and Y. Singer · 2011
Cited alongside, same era.
Learning recurrent neural networks with Hessian-free optimization
J. Martens and I. Sutskever · 2011
Cited alongside, same era.
Hogwild: A lock-free approach to parallelizing stochastic gradient descent
B. Recht, C. Re, S. Wright, and F. Niu · 2011
Cited alongside, same era.
Sample size selection in optimization methods for machine learning
R. H. Byrd, G. M. Chin, J. Nocedal, and Y. Wu · 2012
Revisiting natural gradient for deep networks
R. Pascanu and Y. Bengio · 2014
Later among the works it cites.
Automatic differentiation in machine learning: a survey
A. G. Baydin, B. A. Pearlmutter, A. A. Radul, and J. M. Siskind · 2015
Later among the works it cites.
Lasagne: First release., Aug. 2015
S. Dieleman, J. Schlüter, C. Raffel, E. Olson, S. K. Sønderby, D. Nouri, D. Maturana, M. Thoma, E. Battenberg, J. Kelly, J. D. Fauw, M. Heilman, D. M. de Almeida, B. McFee, H. Weideman, G. Takács, P. de Rivaz, J. Crall, G. Sanders, K. Rasul, C. Liu, G. French, and J. Degrave · 2015
Later among the works it cites.
Optimizing neural networks with Kronecker-factored approximate curvature
J. Martens and R. Grosse · 2015
Later among the works it cites.
Deep learning with elastic averaging SGD
S. Zhang, A. E. Choromanska, and Y. LeCun · 2015
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Large scale distributed deep networks
J. Dean, G. Corrado, R. Monga, K. Chen, M. Devin, M. Mao, A. Senior, P. Tucker, K. Yang, Q. V. Le, et al · 2012
Cited alongside, same era.
Training deep and recurrent networks with Hessian-free optimization
J. Martens and I. Sutskever · 2012
Cited alongside, same era.
Training neural networks with stochastic Hessian-free optimization
R. Kiros · 2013
Cited alongside, same era.
Introductory lectures on convex optimization: A basic course , volume 87
Y. Nesterov · 2013
Cited alongside, same era.
On the importance of initialization and momentum in deep learning
I. Sutskever, J. Martens, G. E. Dahl, and G. E. Hinton · 2013
Cited alongside, same era.
Mini-batch primal and dual methods for SVMs
M. Takác, A. S. Bijral, P. Richtárik, and N. Srebro · 2013
Cited alongside, same era.
R. Al-Rfou, G. Alain, A. Almahairi, C. Angermueller, D. Bahdanau, N. Ballas, F. Bastien, J. Bayer, A. Belikov, A. Belopolsky, et al · 2016
Later among the works it cites.
Revisiting distributed synchronous SGD
J. Chen, R. Monga, S. Bengio, and R. Jozefowicz · 2016
Later among the works it cites.
Distributed deep learning using synchronous stochastic gradient descent
D. Das, S. Avancha, D. Mudigere, K. Vaidynathan, S. Sridharan, D. Kalamkar, B. Kaul, and P. Dubey · 2016
Later among the works it cites.
Accurate, large minibatch SGD: Training imagenet in 1 hour
P. Goyal, P. Dollár, R. Girshick, P. Noordhuis, L. Wesolowski, A. Kyrola, A. Tulloch, Y. Jia, and K. He · 2016
Later among the works it cites.
A Kronecker-factored approximate Fisher matrix for convolution layers
R. Grosse and J. Martens · 2016
Later among the works it cites.
On large-batch training for deep learning: Generalization gap and sharp minima
N. S. Keskar, D. Mudigere, J. Nocedal, M. Smelyanskiy, and P. T. P. Tang · 2016
Later among the works it cites.
Second-order optimization for neural networks
J. Martens · 2016
Later among the works it cites.
Distributed second-order optimization using Kronecker-factored approximations
J. Ba, R. Grosse, and J. Martens · 2017
Closest in time.
Sharp minima can generalize for deep nets
L. Dinh, R. Pascanu, S. Bengio, and Y. Bengio · 2017
Closest in time.