Fetching the paper…
Reading the bibliography…
Second-order optimization methods such as natural gradient descent have the potential to speed up training of neural networks by correcting for the curvature of the loss function.
The Levenberg-Marquardt algorithm: implementation and theory
Moré, J.J · 1978
Earlier work this paper cites.
Learning representations by back-propagating errors
Rumelhart, D.E., Hinton, G.E., and Williams, R.J · 1986
Earlier work this paper cites.
Backpropagation applied to handwritten zip code recognition
LeCun, Y., Boser, B., Denker, J. S., Henderson, D., Howard, R. E., Hubbard, W., and Jackel, L. D · 1989
Earlier work this paper cites.
Acceleration of stochastic approximation by averaging
Polyak, B. T. and Juditsky, A. B · 1992
Earlier work this paper cites.
Applied Numerical Linear Algebra
Demmel, J. W · 1997
Earlier work this paper cites.
Long short-term memory
Hochreiter, S. and Schmidhuber, J · 1997
Earlier work this paper cites.
Natural gradient works efficiently in learning
Amari, Shun-Ichi · 1998
Earlier work this paper cites.
Efficient backprop
LeCun, Y., Bottou, L., Orr, G., and Müller, K · 1998
Earlier work this paper cites.
On “natural” learning and pruning in multilayered perceptrons
Heskes, Tom · 2000
Earlier work this paper cites.
Natural image statistics and neural representation
Simoncelli, E. P. and Olshausen, B. A · 2001
Earlier work this paper cites.
Fast curvature matrix-vector products for second-order gradient descent
Schraudolph, Nicol N · 2002
Earlier work this paper cites.
High performance convolutional neural networks for document processing
Chellapilla, K., Puri, S., and Simard, P · 2006
Earlier work this paper cites.
Numerical optimization
Nocedal, Jorge and Wright, Stephen J · 2006
Earlier work this paper cites.
Topmoumoute online natural gradient algorithm
Le Roux, Nicolas, Manzagol, Pierre-antoine, and Bengio, Yoshua · 2008
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A · 2009
Earlier work this paper cites.
CUDAMat: A CUDA-based matrix class for Python
Mnih, V · 2009
Cited alongside, same era.
Deep learning via Hessian-free optimization
Martens, J · 2010
Cited alongside, same era.
Modeling pixel means and covariances using factorized third-order Boltzmann machines
Ranzato, M. and Hinton, G. E · 2010
Cited alongside, same era.
A tutorial on stochastic approximation algorithms for training restricted Boltzmann machines and deep belief nets
Swersky, K., Chen, Bo, Marlin, B., and de Freitas, N · 2010
Cited alongside, same era.
Improved preconditioner for Hessian-free optimization
Chapelle, O. and Erhan, D · 2011
Cited alongside, same era.
Adaptive subgradient methods for online learning and stochastic optimization
Duchi, J., Hazan, E., and Singer, Y · 2011
Cited alongside, same era.
ADADELTA: An adaptive learning rate method
Zeiler, Matthew D · 2013
Later among the works it cites.
New insights and perspectives on the natural gradient method, 2014
Martens, J · 2014
Later among the works it cites.
On the saddle point problem for non-convex optimization
Pascanu, R., Dauphin, Y. N., Ganguli, S., and Bengio, Y · 2014
Later among the works it cites.
Fast large-scale optimization by unifying stochastic gradient and quasi-Newton methods
Sohl-Dickstein, J., Poole, B., and Ganguli, S · 2014
Later among the works it cites.
Sequence to sequence learning with neural networks
Sutskever, I., Vinyals, O., and Le, Q. V. V · 2014
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Reading digits in natural images with unsupervised feature learning
Netzer, Y., Wang, T., Coates, A., Bissacco, A., Wu, B., and Ng, A. Y · 2011
Cited alongside, same era.
Large scale distributed deep networks
Dean, J., Corrado, G. S., Monga, R., Chen, K., Devin, M., Le, Q. V., Mao, M. Z., Ranzato, M., Senior, A., Tucker, P., Yang, K., and Ng, A. Y · 2012
Cited alongside, same era.
ImageNet classification with deep convolutional neural networks
Krizhevsky, A., Sutskever, I., and Hinton, G. E · 2012
Cited alongside, same era.
Lecture 6.5, RMSProp
Tieleman, T. and Hinton, G · 2012
Cited alongside, same era.
Krylov subspace descent for deep learning
Vinyals, O. and Povey, D · 2012
Cited alongside, same era.
Enhanced gradient for training restricted Boltzmann machines
Cho, K., Raiko, T., and Ilin, A · 2013
Cited alongside, same era.
Yang, Z., Moczulski, M., Denil, M., de Freitas, N., Smola, A., Song, L., and Wang, Z · 2014
Later among the works it cites.
Desjardins, G., Simonyan, K., Pascanu, R., and Kavukcuoglu, K · 2015
Later among the works it cites.
Scaling up natural gradient by sparsely factorizing the inverse Fisher matrix
Grosse, Roger and Salakhutdinov, Ruslan · 2015
Later among the works it cites.
Batch normalization: accelerating deep network training by reducing internal covariate shift
Ioffe, S. and Szegedy, C · 2015
Later among the works it cites.
Adam: a method for stochastic optimization
Kingma, D. P. and Ba, J. L · 2015
Later among the works it cites.
Optimizing neural networks with Kronecker-factored approximate curvature
Martens, J. and Grosse, R · 2015
Later among the works it cites.
Riemannian metrics for neural networks I: feedforward networks
Ollivier, Y · 2015
Later among the works it cites.
Parallel training of DNNs with natural gradient and parameter averaging
Povey, Daniel, Zhang, Xiaohui, and Khudanpur, Sanjeev · 2015
Later among the works it cites.
Toronto Deep Learning ConvNet
Srivastava, N · 2015
Later among the works it cites.