Fetching the paper…
Reading the bibliography…
We evaluate natural gradient, an algorithm originally proposed in Amari (1997), for learning deep models.
Differential geometrical methods in statistics
Amari, S. (1985) · 1985
Earlier work this paper cites.
Information geometry of Boltzmann machines
Amari, S., Kurata, K., and Nagaoka, H. (1992) · 1992
Earlier work this paper cites.
Fast exact multiplication by the hessian
Pearlmutter, B. A. (1994) · 1994
Earlier work this paper cites.
An introduction to the conjugate gradient method without the agonizing pain
Shewchuck, J. (1994) · 1994
Earlier work this paper cites.
Neural learning in structured parameter spaces - natural Riemannian gradient
Amari, S. (1997) · 1997
Earlier work this paper cites.
Natural gradient works efficiently in learning
Amari, S.-I. (1998) · 1998
Earlier work this paper cites.
On natural learning and pruning in multilayered perceptrons
Heskes, T. (2000) · 2000
Earlier work this paper cites.
Numerical Optimization
Nocedal, J. and Wright, S. J. (2000) · 2000
Earlier work this paper cites.
Adaptive natural gradient learning algorithms for various stochastic models
Park, H., Amari, S.-I., and Fukumizu, K. (2000) · 2000
Earlier work this paper cites.
A natural policy gradient
Kakade, S. (2001) · 2001
Earlier work this paper cites.
Fast curvature matrix-vector products for second-order gradient descent
Schraudolph, N. N. (2002) · 2002
Earlier work this paper cites.
Iterative scaled trust-region learning in krylov subspaces via peralmutter’s implicit sparse hessian-vector multiply
Mizutani, E. and Demmel, J. (2003) · 2003
Earlier work this paper cites.
Pattern Recognition and Machine Learning (Information Science and Statistics)
Bishop, C. M. (2006) · 2006
Cited alongside, same era.
Natural conjugate gradient training of multilayer perceptrons
Gonzalez, A. and Dorronsoro, J. (2006) · 2006
Cited alongside, same era.
Optimization Algorithms on Matrix Manifolds
Absil, P.-A., Mahony, R., and Sepulchre, R. (2008) · 2008
Cited alongside, same era.
Natural conjugate gradient in variational inference
Honkela, A., Tornio, M., Raiko, T., and Karhunen, J. (2008) · 2008
Cited alongside, same era.
Topmoumoute online natural gradient algorithm
Le Roux, N., Manzagol, P.-A., and Bengio, Y. (2008) · 2008
Cited alongside, same era.
Natural actor-critic
Peters, J. and Schaal, S. (2008) · 2008
Cited alongside, same era.
Information-geometric optimization algorithms: A unifying picture via invariance principles
Arnold, L., Auger, A., Hansen, N., and Olivier, Y. (2011) · 2011
Later among the works it cites.
Deep learners benefit more from out-of-distribution examples
Bengio, Y., Bastien, F., Bergeron, A., Boulanger-Lewandowski, N., Breuel, T., Chherawala, Y., Cisse, M., Côté, M., Erhan, D., Eustache, J., Glorot, X., Muller, X., Pannetier Lebeuf, S., Pascanu, R., Rifai, S., Savard, F., and Sicard, G. (2011) · 2011
Later among the works it cites.
Improved Preconditioner for Hessian Free Optimization
Chapelle, O. and Erhan, D. (2011) · 2011
Later among the works it cites.
MINRES-QLP: A Krylov subspace method for indefinite or singular symmetric systems
Choi, S.-C. T., Paige, C. C., and Saunders, M. A. (2011) · 2011
Later among the works it cites.
Learning recurrent neural networks with hessian-free optimization
Martens, J. and Sutskever, I. (2011) · 2011
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Stochastic search using the natural gradient
Sun, Y., Wierstra, D., Schaul, T., and Schmidhuber, J. (2009) · 2009
Cited alongside, same era.
Why does unsupervised pre-training help deep learning?
Erhan, D., Courville, A., Bengio, Y., and Vincent, P. (2010) · 2010
Cited alongside, same era.
Approximate riemannian conjugate gradient learning for fixed-form variational bayes
Honkela, A., Raiko, T., Kuusela, M., Tornio, M., and Karhunen, J. (2010) · 2010
Cited alongside, same era.
Deep learning via hessian-free optimization
Martens, J. (2010) · 2010
Cited alongside, same era.
A fast natural newton method
Roux, N. L. and Fitzgibbon, A. W. (2010) · 2010
Cited alongside, same era.
The Toronto face dataset
Susskind, J., Anderson, A., and Hinton, G. E. (2010) · 2010
Cited alongside, same era.
Theano: new features and speed improvements
Bastien, F., Lamblin, P., Pascanu, R., Bergstra, J., Goodfellow, I., Bergeron, A., Bouchard, N., and Bengio, Y. (2012) · 2012
Later among the works it cites.
Natural evolution strategies converge on sphere functions
Schaul, T. (2012) · 2012
Later among the works it cites.
The natural gradient by analogy to signal whitening, and recipes and tricks for its use
Sohl-Dickstein, J. (2012) · 2012
Later among the works it cites.
Krylov Subspace Descent for Deep Learning
Vinyals, O. and Povey, D. (2012) · 2012
Later among the works it cites.
Metric-free natural gradient for joint-training of boltzmann machines
Desjardins, G., Pascanu, R., Courville, A., and Bengio, Y. (2013) · 2013
Closest in time.
Training neural networks with stochastic hessian-free optimization
Kiros, R. (2013) · 2013
Closest in time.