Fetching the paper…
Reading the bibliography…
A central challenge to many fields of science and engineering involves minimizing non-convex error functions over continuous, high dimensional spaces.
On the distribution of the roots of certain symmetric matrices
Wigner, E. P. (1958) · 1958
Earlier work this paper cites.
Neural networks and principal component analysis: Learning from examples without local minima
Baldi, P. and Hornik, K. (1989) · 1989
Earlier work this paper cites.
Learning long-term dependencies with gradient descent is difficult
Bengio, Y., Simard, P., and Frasconi, P. (1994) · 1994
Earlier work this paper cites.
Fast exact multiplication by the hessian
Pearlmutter, B. A. (1994) · 1994
Earlier work this paper cites.
On-line learning in soft committee machines
Saad, D. and Solla, S. A. (1995) · 1995
Earlier work this paper cites.
Natural Gradient Descent for On-Line Learning
Rattray, M., Saad, D., and Amari, S. I. (1998) · 1998
Earlier work this paper cites.
On-line learning theory of soft committee machines with correlated hidden units –steepest gradient descent and natural gradient descent–
Inoue, M., Park, H., and Okada, M. (2003) · 2003
Earlier work this paper cites.
Gaussian Processes for Machine Learning (Adaptive Computation and Machine Learning)
Rasmussen, C. E. and Williams, C. K. I. (2005) · 2005
Earlier work this paper cites.
Numerical Optimization
Nocedal, J. and Wright, S. (2006) · 2006
Earlier work this paper cites.
Statistics of critical points of gaussian fields on large-dimensional spaces
Bray, A. J. and Dean, D. S. (2007) · 2007
Earlier work this paper cites.
Replica symmetry breaking condition exposed by random matrix calculation of landscape complexity
Fyodorov, Y. V. and Williams, I. (2007) · 2007
Cited alongside, same era.
Topmoumoute online natural gradient algorithm
Le Roux, N., Manzagol, P.-A., and Bengio, Y. (2007) · 2007
Cited alongside, same era.
Mean field theory of spin glasses: statistics and dynamics
Parisi, G. (2007) · 2007
Cited alongside, same era.
Theano: a CPU and GPU math expression compiler
Bergstra, J., Breuleux, O., Bastien, F., Lamblin, P., Pascanu, R., Desjardins, G., Turian, J., Warde-Farley, D., and Bengio, Y. (2010) · 2010
Cited alongside, same era.
Advanced Calculus: A Geometric View
Callahan, J. (2010) · 2010
Cited alongside, same era.
Deep learning via hessian-free optimization
Martens, J. (2010) · 2010
Cited alongside, same era.
Random search for hyper-parameter optimization
Bergstra, J. and Bengio, Y. (2012) · 2012
Later among the works it cites.
Krylov Subspace Descent for Deep Learning
Vinyals, O. and Povey, D. (2012) · 2012
Later among the works it cites.
On the difficulty of training recurrent neural networks
Pascanu, R., Mikolov, T., and Bengio, Y. (2013) · 2013
Later among the works it cites.
Learning hierarchical category structure in deep neural networks
Saxe, A., McClelland, J., and Ganguli, S. (2013) · 2013
Later among the works it cites.
On the importance of initialization and momentum in deep learning
Sutskever, I., Martens, J., Dahl, G. E., and Hinton, G. E. (2013) · 2013
Later among the works it cites.
Revisiting natural gradient for deep networks
Pascanu, R. and Bengio, Y. (2014) · 2014
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
An analysis on negative curvature induced by singularity in multi-layer neural-network learning
Mizutani, E. and Dreyfus, S. (2010) · 2010
Cited alongside, same era.
Newton-type methods
Murray, W. (2010) · 2010
Cited alongside, same era.
Theano: new features and speed improvements
Bastien, F., Lamblin, P., Pascanu, R., Bergstra, J., Goodfellow, I. J., Bergeron, A., Bouchard, N., and Bengio, Y. (2012) · 2012
Cited alongside, same era.
On the saddle point problem for non-convex optimization
Pascanu, R., Dauphin, Y., Ganguli, S., and Bengio, Y. (2014) · 2014
Closest in time.
Exact solutions to the nonlinear dynamics of learning in deep linear neural network
Saxe, A., McClelland, J., and Ganguli, S. (2014) · 2014
Closest in time.
Fast large-scale optimization by unifying stochastic gradient and quasi-newton methods
Sohl-Dickstein, J., Poole, B., and Ganguli, S. (2014) · 2014
Closest in time.