Fetching the paper…
Reading the bibliography…
We study the properties of common loss surfaces through their Hessian matrix.
Distribution of eigenvalues for some sets of random matrices
Vladimir A Marčenko and Leonid Andreevich Pastur · 1967
Earlier work this paper cites.
Parallelization of a neural network learning algorithm on a hypercube
J Bourrely · 1989
Earlier work this paper cites.
On the multiple-minima problem in the conformational analysis of molecules: deformation of the potential energy hypersurface by the diffusion equation method
Lucjan Piela, Jaroslaw Kostrowicki, and Harold A Scheraga · 1989
Earlier work this paper cites.
Stochastic gradient learning in neural networks
Léon Bottou · 1991
Earlier work this paper cites.
Optimization methods for computing global minima of nonconvex potential energy functions
Panos M Pardalos, David Shalloway, and Guoliang Xue · 1994
Earlier work this paper cites.
Fast exact multiplication by the hessian
Barak A Pearlmutter · 1994
Earlier work this paper cites.
Flat minima
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Efficient backprop
Yann LeCun, Léon Bottou, GB Orr, and K-R Müller · 1998
Earlier work this paper cites.
Phase transition of the largest eigenvalue for nonnull complex sample covariance matrices
Jinho Baik, Gérard Ben Arous, Sandrine Péché, et al · 2005
Earlier work this paper cites.
Numerical optimization, second edition
Jorge Nocedal and Stephen J Wright · 2006
Earlier work this paper cites.
Learning deep architectures for ai
Yoshua Bengio et al · 2009
Earlier work this paper cites.
Large-scale machine learning with stochastic gradient descent
Léon Bottou · 2010
Earlier work this paper cites.
Deep learning via hessian-free optimization
James Martens · 2010
Cited alongside, same era.
Random matrices and complexity of spin glasses
Antonio Auffinger, Gérard Ben Arous, and Jiří Černỳ · 2013
Cited alongside, same era.
No more pesky learning rates
Tom Schaul, Sixin Zhang, and Yann LeCun · 2013
Cited alongside, same era.
New insights and perspectives on the natural gradient method
James Martens · 2014
Cited alongside, same era.
Explorations on high dimensional landscapes
Levent Sagun, V. Uğur Güney, Gérard Ben Arous, and Yann LeCun · 2014
Cited alongside, same era.
Unreasonable effectiveness of learning neural networks: From accessible states and robust ensembles to basic algorithmic schemes
Gradient descent converges to minimizers
Jason D Lee, Max Simchowitz, Michael I Jordan, and Benjamin Recht · 2016
Later among the works it cites.
The landscape of empirical risk for non-convex losses
Song Mei, Yu Bai, and Andrea Montanari · 2016
Later among the works it cites.
Training recurrent neural networks by diffusion
Hossein Mobahi · 2016
Later among the works it cites.
Gradient descent only converges to minimizers: Non-isolated critical points and invariant regions
Ioannis Panageas and Georgios Piliouras · 2016
Later among the works it cites.
Singularity of the hessian in deep learning
Levent Sagun, Léon Bottou, and Yann LeCun · 2016
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Carlo Baldassi, Christian Borgs, Jennifer T. Chayes, Alessandro Ingrosso, Carlo Lucibello, Luca Saglietti, and Riccardo Zecchina · 2016
Cited alongside, same era.
On the principal components of sample covariance matrices
Alex Bloemendal, Antti Knowles, Horng-Tzer Yau, and Jun Yin · 2016
Cited alongside, same era.
Entropy-sgd: Biasing gradient descent into wide valleys
Pratik Chaudhari, Anna Choromanska, Stefano Soatto, Yann LeCun, Carlo Baldassi, Christian Borgs, Jennifer Chayes, Levent Sagun, and Riccardo Zecchina · 2016
Cited alongside, same era.
Topology and geometry of deep rectified network optimization landscapes
C Daniel Freeman and Joan Bruna · 2016
Cited alongside, same era.
Caglar Gulcehre, Marcin Moczulski, Francesco Visin, and Yoshua Bengio · 2016
Cited alongside, same era.
On graduated optimization for stochastic non-convex problems
Elad Hazan, Kfir Yehuda Levy, and Shai Shalev-Shwartz · 2016
Cited alongside, same era.
On large-batch training for deep learning: Generalization gap and sharp minima
Nitish Shirish Keskar, Dheevatsa Mudigere, Jorge Nocedal, Mikhail Smelyanskiy, and Ping Tak Peter Tang · 2016
Cited alongside, same era.
Later among the works it cites.
Understanding deep learning requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2016
Later among the works it cites.
Energy landscapes for machine learning
Andrew J Ballard, Ritankar Das, Stefano Martiniani, Dhagash Mehta, Levent Sagun, Jacob D Stevenson, and David J Wales · 2017
Closest in time.
Sharp minima can generalize for deep nets
Laurent Dinh, Razvan Pascanu, Samy Bengio, and Yoshua Bengio · 2017
Closest in time.
Accurate, large minibatch sgd: Training imagenet in 1 hour
Priya Goyal, Piotr Dollár, Ross Girshick, Pieter Noordhuis, Lukasz Wesolowski, Aapo Kyrola, Andrew Tulloch, Yangqing Jia, and Kaiming He · 2017
Closest in time.
Three factors influencing minima in sgd
Stanisław Jastrzębski, Zachary Kenton, Devansh Arpit, Nicolas Ballas, Asja Fischer, Yoshua Bengio, and Amos Storkey · 2017
Closest in time.
Opening the black box of deep neural networks via information
Ravid Shwartz-Ziv and Naftali Tishby · 2017
Closest in time.