Fetching the paper…
Reading the bibliography…
We look at the eigenvalues of the Hessian of a loss function before and after training.
Fast exact multiplication by the hessian
Barak A Pearlmutter · 1994
Earlier work this paper cites.
Almost all learning machines are singular
Sumio Watanabe · 2007
Earlier work this paper cites.
On optimization methods for deep learning
Jiquan Ngiam, Adam Coates, Ahbik Lahiri, Bobby Prochnow, Quoc V Le, and Andrew Y Ng · 2011
Earlier work this paper cites.
Efficient backprop
Yann A LeCun, Léon Bottou, Genevieve B Orr, and Klaus-Robert Müller · 2012
Earlier work this paper cites.
Identifying and attacking the saddle point problem in high-dimensional non-convex optimization
Yann N Dauphin, Razvan Pascanu, Caglar Gulcehre, Kyunghyun Cho, Surya Ganguli, and Yoshua Bengio · 2014
Cited alongside, same era.
Explorations on high dimensional landscapes
Levent Sagun, V Uğur Güney, Gérard Ben Arous, and Yann LeCun · 2014
Cited alongside, same era.
Train faster, generalize better: Stability of stochastic gradient descent
Moritz Hardt, Benjamin Recht, and Yoram Singer · 2015
Cited alongside, same era.
Universality in halting time and its applications in optimization
Levent Sagun, Thomas Trogdon, and Yann LeCun · 2015
Cited alongside, same era.
Topology and geometry of deep rectified network optimization landscapes
C Daniel Freeman and Joan Bruna · 2016
Closest in time.
Gradient descent converges to minimizers
Jason D Lee, Max Simchowitz, Michael I Jordan, and Benjamin Recht · 2016
Closest in time.
Gradient descent only converges to minimizers: Non-isolated critical points and invariant regions
Ioannis Panageas and Georgios Piliouras · 2016
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…