Exponentially vanishing sub-optimal local minima in multilayer neural networks
Original
Daniel Soudry and Elad Hoffer · 2017
Later among the works it cites.
The loss surface of deep and wide neural networks
Quynh Nguyen and Matthias Hein · 2017
Later among the works it cites.
An analytical formula of population gradient for two-layered ReLU network and its applications in convergence and critical point analysis
Yuandong Tian · 2017
Later among the works it cites.
Fisher-Rao metric, geometry, and complexity of neural networks
Original
Tengyuan Liang, Tomaso Poggio, Alexander Rakhlin, and James Stokes · 2017
Later among the works it cites.
Empirical analysis of the hessian of over-parametrized neural networks
Original
Levent Sagun, Utku Evci, V Ugur Guney, Yann Dauphin, and Leon Bottou · 2017
Later among the works it cites.
Overcoming catastrophic forgetting in neural networks
James Kirkpatrick, Razvan Pascanu, Neil Rabinowitz, Joel Veness, Guillaume Desjardins, Andrei A Rusu, Kieran Milan, John Quan, Tiago Ramalho, Agnieszka Grabska-Barwinska, et al · 2017
Later among the works it cites.
Deep neural networks as gaussian processes
Original
Jaehoon Lee, Yasaman Bahri, Roman Novak, Samuel S Schoenholz, Jeffrey Pennington, and Jascha Sohl-Dickstein · 2017
Later among the works it cites.
Spectrally-normalized margin bounds for neural networks
Peter L Bartlett, Dylan J Foster, and Matus J Telgarsky · 2017
Later among the works it cites.
Sharp minima can generalize for deep nets
Laurent Dinh, Razvan Pascanu, Samy Bengio, and Yoshua Bengio · 2017
Later among the works it cites.
Understanding generalization and stochastic gradient descent
Original
Samuel L Smith and Quoc V Le · 2017
Later among the works it cites.
The emergence of spectral universality in deep networks
Jeffrey Pennington, Samuel S Schoenholz, and Surya Ganguli · 2018
Closest in time.
Exploring the function space of deep-learning machines
Bo Li and David Saad · 2018
Closest in time.
Dynamical isometry and a mean field theory of CNNs: How to train 10,000-layer vanilla convolutional neural networks
Lechao Xiao, Yasaman Bahri, Jascha Sohl-Dickstein, Samuel S Schoenholz, and Jeffrey Pennington · 2018
Closest in time.
Dynamical isometry and a mean field theory of RNNs: Gating enables signal propagation in recurrent neural networks
Minmin Chen, Jeffrey Pennington, and Samuel S Schoenholz · 2018
Closest in time.
Towards understanding the role of over-parametrization in generalization of neural networks
Original
Behnam Neyshabur, Zhiyuan Li, Srinadh Bhojanapalli, Yann LeCun, and Nathan Srebro · 2018
Closest in time.