Fetching the paper…
Reading the bibliography…
This paper studies the landscape of empirical risk of deep neural networks by theoretically analyzing its convergence behavior to the population risk as well as its stationary points and properties.
On differentiable functions with isolated critical points
D. Gromoll and W. Meyer · 1969
Earlier work this paper cites.
On the capabilities of multilayer perceptrons
E. Baum · 1988
Earlier work this paper cites.
Wave-net: A multiresolution, hierarchical neural network with localized learning
B. Bakshi and G. Stephanopoulos · 1993
Earlier work this paper cites.
For valid generalization, the size of the weights is more important than the size of the network
P. Bartlett · 1997
Earlier work this paper cites.
Statistical learning theory
V. N. Vapnik and V. Vapnik · 1998
Earlier work this paper cites.
Vapnik-chervonenkis dimension of neural nets
P. Bartlett and W. Maass · 2003
Earlier work this paper cites.
A fast learning algorithm for deep belief nets
G. Hinton, S. Osindero, and Y. Teh · 2006
Earlier work this paper cites.
Replica symmetry breaking condition exposed by random matrix calculation of landscape complexity
Y. Fyodorov and I. Williams · 2007
Earlier work this paper cites.
A unified architecture for natural language processing: Deep neural networks with multitask learning
R. Collobert and J. Weston · 2008
Earlier work this paper cites.
A unified framework for high-dimensional analysis of M-estimators with decomposable regularizers
S. Negahban, B. Yu, M. Wainwright, and P. Ravikumar · 2009
Earlier work this paper cites.
Learnability, stability and uniform convergence
S. Shalev-Shwartz, O. Shamir, N. Srebro, and K. Sridharan · 2010
Earlier work this paper cites.
Deep neural networks for acoustic modeling in speech recognition: The shared views of four research groups
G. Hinton, L. Deng, D. Yu, G. Dahl, A. Mohamed, N. Jaitly, A. Senior, V. Vanhoucke, P. Nguyen, T. Sainath, et al · 2012
Earlier work this paper cites.
Introduction to the non-asymptotic analysis of random matrices, compressed sensing
R. Vershynin · 2012
Earlier work this paper cites.
Robustness and generalization
H. Xu and S. Mannor · 2012
Cited alongside, same era.
Modern geometry—methods and applications: Part II: The geometry and topology of manifolds
B. Dubrovin, A. Fomenko, and S. Novikov · 2012
Cited alongside, same era.
Speech recognition with deep recurrent neural networks
A. Graves, A. Mohamed, and G. Hinton · 2013
Cited alongside, same era.
Hanson-wright inequality and sub-gaussian concentration
M. Rudelson and R. Vershynin · 2013
Cited alongside, same era.
Identifying and attacking the saddle point problem in high-dimensional non-convex optimization
Y. Dauphin, R. Pascanu, C. Gulcehre, K. Cho, S. Ganguli, and Y. Bengio · 2014
Cited alongside, same era.
Understanding machine learning: From theory to algorithms
S. Shalev-Shwartz and S. Ben-David · 2014
Cited alongside, same era.
Deep learning without poor local minima
K. Kawaguchi · 2016
Later among the works it cites.
No bad local minima: Data independent training error guarantees for multilayer neural networks
D. Soudry and Y. Carmon · 2016
Later among the works it cites.
ℓ 1 \ell_{1} -regularized neural networks are improperly learnable in polynomial time
Y. Zhang, J. Lee, and M. Jordan · 2016
Later among the works it cites.
Convexified convolutional neural networks
Y. Zhang, P. Liang, and M. Wainwright · 2016
Later among the works it cites.
The landscape of empirical risk for non-convex losses
S. Mei, Y. Bai, and A. Montanari · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Understanding machine learning: From theory to algorithms
S. Shalev-Shwartz and S. Ben-David · 2014
Cited alongside, same era.
Going deeper with convolutions
C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich · 2015
Cited alongside, same era.
The loss surfaces of multilayer networks
A. Choromanska, M. Henaff, M. Mathieu, G. Arous, and Y. LeCun · 2015
Cited alongside, same era.
Open problem: The landscape of the loss surfaces of multilayer networks
A. Choromanska, Y. LeCun, and G. Arous · 2015
Cited alongside, same era.
Norm-based capacity control in neural networks
B. Neyshabur, R. Tomioka, and N. Srebro · 2015
Cited alongside, same era.
Statistic s997 lecture notes, MIT mathematics
P. Rigollet · 2015
Cited alongside, same era.
The power of depth for feedforward neural networks
R. Eldan and O. Shamir · 2016
Later among the works it cites.
Lecture notes of advanced statistical theory I, CMU
R. Alessandro · 2016
Later among the works it cites.
S. Shalev-Shwartz, O. Shamir, and S. Shammah · 2017
Closest in time.
Symmetry-breaking convergence analysis of certain two-layered neural networks with ReLU nonlinearity
Y. Tian · 2017
Closest in time.
The loss surface of deep and wide neural networks
Q. Nguyen and M. Hein · 2017
Closest in time.
Fast rates for empirical risk minimization of strict saddle problems
A. Gonen and S. Shalev-Shwartz · 2017
Closest in time.
Exponentially vanishing sub-optimal local minima in multilayer neural networks
D. Soudry and E. Hoffer · 2017
Closest in time.