Fetching the paper…
Reading the bibliography…
An emerging design principle in deep learning is that each layer of a deep artificial neural network should be able to easily express the identity transformation.
Extensions of lipschitz mappings into a hilbert space
William B Johnson and Joram Lindenstrauss · 1984
Earlier work this paper cites.
Neural networks and principal component analysis: Learning from examples without local minima
P. Baldi and K. Hornik · 1989
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton · 2012
Earlier work this paper cites.
Random matrices and complexity of spin glasses
Antonio Auffinger, Gérard Ben Arous, and Jiří Černỳ · 2013
Earlier work this paper cites.
Identifying and attacking the saddle point problem in high-dimensional non-convex optimization
Yann N Dauphin, Razvan Pascanu, Caglar Gulcehre, Kyunghyun Cho, Surya Ganguli, and Yoshua Bengio · 2014
Earlier work this paper cites.
Qualitatively characterizing neural network optimization problems
I. J. Goodfellow, O. Vinyals, and A. M. Saxe · 2014
Earlier work this paper cites.
Striving for Simplicity: The All Convolutional Net
J. T. Springenberg, A. Dosovitskiy, T. Brox, and M. Riedmiller · 2014
Cited alongside, same era.
The loss surfaces of multilayer networks
Anna Choromanska, Mikael Henaff, Michael Mathieu, Gérard Ben Arous, and Yann LeCun · 2015
Cited alongside, same era.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2015
Cited alongside, same era.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Sergey Ioffe and Christian Szegedy · 2015
Cited alongside, same era.
Identity mappings in deep residual networks
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Cited alongside, same era.
Linear Convergence of Gradient and Proximal-Gradient Methods Under the Polyak- \ \backslash L { \{ } \} ojasiewicz Condition
H. Karimi, J. Nutini, and M. Schmidt · 2016
Closest in time.
Deep Learning without Poor Local Minima
K. Kawaguchi · 2016
Closest in time.
No bad local minima: Data independent training error guarantees for multilayer neural networks
D. Soudry and Y. Carmon · 2016
Closest in time.
Normal matrix, from mathworld–a wolfram web resource., 2016
Eric W. Weisstein · 2016
Closest in time.
Johnson–lindenstrauss lemma — wikipedia, the free encyclopedia, 2016
Wikipedia · 2016
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Gao Huang, Zhuang Liu, and Kilian Q. Weinberger · 2016
Cited alongside, same era.