Fetching the paper…
Reading the bibliography…
Neural network training relies on our ability to find "good" minimizers of highly non-convex loss functions.
Training a 3-node neural network is np-complete
Avrim Blum and Ronald L Rivest · 1989
Earlier work this paper cites.
A simple weight decay can improve generalization
Anders Krogh and John A Hertz · 1992
Earlier work this paper cites.
Flat minima
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Visualization of learning in multilayer perceptron networks using principal component analysis
Marcus Gallagher and Tom Downs · 2003
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
Xavier Glorot and Yoshua Bengio · 2010
Earlier work this paper cites.
Identifying and attacking the saddle point problem in high-dimensional non-convex optimization
Yann N Dauphin, Razvan Pascanu, Caglar Gulcehre, Kyunghyun Cho, Surya Ganguli, and Yoshua Bengio · 2014
Earlier work this paper cites.
The loss surfaces of multilayer networks
Anna Choromanska, Mikael Henaff, Michael Mathieu, Gérard Ben Arous, and Yann LeCun · 2015
Earlier work this paper cites.
Qualitatively characterizing neural network optimization problems
Ian J Goodfellow, Oriol Vinyals, and Andrew M Saxe · 2015
Earlier work this paper cites.
Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
Sergey Ioffe and Christian Szegedy · 2015
Earlier work this paper cites.
Very Deep Convolutional Networks for Large-Scale Image Recognition
Karen Simonyan and Andrew Zisserman · 2015
Earlier work this paper cites.
Phasemax: Convex phase retrieval via basis pursuit
Tom Goldstein and Christoph Studer · 2016
Earlier work this paper cites.
Deep Residual Learning for Image Recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
An empirical analysis of deep network loss surfaces
Daniel Jiwoong Im, Michael Tao, and Kristin Branson · 2016
Earlier work this paper cites.
Stuck in a what? adventures in weight space
Zachary C Lipton · 2016
Earlier work this paper cites.
Visualizing deep network training trajectories with pca
Eliana Lorch · 2016
Cited alongside, same era.
On the quality of the initial basin in overspecified neural networks
Itay Safran and Ohad Shamir · 2016
Cited alongside, same era.
Local minima in training of deep networks
Grzegorz Swirszcz, Wojciech Marian Czarnecki, and Razvan Pascanu · 2016
Cited alongside, same era.
Wide residual networks
Sergey Zagoruyko and Nikos Komodakis · 2016
Cited alongside, same era.
The shattered gradients problem: If resnets are the answer, then what is the question?
David Balduzzi, Marcus Frean, Lennox Leary, JP Lewis, Kurt Wan-Duo Ma, and Brian McWilliams · 2017
Cited alongside, same era.
Entropy-sgd: Biasing gradient descent into wide valleys
Pratik Chaudhari, Anna Choromanska, Stefano Soatto, and Yann LeCun · 2017
Densely connected convolutional networks
Gao Huang, Zhuang Liu, Kilian Q Weinberger, and Laurens van der Maaten · 2017
Closest in time.
Generalization in deep learning
Kenji Kawaguchi, Leslie Pack Kaelbling, and Yoshua Bengio · 2017
Closest in time.
On large-batch training for deep learning: Generalization gap and sharp minima
Nitish Shirish Keskar, Dheevatsa Mudigere, Jorge Nocedal, Mikhail Smelyanskiy, and Ping Tak Peter Tang · 2017
Closest in time.
Convergence analysis of two-layer neural networks with relu activation
Yuanzhi Li and Yang Yuan · 2017
Closest in time.
Theory of deep learning ii: Landscape of the empirical risk in deep learning
Qianli Liao and Tomaso Poggio · 2017
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Automated inference with adaptive batches
Soham De, Abhay Yadav, David Jacobs, and Tom Goldstein · 2017
Cited alongside, same era.
Sharp minima can generalize for deep nets
Laurent Dinh, Razvan Pascanu, Samy Bengio, and Yoshua Bengio · 2017
Cited alongside, same era.
Computing nonvacuous generalization bounds for deep (stochastic) neural networks with many more parameters than training data
Gintare Karolina Dziugaite and Daniel M Roy · 2017
Cited alongside, same era.
Topology and geometry of half-rectified network optimization
C Daniel Freeman and Joan Bruna · 2017
Cited alongside, same era.
Accurate, large minibatch sgd: Training imagenet in 1 hour
Priya Goyal, Piotr Dollár, Ross Girshick, Pieter Noordhuis, Lukasz Wesolowski, Aapo Kyrola, Andrew Tulloch, Yangqing Jia, and Kaiming He · 2017
Cited alongside, same era.
Global optimality in neural network training
Benjamin D Haeffele and René Vidal · 2017
Cited alongside, same era.
Exploring generalization in deep learning
Behnam Neyshabur, Srinadh Bhojanapalli, David McAllester, and Nati Srebro · 2017
Closest in time.
The loss surface of deep and wide neural networks
Quynh Nguyen and Matthias Hein · 2017
Closest in time.
Exploring loss function topology with cyclical learning rates
Leslie N Smith and Nicholay Topin · 2017
Closest in time.
Theoretical insights into the optimization landscape of over-parameterized shallow neural networks
Mahdi Soltanolkotabi, Adel Javanmard, and Jason D Lee · 2017
Closest in time.
Exponentially vanishing sub-optimal local minima in multilayer neural networks
Daniel Soudry and Elad Hoffer · 2017
Closest in time.
An analytical formula of population gradient for two-layered relu network and its applications in convergence and critical point analysis
Yuandong Tian · 2017
Closest in time.
Diverse neural network learns true target functions
Bo Xie, Yingyu Liang, and Le Song · 2017
Closest in time.
Global optimality conditions for deep neural networks
Chulhee Yun, Suvrit Sra, and Ali Jadbabaie · 2017
Closest in time.
Understanding deep learning requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2017
Closest in time.