Fetching the paper…
Reading the bibliography…
We use smoothed analysis techniques to provide guarantees on the training loss of Multilayer Neural Networks (MNNs) at differentiable local minima.
Learning machines
Nils J. Nilsson · 1965
Earlier work this paper cites.
Directionally Lipschitzian Functions and Subdifferential Calculus
R T Rockafellarf · 1979
Earlier work this paper cites.
On the capabilities of multilayer perceptrons
Eric B. Baum · 1988
Earlier work this paper cites.
Nonconvergence to unstable points in urn models and stochastic approximations
R Pemantle · 1990
Earlier work this paper cites.
On the problem of local minima in backpropagation, 1992
Marco Gori and Alberto Tesi · 1992
Earlier work this paper cites.
Online learning and stochastic approximations
L Bottou · 1998
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Y LeCun, L Bottou, Y Bengio, and P Haffner · 1998
Earlier work this paper cites.
Local minima and plateaus in hierarchical structures of multilayer perceptrons
K. Fukumizu and S. Amari · 2000
Earlier work this paper cites.
Training a single sigmoidal neuron is hard
Jirí Síma · 2002
Earlier work this paper cites.
Extreme learning machine: Theory and applications
Guang-Bin Huang, Qin-Yu Zhu, and Chee-Kheong Siew · 2006
Earlier work this paper cites.
Identifiability of parameters in latent structure models with many observed variables
Elizabeth S. Allman, Catherine Matias, and John A. Rhodes · 2009
Earlier work this paper cites.
Smoothed analysis: an attempt to explain the behavior of algorithms in practice
Daniel A Spielman and Shang-hua Teng · 2009
Cited alongside, same era.
Improving neural networks by preventing co-adaptation of feature detectors
Geoffrey E Hinton, Nitish Srivastava, Alex Krizhevsky, Ilya Sutskever, and Ruslan R. Salakhutdinov · 2012
Cited alongside, same era.
Imagenet classification with deep convolutional neural networks
A Krizhevsky, I Sutskever, and G E Hinton · 2012
Cited alongside, same era.
Smoothed Analysis of Tensor Decompositions
Aditya Bhaskara, Moses Charikar, Ankur Moitra, and Aravindan Vijayaraghavan · 2013
Cited alongside, same era.
Learning Polynomials with Neural Networks
A Andoni, R Panigrahy, G Valiant, and L Zhang · 2014
Cited alongside, same era.
Global Optimality in Tensor Factorization, Deep Learning, and Beyond
Benjamin D Haeffele and René Vidal · 2015
Later among the works it cites.
Train faster, generalize better: Stability of stochastic gradient descent
Moritz Hardt, Benjamin Recht, and Y Singer · 2015
Later among the works it cites.
Deep Residual Learning for Image Recognition
K He, X Zhang, S Ren, and J. Sun · 2015
Later among the works it cites.
Delving Deep into Rectifiers: Surpassing Human-Level Performance on ImageNet Classification
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2015
Later among the works it cites.
Beating the Perils of Non-Convexity: Guaranteed Training of Neural Networks using Tensor Methods
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
YN Dauphin, Razvan Pascanu, and Caglar Gulcehre · 2014
Cited alongside, same era.
On the Computational Efficiency of Training Neural Networks
Roi Livni, S Shalev-Shwartz, and Ohad Shamir · 2014
Cited alongside, same era.
Exact solutions to the nonlinear dynamics of learning in deep linear neural networks
A M Saxe, J L. McClelland, and S Ganguli · 2014
Cited alongside, same era.
Dropout : A Simple Way to Prevent Neural Networks from Overfitting
Nitish Srivastava, Geoffrey E. Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov · 2014
Cited alongside, same era.
The Loss Surfaces of Multilayer Networks
Anna Choromanska, Mikael Henaff, Michael Mathieu, Gérard Ben Arous, and Y LeCun · 2015
Cited alongside, same era.
Qualitatively characterizing neural network optimization problems
Ian J. Goodfellow, Oriol Vinyals, and Andrew M. Saxe · 2015
Cited alongside, same era.
M Janzamin, H Sedghi, and A Anandkumar · 2015
Later among the works it cites.
Adam: a Method for Stochastic Optimization
Diederik P Kingma and Jimmy Lei Ba · 2015
Later among the works it cites.
Deep learning
Y LeCun, Yoshua Bengio, and Geoffrey Hinton · 2015
Later among the works it cites.
On the Quality of the Initial Basin in Overspecified Neural Networks
Itay Safran and Ohad Shamir · 2015
Later among the works it cites.
Empirical Evaluation of Rectified Activations in Convolution Network
Bing Xu, Naiyan Wang, Tianqi Chen, and Mu Li · 2015
Later among the works it cites.
Gradient Descent Converges to Minimizers
Jason D. Lee, Max Simchowitz, Michael I. Jordan, and Benjamin Recht · 2016
Closest in time.