Fetching the paper…
Reading the bibliography…
We describe a layer-by-layer algorithm for training deep convolutional networks, where each step involves gradient updates for a two layer network followed by a simple clustering algorithm.
On the problem of local minima in backpropagation
Marco Gori and Alberto Tesi · 1992
Earlier work this paper cites.
Deep mixtures of factor analysers
Yichuan Tang, Ruslan Salakhutdinov, and Geoffrey Hinton · 2012
Earlier work this paper cites.
Learning polynomials with neural networks
Alexandr Andoni, Rina Panigrahy, Gregory Valiant, and Li Zhang · 2014
Earlier work this paper cites.
Provable bounds for learning some deep representations
Sanjeev Arora, Aditya Bhaskara, Rong Ge, and Tengyu Ma · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
On the computational efficiency of training neural networks
Roi Livni, Shai Shalev-Shwartz, and Ohad Shamir · 2014
Earlier work this paper cites.
Factoring variations in natural images with deep gaussian mixture models
Aaron Van den Oord and Benjamin Schrauwen · 2014
Earlier work this paper cites.
Learning halfspaces and neural networks with random initialization
Yuchen Zhang, Jason D Lee, Martin J Wainwright, and Michael I Jordan · 2015
Cited alongside, same era.
Tensorflow
G. Google-Brain · 2016
Cited alongside, same era.
Deep learning and hierarchal generative models
Elchanan Mossel · 2016
Cited alongside, same era.
A probabilistic framework for deep learning
Ankit B Patel, Minh Tan Nguyen, and Richard Baraniuk · 2016
Cited alongside, same era.
l1-regularized neural networks are improperly learnable in polynomial time
Yuchen Zhang, Jason D Lee, and Michael I Jordan · 2016
Cited alongside, same era.
Sgd learns over-parameterized networks that provably generalize on linearly separable data
Alon Brutzkus, Amir Globerson, Eran Malach, and Shai Shalev-Shwartz · 2017
Later among the works it cites.
Sgd learns the conjugate kernel class of the network
Amit Daniely · 2017
Later among the works it cites.
When is a convolutional filter easy to learn?
Simon S Du, Jason D Lee, and Yuandong Tian · 2017
Later among the works it cites.
Convergence analysis of two-layer neural networks with relu activation
Yuanzhi Li and Yang Yuan · 2017
Later among the works it cites.
Yuandong Tian · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yuchen Zhang, Percy Liang, and Martin J Wainwright · 2016
Cited alongside, same era.
Globally optimal gradient descent for a convnet with gaussian inputs
Alon Brutzkus and Amir Globerson · 2017
Cited alongside, same era.
Recovery guarantees for one-hidden-layer neural networks
Kai Zhong, Zhao Song, Prateek Jain, Peter L Bartlett, and Inderjit S Dhillon · 2017
Later among the works it cites.