Fetching the paper…
Reading the bibliography…
How can local-search methods such as stochastic gradient descent (SGD) avoid bad local minima in training multi-layer neural networks? Why can they fit random labels even given non-convex and non-smooth architectures? Most existing theory only covers networks with one hidden layer, so can we go deeper? In this paper, we focus on recurrent neural networks (RNNs) which are multi-layer networks widely used in natural language processing.
Can SGD Learn Recurrent Neural Networks with Provable Generalization?
Zeyuan Allen-Zhu and Yuanzhi Li · 1902
Earlier work this paper cites.
A topological property of real analytic subsets
S Lojasiewicz · 1963
Earlier work this paper cites.
Gradient methods for minimizing functionals
Boris Teodorovich Polyak · 1963
Earlier work this paper cites.
Finding structure in time
Jeffrey L Elman · 1990
Earlier work this paper cites.
Turing computability with neural nets
Hava T Siegelmann and Eduardo D Sontag · 1991
Earlier work this paper cites.
A model of multiplicative neural responses in parietal cortex
Emilio Salinas and Laurence F. Abbott · 1996
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
On the piecewise analysis of networks of linear threshold neurons
Richard LT Hahnloser · 1998
Earlier work this paper cites.
Adaptive estimation of a quadratic functional by model selection
Beatrice Laurent and Pascal Massart · 2000
Earlier work this paper cites.
Smoothed analysis of algorithms: Why the simplex algorithm usually takes polynomial time
Daniel A. Spielman and Shang-Hua Teng · 2004
Earlier work this paper cites.
Recurrent neural network based language model
Tomáš Mikolov, Martin Karafiát, Lukáš Burget, Jan Černockỳ, and Sanjeev Khudanpur · 2010
Earlier work this paper cites.
Extensions of recurrent neural network language model
Tomáš Mikolov, Stefan Kombrink, Lukáš Burget, Jan Černockỳ, and Sanjeev Khudanpur · 2011
Earlier work this paper cites.
Generating text with recurrent neural networks
Ilya Sutskever, James Martens, and Geoffrey Hinton · 2011
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton · 2012
Earlier work this paper cites.
Lstm neural networks for language modeling
Martin Sundermeyer, Ralf Schlüter, and Hermann Ney · 2012
Earlier work this paper cites.
Speech recognition with deep recurrent neural networks
Alex Graves, Abdel-rahman Mohamed, and Geoffrey Hinton · 2013
Earlier work this paper cites.
Recurrent continuous translation models
Nal Kalchbrenner and Phil Blunsom · 2013
Earlier work this paper cites.
Provable bounds for learning some deep representations
Sanjeev Arora, Aditya Bhaskara, Rong Ge, and Tengyu Ma · 2014
Earlier work this paper cites.
On the properties of neural machine translation: Encoder-decoder approaches
Kyunghyun Cho, Bart Van Merriënboer, Dzmitry Bahdanau, and Yoshua Bengio · 2014
Earlier work this paper cites.
Empirical evaluation of gated recurrent neural networks on sequence modeling
Junyoung Chung, Caglar Gulcehre, KyungHyun Cho, and Yoshua Bengio · 2014
Earlier work this paper cites.
Alex Graves, Greg Wayne, and Ivo Danihelka · 2014
Earlier work this paper cites.
Long short-term memory recurrent neural network architectures for large scale acoustic modeling
Haşim Sak, Andrew Senior, and Françoise Beaufays · 2014
Earlier work this paper cites.
Sequence to sequence learning with neural networks
Ilya Sutskever, Oriol Vinyals, and Quoc V Le · 2014
Earlier work this paper cites.
Qualitatively characterizing neural network optimization problems
Ian J Goodfellow, Oriol Vinyals, and Andrew M Saxe · 2015
Cited alongside, same era.
Beating the perils of non-convexity: Guaranteed training of neural networks using tensor methods
Majid Janzamin, Hanie Sedghi, and Anima Anandkumar · 2015
Cited alongside, same era.
Continuous control with deep reinforcement learning
Timothy P Lillicrap, Jonathan J Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra · 2015
Cited alongside, same era.
Provable methods for training neural networks with sparse connectivity
Hanie Sedghi and Anima Anandkumar · 2015
Cited alongside, same era.
Deep speech 2: End-to-end speech recognition in English and Mandarin
Dario Amodei, Sundaram Ananthanarayanan, Rishita Anubhai, Jingliang Bai, Eric Battenberg, Carl Case, Jared Casper, Bryan Catanzaro, Qiang Cheng, Guoliang Chen, et al · 2016
Learning ReLUs via gradient descent
Mahdi Soltanolkotabi · 2017
Later among the works it cites.
Yuandong Tian · 2017
Later among the works it cites.
Mean field residual networks: On the edge of chaos
Greg Yang and Samuel Schoenholz · 2017
Later among the works it cites.
Understanding deep learning requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2017
Later among the works it cites.
Critical points of neural networks: Analytical forms and landscape properties
Yi Zhou and Yingbin Liang · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Toward deeper understanding of neural networks: The power of initialization and a dual view on expressivity
Amit Daniely, Roy Frostig, and Yoram Singer · 2016
Cited alongside, same era.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Cited alongside, same era.
Exponential expressivity in deep neural networks through transient chaos
Ben Poole, Subhaneil Lahiri, Maithreyi Raghu, Jascha Sohl-Dickstein, and Surya Ganguli · 2016
Cited alongside, same era.
Mastering the game of Go with deep neural networks and tree search
David Silver, Aja Huang, Chris J Maddison, Arthur Guez, Laurent Sifre, George Van Den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Veda Panneershelvam, Marc Lanctot, et al · 2016
Cited alongside, same era.
Globally optimal gradient descent for a convnet with gaussian inputs
Alon Brutzkus and Amir Globerson · 2017
Cited alongside, same era.
SGD learns the conjugate kernel class of the network
Amit Daniely · 2017
Cited alongside, same era.
Topology and geometry of half-rectified network optimization
C. Daniel Freeman and Joan Bruna · 2017
Cited alongside, same era.
Learning and Generalization in Overparameterized Neural Networks, Going Beyond Two Layers
Zeyuan Allen-Zhu, Yuanzhi Li, and Yingyu Liang · 2018
Closest in time.
A convergence analysis of gradient descent for deep linear neural networks
Sanjeev Arora, Nadav Cohen, Noah Golowich, and Wei Hu · 2018
Closest in time.
Gradient descent with identity initialization efficiently learns positive definite linear transformations
Peter Bartlett, Dave Helmbold, and Phil Long · 2018
Closest in time.
SGD learns over-parameterized networks that provably generalize on linearly separable data
Alon Brutzkus, Amir Globerson, Eran Malach, and Shai Shalev-Shwartz · 2018
Closest in time.
Minmin Chen, Jeffrey Pennington, and Samuel S. Schoenholz · 2018
Closest in time.
Gradient descent learns linear dynamical systems
Moritz Hardt, Tengyu Ma, and Benjamin Recht · 2018
Closest in time.
Spectral filtering for general linear dynamical systems
Elad Hazan, Holden Lee, Karan Singh, Cyril Zhang, and Yi Zhang · 2018
Closest in time.
Expressive power of recurrent neural networks
Valentin Khrulkov, Alexander Novikov, and Ivan Oseledets · 2018
Closest in time.
Learning overparameterized neural networks via stochastic gradient descent on structured data
Yuanzhi Li and Yingyu Liang · 2018
Closest in time.
A mean field view of the landscape of two-layer neural networks
Song Mei, Andrea Montanari, and Phan-Minh Nguyen · 2018
Closest in time.
Convergence results for neural networks via electrodynamics
Rina Panigrahy, Ali Rahimi, Sushant Sachdeva, and Qiuyi Zhang · 2018
Closest in time.
Spurious local minima are common in two-layer ReLU neural networks
Itay Safran and Ohad Shamir · 2018
Closest in time.
Polynomial convergence of gradient descent for training one-hidden-layer neural networks
Santosh Vempala and John Wilmes · 2018
Closest in time.
Dynamical isometry and a mean field theory of CNNs: How to train 10,000-layer vanilla convolutional neural networks
Lechao Xiao, Yasaman Bahri, Jascha Sohl-Dickstein, Samuel S. Schoenholz, and Jeffrey Pennington · 2018
Closest in time.
Deep mean field theory: Layerwise variance and width variation as methods to control gradient explosion
Greg Yang and Sam S. Schoenholz · 2018
Closest in time.
A convergence theory for deep learning via over-parameterization
Zeyuan Allen-Zhu, Yuanzhi Li, and Zhao Song · 2019
Closest in time.
A Generalization Theory of Gradient Descent for Learning Over-parameterized Deep ReLU Networks
Yuan Cao and Quanquan Gu · 2019
Closest in time.