Fetching the paper…
Reading the bibliography…
It is widely believed that the practical success of Convolutional Neural Networks (CNNs) and Recurrent Neural Networks (RNNs) owes to the fact that CNNs and RNNs use a more compact parametric representation than their Fully-Connected Neural Network (FNN) counterparts, and consequently require fewer training examples to accurately estimate their parameters.
A new approach to linear filtering and prediction problems
Rudolph Emil Kalman · 1960
Earlier work this paper cites.
Probability inequalities for sums of bounded random variables
Wassily Hoeffding · 1963
Earlier work this paper cites.
The sizes of compact subsets of hilbert space and continuity of gaussian processes
R. M. Dudley · 1967
Earlier work this paper cites.
Lower bounds for constant weight codes
Ron Graham and Neil Sloane · 1980
Earlier work this paper cites.
Learning representations by back-propagating errors
David E Rumelhart, Geoffrey E Hinton, Ronald J Williams, et al · 1988
Earlier work this paper cites.
Convolutional networks for images, speech, and time series
Yann LeCun, Yoshua Bengio, et al · 1995
Earlier work this paper cites.
Empirical Processes in M-estimation , volume 6
Sara A van de Geer · 2000
Earlier work this paper cites.
Theory of point estimation
Erich L Lehmann and George Casella · 2006
Earlier work this paper cites.
Neural network learning: Theoretical foundations
Martin Anthony and Peter L Bartlett · 2009
Earlier work this paper cites.
Simultaneous analysis of lasso and dantzig selector
Peter J Bickel, Ya’acov Ritov, and Alexandre B Tsybakov · 2009
Earlier work this paper cites.
Introduction to nonparametric estimation, 2009
Alexandre B Tsybakov · 2009
Earlier work this paper cites.
Sharp thresholds for high-dimensional and noisy sparsity recovery using ℓ 1 \ell_{1} -constrained quadratic programming (lasso)
Martin J Wainwright · 2009
Earlier work this paper cites.
Recurrent neural network based language model
Tomáš Mikolov, Martin Karafiát, Lukáš Burget, Jan Černockỳ, and Sanjeev Khudanpur · 2010
Earlier work this paper cites.
All of statistics: a concise course in statistical inference
Larry Wasserman · 2013
Earlier work this paper cites.
The loss surfaces of multilayer networks
Anna Choromanska, Mikael Henaff, Michael Mathieu, Gérard Ben Arous, and Yann LeCun · 2015
Earlier work this paper cites.
Global optimality in tensor factorization, deep learning, and beyond
Benjamin D Haeffele and René Vidal · 2015
Earlier work this paper cites.
Norm-based capacity control in neural networks
Behnam Neyshabur, Ryota Tomioka, and Nathan Srebro · 2015
Cited alongside, same era.
Learning halfspaces and neural networks with random initialization
Yuchen Zhang, Jason D Lee, Martin J Wainwright, and Michael I Jordan · 2015
Cited alongside, same era.
Topology and geometry of half-rectified network optimization
C Daniel Freeman and Joan Bruna · 2016
Cited alongside, same era.
Reliably learning the ReLU in polynomial time
Surbhi Goel, Varun Kanade, Adam Klivans, and Justin Thaler · 2016
Cited alongside, same era.
Identity matters in deep learning
Moritz Hardt and Tengyu Ma · 2016
On the global geometry of sphere-constrained sparse blind deconvolution
Yuqian Zhang, Yenson Lau, Han-wen Kuo, Sky Cheung, Abhay Pasupathy, and John Wright · 2017
Later among the works it cites.
The landscape of deep learning algorithms
Pan Zhou and Jiashi Feng · 2017
Later among the works it cites.
Stronger generalization bounds for deep nets via a compression approach
Sanjeev Arora, Rong Ge, Behnam Neyshabur, and Yi Zhang · 2018
Closest in time.
On the power of over-parametrization in neural networks with quadratic activation
Simon S Du and Jason D Lee · 2018
Closest in time.
Learning one-hidden-layer neural networks with landscape design
Rong Ge, Jason D. Lee, and Tengyu Ma · 2018
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Deep learning without poor local minima
Kenji Kawaguchi · 2016
Cited alongside, same era.
On the quality of the initial basin in overspecified neural networks
Itay Safran and Ohad Shamir · 2016
Cited alongside, same era.
Noise-adaptive margin-based active learning and lower bounds under tsybakov noise condition
Yining Wang and Aarti Singh · 2016
Cited alongside, same era.
Globally optimal gradient descent for a convnet with gaussian inputs
Alon Brutzkus and Amir Globerson · 2017
Cited alongside, same era.
Learning linear dynamical systems via spectral filtering
Elad Hazan, Karan Singh, and Cyril Zhang · 2017
Cited alongside, same era.
PAC-Bayesian margin bounds for convolutional neural networks-technical report
Pitas Konstantinos, Mike Davies, and Pierre Vandergheynst · 2017
Cited alongside, same era.
Convergence analysis of two-layer neural networks with relu activation
Yuanzhi Li and Yang Yuan · 2017
Cited alongside, same era.
Surbhi Goel, Adam Klivans, and Raghu Meka · 2018
Closest in time.
Gradient descent learns linear dynamical systems
Moritz Hardt, Tengyu Ma, and Benjamin Recht · 2018
Closest in time.
On tighter generalization bound for deep neural networks: Cnns, resnets, and beyond
Xingguo Li, Junwei Lu, Zhaoran Wang, Jarvis Haupt, and Tuo Zhao · 2018
Closest in time.
Learning overparameterized neural networks via stochastic gradient descent on structured data
Yuanzhi Li and Yingyu Liang · 2018
Closest in time.
Non-asymptotic identification of lti systems from a single trajectory
Samet Oymak and Necmiye Ozay · 2018
Closest in time.
Learning without mixing: Towards a sharp analysis of linear system identification
Max Simchowitz, Horia Mania, Stephen Tu, Michael I Jordan, and Benjamin Recht · 2018
Closest in time.
Minimax reconstruction risk of convolutional sparse dictionary learning
Shashank Singh, Barnabás Póczos, and Jian Ma · 2018
Closest in time.
Stochastic gradient descent optimizes over-parameterized deep ReLU networks
Difan Zou, Yuan Cao, Dongruo Zhou, and Quanquan Gu · 2018
Closest in time.
Can SGD learn recurrent neural networks with provable generalization?
Zeyuan Allen-Zhu and Yuanzhi Li · 2019
Closest in time.
Sanjeev Arora, Simon S Du, Wei Hu, Zhiyuan Li, and Ruosong Wang · 2019
Closest in time.
Gradient descent provably optimizes over-parameterized neural networks
Simon S. Du, Xiyu Zhai, Barnabas Poczos, and Aarti Singh · 2019
Closest in time.