Fetching the paper…
Reading the bibliography…
It is well-known that overparametrized neural networks trained using gradient-based methods quickly achieve small training error with appropriate hyperparameter settings.
Sanjeev Arora, Simon S. Du, Wei Hu, Zhi yuan Li, and Ruosong Wang · 1901
Earlier work this paper cites.
On the impact of the activation function on deep neural networks training
Soufiane Hayou, Arnaud Doucet, and Judith Rousseau · 1902
Earlier work this paper cites.
Das asymptotische Verteilungsgesetz der Eigenwerte linearer partieller Differentialgleichungen (mit einer Anwendung auf die Theorie der Hohlraumstrahlung)
Hermann Weyl · 1912
Earlier work this paper cites.
Contributions to the theory of Hermitian series. II. The representation problem
Einar Hille · 1940
Earlier work this paper cites.
Solutions to some functional equations and their applications to characterization of probability distributions
CG Khatri and C Radhakrishna Rao · 1968
Earlier work this paper cites.
Special functions and their applications
N. N. Lebedev · 1972
Earlier work this paper cites.
Orthogonal polynomials
Gábor Szegő · 1975
Earlier work this paper cites.
Asymptotic coefficients of hermite function series
John P Boyd · 1984
Earlier work this paper cites.
Multilayer feedforward networks with a nonpolynomial activation function can approximate any function
Moshe Leshno, Vladimir Ya. Lin, Allan Pinkus, and Shimon Schocken · 1993
Earlier work this paper cites.
Lectures on Hermite and Laguerre expansions , volume 42 of Mathematical Notes
Sundaram Thangavelu · 1993
Earlier work this paper cites.
Approximation theory of the MLP model in neural networks
Allan Pinkus · 1999
Earlier work this paper cites.
Adaptive estimation of a quadratic functional by model selection
Beatrice Laurent and Pascal Massart · 2000
Earlier work this paper cites.
Distributional and L q L^{q} norm inequalities for polynomials over convex bodies in ℝ n \mathbb{R}^{n}
Anthony Carbery and James Wright · 2001
Earlier work this paper cites.
Chebyshev Polynomials
J.C. Mason and D.C. Handscomb · 2002
Earlier work this paper cites.
Smoothed analysis of algorithms: Why the simplex algorithm usually takes polynomial time
Daniel A. Spielman and Shang-Hua Teng · 2004
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky and Geoffrey Hinton · 2009
Earlier work this paper cites.
Smallest singular value of a random rectangular matrix
Mark Rudelson and Roman Vershynin · 2009
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
Xavier Glorot and Yoshua Bengio · 2010
Earlier work this paper cites.
Geršgorin and his circles , volume 36
Richard S Varga · 2010
Earlier work this paper cites.
Learning kernel-based halfspaces with the 0-1 loss
Shai Shalev-Shwartz, Ohad Shamir, and Karthik Sridharan · 2011
Cited alongside, same era.
Concentration inequalities
Stéphane Boucheron, Gábor Lugosi, and Pascal Massart · 2013
Cited alongside, same era.
The more, the merrier: the blessing of dimensionality for learning large gaussian mixtures
Joseph Anderson, Mikhail Belkin, Navin Goyal, Luis Rademacher, and James R. Voss · 2014
Cited alongside, same era.
Analysis of boolean functions
Ryan O’Donnell · 2014
Cited alongside, same era.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2015
Cited alongside, same era.
In search of the real inductive bias: On the role of implicit regularization in deep learning
On the global convergence of gradient descent for over-parameterized models using optimal transport
Lénaïc Chizat and Francis Bach · 2018
Later among the works it cites.
On the power of over-parametrization in neural networks with quadratic activation
Simon S. Du and Jason D. Lee · 2018
Later among the works it cites.
Is it time to swish? comparing deep learning activation functions across nlp tasks
Steffen Eger, Paul Youssef, and Iryna Gurevych · 2018
Later among the works it cites.
Which neural net architectures give rise to exploding and vanishing gradients?
Boris Hanin · 2018
Later among the works it cites.
How to start training: The effect of initialization and architecture
Boris Hanin and David Rolnick · 2018
Later among the works it cites.
Neural tangent kernel: Convergence and generalization in neural networks
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Behnam Neyshabur, Ryota Tomioka, and Nathan Srebro · 2015
Cited alongside, same era.
Neural network with unbounded activation functions is universal approximator
Sho Sonoda and Noboru Murata · 2015
Cited alongside, same era.
Fast and accurate deep network learning by exponential linear units (elus)
Djork-Arné Clevert, Thomas Unterthiner, and Sepp Hochreiter · 2016
Cited alongside, same era.
Toward deeper understanding of neural networks: The power of initialization and a dual view on expressivity
Amit Daniely, Roy Frostig, and Yoram Singer · 2016
Cited alongside, same era.
Deep Learning
Ian Goodfellow, Yoshua Bengio, and Aaron Courville · 2016
Cited alongside, same era.
Revise saturated activation functions
Bing Xu, Ruitong Huang, and Mu Li · 2016
Cited alongside, same era.
Sigmoid-weighted linear units for neural network function approximation in reinforcement learning
Stefan Elfwing, Eiji Uchibe, and Kenji Doya · 2017
Cited alongside, same era.
Arthur Jacot, Franck Gabriel, and Clement Hongler · 2018
Later among the works it cites.
Learning overparameterized neural networks via stochastic gradient descent on structured data
Yuanzhi Li and Yingyu Liang · 2018
Later among the works it cites.
A random matrix approach to neural networks
Cosme Louart, Zhenyu Liao, and Romain Couillet · 2018
Later among the works it cites.
A mean field view of the landscape of two-layer neural networks
Song Mei, Andrea Montanari, and Phan-Minh Nguyen · 2018
Later among the works it cites.
Activation functions: Comparison of trends in practice and research for deep learning
Chigozie Nwankpa, Winifred Ijomah, Anthony Gachagan, and Stephen Marshall · 2018
Later among the works it cites.
The spectrum of the fisher information matrix of a single-hidden-layer neural network
Jeffrey Pennington and Pratik Worah · 2018
Later among the works it cites.
The emergence of spectral universality in deep networks
Jeffrey Pennington, Samuel S. Schoenholz, and Surya Ganguli · 2018
Later among the works it cites.
Searching for activation functions
Prajit Ramachandran, Barret Zoph, and Quoc V. Le · 2018
Later among the works it cites.
Invariance of weight distributions in rectified mlps
Russell Tsuchida, Farbod Roosta-Khorasani, and Marcus Gallagher · 2018
Later among the works it cites.
A convergence theory for deep learning via over-parameterization
Zeyuan Allen-Zhu, Yuanzhi Li, and Zhao Song · 2019
Closest in time.
A convergence analysis of gradient descent for deep linear neural networks
Sanjeev Arora, Nadav Cohen, Noah Golowich, and Wei Hu · 2019
Closest in time.
On the expressive power of deep polynomial neural networks
Joe Kileel, Matthew Trager, and Joan Bruna · 2019
Closest in time.
Samet Oymak and Mahdi Soltanolkotabi · 2019
Closest in time.