Fetching the paper…
Reading the bibliography…
Deep neural networks can empirically perform efficient hierarchical learning, in which the layers learn useful representations of the data.
S. Arora, S. S. Du, W. Hu, Z. Li, and R. Wang · 1901
Earlier work this paper cites.
Linearized two-layers neural networks in high dimension
B. Ghorbani, S. Mei, T. Misiakiewicz, and A. Montanari · 1904
Earlier work this paper cites.
Computational limitations of small-depth circuits
J. Håstad · 1987
Earlier work this paper cites.
Capabilities of three-layered perceptrons
B. Irie and S. Miyake · 1988
Earlier work this paper cites.
Approximation by superpositions of a sigmoidal function
G. Cybenko · 1989
Earlier work this paper cites.
On the approximate realization of continuous mappings by neural networks
K.-I. Funahashi · 1989
Earlier work this paper cites.
Approximation capabilities of multilayer feedforward networks
K. Hornik · 1991
Earlier work this paper cites.
Approximation by ridge functions and neural networks with one hidden layer
C. K. Chui and X. Li · 1992
Earlier work this paper cites.
Universal approximation bounds for superpositions of a sigmoidal function
A. R. Barron · 1993
Earlier work this paper cites.
Multilayer feedforward networks with a nonpolynomial activation function can approximate any function
M. Leshno, V. Y. Lin, A. Pinkus, and S. Schocken · 1993
Earlier work this paper cites.
Random approximants and neural networks
Y. Makovoz · 1996
Earlier work this paper cites.
Neural networks for optimal approximation of smooth and analytic functions
H. N. Mhaskar · 1996
Earlier work this paper cites.
Agnostically learning halfspaces
A. T. Kalai, A. R. Klivans, Y. Mansour, and R. A. Servedio · 2008
Earlier work this paper cites.
Introduction to the non-asymptotic analysis of random matrices
R. Vershynin · 2010
Earlier work this paper cites.
Shallow vs. deep sum-product networks
O. Delalleau and Y. Bengio · 2011
Earlier work this paper cites.
Learning polynomials with neural networks
A. Andoni, R. Panigrahy, G. Valiant, and L. Zhang · 2014
Earlier work this paper cites.
Rich feature hierarchies for accurate object detection and semantic segmentation
R. Girshick, J. Donahue, T. Darrell, and J. Malik · 2014
Earlier work this paper cites.
Square deal: Lower bounds and improved relaxations for tensor recovery
C. Mu, B. Huang, J. Wright, and D. Goldfarb · 2014
Earlier work this paper cites.
Deep learning
Y. LeCun, Y. Bengio, and G. Hinton · 2015
Cited alongside, same era.
An average-case depth hierarchy theorem for boolean circuits
B. Rossman, R. A. Servedio, and L.-Y. Tan · 2015
Cited alongside, same era.
The power of depth for feedforward neural networks
R. Eldan and O. Shamir · 2016
Cited alongside, same era.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Cited alongside, same era.
Gradient descent only converges to minimizers
J. D. Lee, M. Simchowitz, M. I. Jordan, and B. Recht · 2016
Cited alongside, same era.
Benefits of depth in neural networks
M. Telgarsky · 2016
Cited alongside, same era.
Generalization error bounds of gradient descent for learning overparameterized deep relu networks
Y. Cao and Q. Gu · 2019
Later among the works it cites.
Efficient approximation of deep relu networks for functions on low dimensional manifolds
M. Chen, H. Jiang, W. Liao, and T. Zhao · 2019
Later among the works it cites.
On lazy training in differentiable programming
L. Chizat, E. Oyallon, and F. Bach · 2019
Later among the works it cites.
Gradient descent finds global minima of deep neural networks
S. Du, J. Lee, H. Li, L. Wang, and X. Zhai · 2019
Later among the works it cites.
Asymptotics of wide networks from feynman diagrams
E. Dyer and G. Gur-Ari · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Error bounds for approximations with deep relu networks
D. Yarotsky · 2017
Cited alongside, same era.
A convergence theory for deep learning via over-parameterization
Z. Allen-Zhu, Y. Li, and Z. Song · 2018
Cited alongside, same era.
A note on lazy training in supervised differentiable programming
L. Chizat and F. Bach · 2018
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova · 2018
Cited alongside, same era.
Neural tangent kernel: Convergence and generalization in neural networks
A. Jacot, F. Gabriel, and C. Hongler · 2018
Cited alongside, same era.
Learning overparameterized neural networks via stochastic gradient descent on structured data
Y. Li and Y. Liang · 2018
Cited alongside, same era.
J. Huang and H.-T. Yau · 2019
Later among the works it cites.
Stochastic gradient descent escapes saddle points efficiently
C. Jin, P. Netrapalli, R. Ge, S. M. Kakade, and M. I. Jordan · 2019
Later among the works it cites.
Wide neural networks of any depth evolve as linear models under gradient descent
J. Lee, L. Xiao, S. Schoenholz, Y. Bahri, R. Novak, J. Sohl-Dickstein, and J. Pennington · 2019
Later among the works it cites.
High-dimensional statistics: A non-asymptotic viewpoint , volume 48
M. J. Wainwright · 2019
Later among the works it cites.
Regularization matters: Generalization and optimization of neural nets vs their induced kernel
C. Wei, J. D. Lee, Q. Liu, and T. Ma · 2019
Later among the works it cites.
On the power and limitations of random features for understanding neural networks
G. Yehudai and O. Shamir · 2019
Later among the works it cites.
Backward feature correction: How deep learning performs deep learning
Z. Allen-Zhu and Y. Li · 2020
Closest in time.
Beyond linearization: On quadratic and higher-order approximation of wide neural networks
Y. Bai and J. D. Lee · 2020
Closest in time.
Taylorized training: Towards better approximation of neural network training at finite width
Y. Bai, B. Krause, H. Wang, C. Xiong, and R. Socher · 2020
Closest in time.
Learning polynomials of few relevant dimensions
S. Chen and R. Meka · 2020
Closest in time.
K. Huang, Y. Wang, M. Tao, and T. Zhao · 2020
Closest in time.
Kernel and rich regimes in overparametrized models
B. Woodworth, S. Gunasekar, J. D. Lee, E. Moroshko, P. Savarese, I. Golan, D. Soudry, and N. Srebro · 2020
Closest in time.