Fetching the paper…
Reading the bibliography…
A common lens to theoretically study neural net architectures is to analyze the functions they can approximate.
Data-dependent sample complexity of deep neural networks via lipschitz augmentation
C. Wei and T. Ma · 1905
Earlier work this paper cites.
Improved sample complexities for deep networks and robust classification via an all-layer margin
C. Wei and T. Ma · 1910
Earlier work this paper cites.
Approximation theory and methods
M. J. D. Powell et al · 1981
Earlier work this paper cites.
Approximation by superpositions of a sigmoidal function
G. Cybenko · 1989
Earlier work this paper cites.
Multilayer feedforward networks are universal approximators
K. Hornik, M. Stinchcombe, and H. White · 1989
Earlier work this paper cites.
Universal approximation using radial-basis-function networks
J. Park and I. W. Sandberg · 1991
Earlier work this paper cites.
Universal approximation bounds for superpositions of a sigmoidal function
A. R. Barron · 1993
Earlier work this paper cites.
Multilayer feedforward networks with a nonpolynomial activation function can approximate any function
M. Leshno, V. Y. Lin, A. Pinkus, and S. Schocken · 1993
Earlier work this paper cites.
On the computational power of neural nets
H. T. Siegelmann and E. D. Sontag · 1995
Earlier work this paper cites.
For valid generalization the size of the weights is more important than the size of the network
P. Bartlett · 1996
Earlier work this paper cites.
Vapnik-chervonenkis dimension of recurrent neural networks
P. Koiran and E. D. Sontag · 1998
Earlier work this paper cites.
Models of computation - exploring the power of computing
J. Savage · 1998
Earlier work this paper cites.
The space complexity of approximating the frequency moments
N. Alon, Y. Matias, and M. Szegedy · 1999
Earlier work this paper cites.
Recurrent neural networks are universal approximators
A. M. Schäfer and H.-G. Zimmermann · 2007
Earlier work this paper cites.
Computational capabilities of graph neural networks
F. Scarselli, M. Gori, A. C. Tsoi, M. Hagenbuchner, and G. Monfardini · 2008
Earlier work this paper cites.
On the expressive power of deep architectures
Y. Bengio and O. Delalleau · 2011
Earlier work this paper cites.
Introduction to the Theory of Computation
M. Sipser · 2013
Earlier work this paper cites.
Approximation analysis of convolutional neural networks
C. Bao, Q. Li, Z. Shen, C. Tai, L. Wu, and X. Xiang · 2014
Earlier work this paper cites.
Norm-based capacity control in neural networks
B. Neyshabur, R. Tomioka, and N. Srebro · 2015
Earlier work this paper cites.
Understanding deep neural networks with rectified linear units
R. Arora, A. Basu, P. Mianjy, and A. Mukherjee · 2016
Earlier work this paper cites.
Convolutional rectifier networks as generalized tensor decompositions
N. Cohen and A. Shashua · 2016
Earlier work this paper cites.
On the expressive power of deep learning: A tensor analysis
N. Cohen, O. Sharir, and A. Shashua · 2016
Earlier work this paper cites.
The power of depth for feedforward neural networks
R. Eldan and O. Shamir · 2016
Cited alongside, same era.
Identity matters in deep learning
M. Hardt and T. Ma · 2016
Cited alongside, same era.
benefits of depth in neural networks
M. Telgarsky · 2016
Cited alongside, same era.
Understanding deep learning requires rethinking generalization
C. Zhang, S. Bengio, M. Hardt, B. Recht, and O. Vinyals · 2016
Cited alongside, same era.
Spectrally-normalized margin bounds for neural networks
P. Bartlett, D. J. Foster, and M. Telgarsky · 2017
Cited alongside, same era.
Depth separation for neural networks
Limitations of lazy training of two-layers neural network
B. Ghorbani, S. Mei, T. Misiakiewicz, and A. Montanari · 2019
Later among the works it cites.
On the computational power of rnns
S. A. Korsky and R. C. Berwick · 2019
Later among the works it cites.
Deterministic pac-bayesian generalization bounds for deep networks via generalizing noise-resilience
V. Nagarajan and J. Z. Kolter · 2019
Later among the works it cites.
On the turing completeness of modern neural network architectures
J. Pérez, J. Marinković, and P. Barceló · 2019
Later among the works it cites.
Universal approximations of permutation invariant/equivariant functions by deep neural networks
A. Sannai, Y. Takai, and M. Cordonnier · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A. Daniely · 2017
Cited alongside, same era.
Nearly-tight vc-dimension bounds for piecewise linear neural networks
N. Harvey, C. Liaw, and A. Mehrabian · 2017
Cited alongside, same era.
On the ability of neural nets to express distributions
H. Lee, R. Ge, T. Ma, A. Risteski, and S. Arora · 2017
Cited alongside, same era.
A pac-bayesian approach to spectrally-normalized margin bounds for neural networks
B. Neyshabur, S. Bhojanapalli, and N. Srebro · 2017
Cited alongside, same era.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin · 2017
Cited alongside, same era.
Recurrent neural networks as weighted language recognizers
Y. Chen, S. Gilroy, A. Maletti, J. May, and K. Knight · 2018
Cited alongside, same era.
Gradient descent provably optimizes over-parameterized neural networks
S. S. Du, X. Zhai, B. Poczos, and A. Singh · 2018
Cited alongside, same era.
Later among the works it cites.
Regularization matters: Generalization and optimization of neural nets vs their induced kernel
C. Wei, J. D. Lee, Q. Liu, and T. Ma · 2019
Later among the works it cites.
Are transformers universal approximators of sequence-to-sequence functions?
C. Yun, S. Bhojanapalli, A. S. Rawat, S. J. Reddi, and S. Kumar · 2019
Later among the works it cites.
On the computational power of transformers and its implications in sequence modeling
S. Bhattamishra, A. Patel, and N. Goyal · 2020
Later among the works it cites.
Better depth-width trade-offs for neural networks through the lens of dynamical systems
V. Chatziafratis, S. G. Nagarajan, and I. Panageas · 2020
Later among the works it cites.
Rnns can generate bounded hierarchical languages with optimal memory
J. Hewitt, M. Hahn, S. Ganguli, P. Liang, and C. D. Manning · 2020
Later among the works it cites.
Neural tangent kernels, transportation mappings, and universal approximation
Z. Ji, M. Telgarsky, and R. Xian · 2020
Later among the works it cites.
Why do local methods solve nonconvex problems?
T. Ma · 2020
Later among the works it cites.
Provable memorization via deep neural networks using sub-linear parameters
S. Park, J. Lee, C. Yun, and J. Shin · 2020
Later among the works it cites.
Universal approximation property of neural ordinary differential equations
T. Teshima, K. Tojo, M. Ikeda, I. Ishikawa, and K. Oono · 2020
Later among the works it cites.
Neural networks with small weights and depth-separation barriers
G. Vardi and O. Shamir · 2020
Later among the works it cites.
Memory capacity of neural networks with threshold and relu activations
R. Vershynin · 2020
Later among the works it cites.
Approximation capabilities of neural odes and invertible residual networks
H. Zhang, X. Gao, J. Unterman, and T. Arodz · 2020
Later among the works it cites.
Universality of deep convolutional neural networks
D.-X. Zhou · 2020
Later among the works it cites.
Turing completeness of bounded-precision recurrent neural networks
S. Chung and H. Siegelmann · 2021
Closest in time.
The connection between approximation, depth separation and learnability in neural networks
E. Malach, G. Yehudai, S. Shalev-Shwartz, and O. Shamir · 2021
Closest in time.
Universal approximations of invariant maps by neural networks
D. Yarotsky · 2021
Closest in time.