Fetching the paper…
Reading the bibliography…
The universal approximation theorem, in one of its most general versions, says that if we consider only continuous activation functions $\sigma$, then a standard feedforward neural network with one hidden layer is able to approximate any continuous multivariate function $f$ to any given approximation threshold $\varepsilon$, if and only if $\sigma$ is non-polynomial.
Topics in number theory. Vols. 1 and 2
William Judson LeVeque · 1956
Earlier work this paper cites.
On the power and limitations of random features for understanding neural networks
Gilad Yehudai and Ohad Shamir · 1956
Earlier work this paper cites.
On the distribution of points of maximum deviation in the approximation of continuous functions by polynomials
M. I. Kadec · 1960
Earlier work this paper cites.
On the distribution of points of maximum deviation in the approximation of continuous functions by polynomials
M. I. Kadec · 1963
Earlier work this paper cites.
An introduction to the approximation of functions
Theodore J. Rivlin · 1969
Earlier work this paper cites.
Approximation by superpositions of a sigmoidal function
G. Cybenko · 1989
Earlier work this paper cites.
Optimal nonlinear approximation
Ronald A. DeVore, Ralph Howard, and Charles Micchelli · 1989
Earlier work this paper cites.
On the approximate realization of continuous mappings by neural networks
K. Funahashi · 1989
Earlier work this paper cites.
Multilayer feedforward networks are universal approximators
K. Hornik, M. Stinchcombe, and H. White · 1989
Earlier work this paper cites.
Linear dependence of a function set of m m variables with vanishing generalized Wronskians
K. Wolsson · 1989
Earlier work this paper cites.
Approximation capabilities of multilayer feedforward networks
Kurt Hornik · 1991
Earlier work this paper cites.
Multilayer feedforward networks with a nonpolynomial activation function can approximate any function
Moshe Leshno, Vladimir Ya. Lin, Allan Pinkus, and Shimon Schocken · 1993
Cited alongside, same era.
Commutative algebra , volume 150 of Graduate Texts in Mathematics
David Eisenbud · 1995
Cited alongside, same era.
Neural networks for optimal approximation of smooth and analytic functions
H. N. Mhaskar · 1996
Cited alongside, same era.
Lower bounds for approximation by mlp neural networks
Vitaly Maiorov and Allan Pinkus · 1999
Cited alongside, same era.
Approximation theory of the mlp model in neural networks
Allan Pinkus · 1999
Cited alongside, same era.
Enumerative combinatorics. Vol. 2 , volume 62 of Cambridge Studies in Advanced Mathematics
Richard P. Stanley · 1999
On the number of linear regions of deep neural networks
Guido Montúfar, Razvan Pascanu, Kyunghyun Cho, and Yoshua Bengio · 2014
Later among the works it cites.
The power of depth for feedforward neural networks
Ronen Eldan and Ohad Shamir · 2016
Later among the works it cites.
Shiyu Liang and R. Srikant · 2016
Later among the works it cites.
benefits of depth in neural networks
Matus Telgarsky · 2016
Later among the works it cites.
Universal function approximation by deep neural nets with bounded width and relu activations
Boris Hanin · 2017
Later among the works it cites.
Why does deep and cheap learning work so well?
Henry W. Lin, Max Tegmark, and David Rolnick · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
On the near optimality of the stochastic approximation of smooth functions by neural networks
V. E. Maiorov and R. Meir · 2000
Cited alongside, same era.
Random features for large-scale kernel machines
Ali Rahimi and Benjamin Recht · 2007
Cited alongside, same era.
Uniform approximation of functions with random bases
A. Rahimi and B. Recht · 2008
Cited alongside, same era.
Tropicalization and irreducibility of generalized Vandermonde determinants
Carlos D’Andrea and Luis Felipe Tabera · 2009
Cited alongside, same era.
Shallow vs. deep sum-product networks
Olivier Delalleau and Yoshua Bengio · 2011
Cited alongside, same era.
Later among the works it cites.
The expressive power of neural networks: A view from the width
Zhou Lu, Hongming Pu, Feicheng Wang, Zhiqiang Hu, and Liwei Wang · 2017
Later among the works it cites.
When and why are deep networks better than shallow ones?, 2017
Hrushikesh Mhaskar, Qianli Liao, and Tomaso Poggio · 2017
Later among the works it cites.
Error bounds for approximations with deep relu networks
Dmitry Yarotsky · 2017
Later among the works it cites.
Understanding deep neural networks with rectified linear units
Raman Arora, Amitabh Basu, Poorya Mianjy, and Anirbit Mukherjee · 2018
Later among the works it cites.
On the approximation properties of random ReLU features
Yitong Sun, Anna Gilbert, and Ambuj Tewari · 2019
Later among the works it cites.