Fetching the paper…
Reading the bibliography…
We call a finite family of activation functions superexpressive if any multivariate continuous function can be approximated by a neural network that uses these activations and has a fixed architecture only depending on the number of input variables (i.e., to achieve any accuracy we only need to adjust the weights, without increasing the number of neurons).
On the representation of continuous functions of many variables by superposition of continuous functions of one variable and addition
Kolmogorov, A. N · 1957
Earlier work this paper cites.
Optimal nonlinear approximation
DeVore, R. A., Howard, R., and Micchelli, C · 1989
Earlier work this paper cites.
Fewnomials
Khovanskii, A. G · 1991
Earlier work this paper cites.
Kolmogorov’s theorem is relevant
Kůrková, V · 1991
Earlier work this paper cites.
Kolmogorov’s theorem and multilayer neural networks
Kůrková, V · 1992
Earlier work this paper cites.
Polynomial bounds for VC dimension of sigmoidal and general Pfaffian neural networks
Karpinski, M. and Macintyre, A · 1997
Earlier work this paper cites.
Lower bounds for approximation by mlp neural networks
Maiorov, V. and Pinkus, A · 1999
Earlier work this paper cites.
Betti numbers of semi-Pfaffian sets
Zell, T · 1999
Cited alongside, same era.
Kolmogorov’s spline network
Igelnik, B. and Parikh, N · 2003
Cited alongside, same era.
Complexity of computations with Pfaffian and Noetherian functions
Gabrielov, A. and Vorobjov, N · 2004
Cited alongside, same era.
Deep network approximation with discrepancy being reciprocal of width to power of depth
Shen, Z., Yang, H., and Zhang, S · 2006
Cited alongside, same era.
Neural network approximation: Three hidden layers are enough
Shen, Z., Yang, H., and Zhang, S · 2010
Cited alongside, same era.
Deep sparse rectifier neural networks
Glorot, X., Bordes, A., and Bengio, Y · 2011
On the approximation by neural networks with bounded number of neurons in hidden layers
Ismailov, V. E · 2014
Later among the works it cites.
Fast and accurate deep network learning by exponential linear units (elus)
Clevert, D.-A., Unterthiner, T., and Hochreiter, S · 2015
Later among the works it cites.
A single hidden layer feedforward network with only one neuron in the hidden layer can approximate any univariate function
Guliyev, N. J. and Ismailov, V. E · 2016
Later among the works it cites.
The phase diagram of approximation rates for deep neural networks
Yarotsky, D. and Zhevnerchuk, A · 2019
Later among the works it cites.
Error bounds for deep ReLU networks using the Kolmogorov–Arnold superposition theorem
Montanelli, H. and Yang, H · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
On the complexity of neural network classifiers: A comparison between shallow and deep architectures
Bianchini, M. and Scarselli, F · 2014
Cited alongside, same era.
Approximation capability of two hidden layer feedforward neural networks with fixed weights
Guliyev, N. J. and Ismailov, V. E
Cited in the paper.
On the approximation by single hidden layer feedforward neural networks with fixed weights
Guliyev, N. J. and Ismailov, V. E
Cited in the paper.
The Kolmogorov-Arnold representation theorem revisited
Schmidt-Hieber, J · 2020
Later among the works it cites.