Fetching the paper…
Reading the bibliography…
For any positive integer $k$, there exist neural networks with $\Theta(k^3)$ layers, $\Theta(1)$ nodes per layer, and $\Theta(1)$ distinct parameters which can not be approximated by networks with $\mathcal{O}(k)$ layers unless they are exponentially large --- they must possess $\Omega(2^k)$ nodes.
Über die beste annäherung von funktionen einer gegebenen funktionenklasse
Andrei Kolmogorov · 1936
Earlier work this paper cites.
On multidimensional variations
Anatoli Vitushkin · 1955
Earlier work this paper cites.
On the representation of continuous functions of several variables by superpositions of continuous functions of one variable and addition
Andrey Nikolaevich Kolmogorov · 1957
Earlier work this paper cites.
Estimation of the complexity of the tabulation problem
Anatoli Vitushkin · 1959
Earlier work this paper cites.
Lower bounds for approximation by nonlinear manifolds
Hugh E. Warren · 1968
Earlier work this paper cites.
Neocognitron: A self-organizing neural network model for a mechanism of pattern recognition unaffected by shift in position
Kunihiko Fukushima · 1980
Earlier work this paper cites.
Computational Limitations of Small Depth Circuits
Johan Håstad · 1986
Earlier work this paper cites.
Real Algebraic Geometry
Jacek Bochnak, Michal Coste, and Marie-Françoise Roy · 1998
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Yann LeCun, Léon Bottou, Yoshua Bengio, and Patrick Haffner · 1998
Cited alongside, same era.
Neural Network Learning: Theoretical Foundations
Martin Anthony and Peter L. Bartlett · 1999
Cited alongside, same era.
Lower bounds on the complexity of approximating continuous functions by sigmoidal neural networks
Michael Schmitt · 2000
Cited alongside, same era.
An empirical comparison of supervised learning algorithms
Rich Caruana and Alexandru Niculescu-Mizil · 2006
Cited alongside, same era.
Shallow vs. deep sum-product networks
Yoshua Bengio and Olivier Delalleau · 2011
Cited alongside, same era.
Sum-product networks: A new deep architecture
Hoifung Poon and Pedro M. Domingos · 2011
Cited alongside, same era.
Deep networks are effective encoders of periodicity
Lech Szymanski and Brendan McCane · 2014
Later among the works it cites.
The power of depth for feedforward neural networks
Ronen Eldan and Ohad Shamir · 2015
Later among the works it cites.
Daniel Kane and Ryan Williams · 2015
Later among the works it cites.
On the expressive efficiency of sum product networks
James Martens and Venkatesh Medabalimi · 2015
Later among the works it cites.
An average-case depth hierarchy theorem for boolean circuits
Benjamin Rossman, Rocco A. Servedio, and Li-Yang Tan · 2015
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffery Hinton · 2012
Cited alongside, same era.
On the number of linear regions of deep neural networks
Guido Montúfar, Razvan Pascanu, Kyunghyun Cho, and Yoshua Bengio · 2014
Cited alongside, same era.
Über die analytische darstellbarkeit sogenannter willkürlicher functionen einer reellen veränderlichen
Karl Weierstrass
Cited in the paper.
Matus Telgarsky · 2015
Later among the works it cites.
Statistics 311/electrical engineering 377: Information theory and statistics
John Duchi · 2016
Closest in time.