Fetching the paper…
Reading the bibliography…
We demonstrate that a very deep ResNet with stacked modules with one neuron per hidden layer and ReLU activation functions can uniformly approximate any Lebesgue integrable function in $d$ dimensions, i.e.
Approximation by superpositions of a sigmoidal function
G. Cybenko · 1989
Earlier work this paper cites.
On the approximate realization of continuous mappings by neural networks
K. Funahashi · 1989
Earlier work this paper cites.
Multilayer feedforward networks are universal approximators
K. Hornik, M. Stinchcombe, and H. White · 1989
Earlier work this paper cites.
Kolmogorov’s theorem and multilayer neural networks
V. Kurková · 1992
Earlier work this paper cites.
Universal approximation bounds for superpositions of a sigmoidal function
A. R. Barron · 1993
Earlier work this paper cites.
Neural networks for optimal approximation of smooth and analytic functions
H. N. Mhaskar · 1996
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner · 1998
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
A. Krizhevsky, I. Sutskever, and G. Hinton · 2012
Earlier work this paper cites.
Dropout: A simple way to prevent neural networks from overfitting
N. Srivastava, G. Hinton, A. Krizhevsky, I. Sutskever, and R. Salakhutdinov · 2014
Earlier work this paper cites.
Deep networks are effective encoders of periodicity
L. Szymanski and B. McCane · 2014
Earlier work this paper cites.
The loss surfaces of multilayer networks
A. Choromanska, M. Henaff, M. Mathieu, G. B. Arous, and Y. LeCun · 2015
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
K. Simonyan and A. Zisserman · 2015
Earlier work this paper cites.
Going deeper with convolutions
C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, A. Rabinovich, J. Rick Chang, et al · 2015
Cited alongside, same era.
On the expressive power of deep learning: A tensor analysis
N. Cohen, O. Sharir, and A. Shashua · 2016
Cited alongside, same era.
The power of depth for feedforward neural networks
R. Eldan and O. Shamir · 2016
Cited alongside, same era.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Cited alongside, same era.
Identity mappings in deep residual networks
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Cited alongside, same era.
Deep learning without poor local minima
K. Kawaguchi · 2016
Cited alongside, same era.
Why deep neural networks for function approximation?
S. Liang and R. Srikant · 2017
Later among the works it cites.
The expressive power of neural networks: A view from the width
Z. Lu, H. Pu, F. Wang, Z. Hu, and L. Wang · 2017
Later among the works it cites.
The loss surface of deep and wide neural networks
Q. Nguyen and M. Hein · 2017
Later among the works it cites.
Weight sharing is crucial to succesful optimization
S. Shalev-Shwartz, O. Shamir, and S. Shammah · 2017
Later among the works it cites.
Error bounds for approximations with deep relu networks
D. Yarotsky · 2017
Later among the works it cites.
On the optimization of deep networks: Implicit acceleration by overparameterization
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
H. N. Mhaskar and T. Poggio · 2016
Cited alongside, same era.
Benefits of depth in neural networks
M. Telgarsky · 2016
Cited alongside, same era.
Understanding deep learning requires rethinking generalization
C. Zhang, S. Bengio, M. Hardt, B. Recht, and O. Vinyals · 2016
Cited alongside, same era.
Gradient descent learns one-hidden-layer cnn: Don’t be afraid of spurious local minima
S. S. Du, J. Lee, Y. Tian, B. Poczos, and A. Singh · 2017
Cited alongside, same era.
Approximating continuous functions by relu nets of minimal width
B. Hanin and M. Sellke · 2017
Cited alongside, same era.
Identity matters in deep learning
M. Hardt and T. Ma · 2017
Cited alongside, same era.
S. Arora, N. Cohen, and E. Hazan · 2018
Closest in time.
On decision regions of narrow deep neural networks
H. Beise, S. D. Da Cruz, and U. Schroder · 2018
Closest in time.
Sgd learns over-parameterized networks that provably generalize on linearly separable data
A. Brutzkus, A. Globerson, E. Malach, and S. Shalev-Shwartz · 2018
Closest in time.
The power of deeper networks for expressing natural functions
D. Rolnick and M. Tegmark · 2018
Closest in time.
Are resnets provably better than linear predictors?
O. Shamir · 2018
Closest in time.
Optimal approximation of continuous functions by very deep relu networks
D. Yarotsky · 2018
Closest in time.
Global optimality conditions for deep neural networks
C. Yun, S. Sra, and A. Jadbabaie · 2018
Closest in time.