Fetching the paper…
Reading the bibliography…
We develop a corrective mechanism for neural network approximation: the total available non-linear units are divided into multiple groups and the first group approximates the function under consideration, the second group approximates the error in approximation produced by the first group and corrects it, the third group approximates the error produced by the first and second groups together and so on.
Approximation by superpositions of a sigmoidal function
George Cybenko · 1989
Earlier work this paper cites.
On the approximate realization of continuous mappings by neural networks
Ken-Ichi Funahashi · 1989
Earlier work this paper cites.
Multilayer feedforward networks are universal approximators
Kurt Hornik, Maxwell Stinchcombe, and Halbert White · 1989
Earlier work this paper cites.
Universal approximation bounds for superpositions of a sigmoidal function
Andrew R Barron · 1993
Earlier work this paper cites.
Uniform approximation of functions with random bases
Ali Rahimi and Benjamin Recht · 2008
Earlier work this paper cites.
Weighted sums of random kitchen sinks: Replacing minimization with randomization in learning
Ali Rahimi and Benjamin Recht · 2009
Earlier work this paper cites.
Shallow vs. deep sum-product networks
Olivier Delalleau and Yoshua Bengio · 2011
Earlier work this paper cites.
Learning polynomials with neural networks
Alexandr Andoni, Rina Panigrahy, Gregory Valiant, and Li Zhang · 2014
Earlier work this paper cites.
Convex optimization: Algorithms and complexity
Sébastien Bubeck et al · 2015
Earlier work this paper cites.
Improving deep neural networks using softplus units
Hao Zheng, Zhanlei Yang, Wenju Liu, Jizhong Liang, and Yanpeng Li · 2015
Earlier work this paper cites.
Deep learning
Ian Goodfellow, Yoshua Bengio, and Aaron Courville · 2016
Earlier work this paper cites.
Why deep neural networks for function approximation?
Shiyu Liang and Rayadurgam Srikant · 2016
Earlier work this paper cites.
benefits of depth in neural networks
Matus Telgarsky · 2016
Earlier work this paper cites.
Understanding deep learning requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2016
Earlier work this paper cites.
Approximating continuous functions by relu nets of minimal width
Boris Hanin and Mark Sellke · 2017
Cited alongside, same era.
The expressive power of neural networks: A view from the width
Zhou Lu, Hongming Pu, Feicheng Wang, Zhiqiang Hu, and Liwei Wang · 2017
Cited alongside, same era.
Depth-width tradeoffs in approximating natural functions with neural networks
Itay Safran and Ohad Shamir · 2017
Cited alongside, same era.
Error bounds for approximations with deep relu networks
Dmitry Yarotsky · 2017
Cited alongside, same era.
A convergence theory for deep learning via over-parameterization
Zeyuan Allen-Zhu, Yuanzhi Li, and Zhao Song · 2018
Cited alongside, same era.
Gradient descent finds global minima of deep neural networks
Simon Du, Jason Lee, Haochuan Li, Liwei Wang, and Xiyu Zhai · 2019
Later among the works it cites.
Ziwei Ji and Matus Telgarsky · 2019
Later among the works it cites.
Neural tangent kernels, transportation mappings, and universal approximation
Ziwei Ji, Matus Telgarsky, and Ruicheng Xian · 2019
Later among the works it cites.
Gradient descent finds global minima for generalizable deep neural networks of practical sizes
Kenji Kawaguchi and Jiaoyang Huang · 2019
Later among the works it cites.
Bo Li, Shanshan Tang, and Haijun Yu · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Mikhail Belkin, Siyuan Ma, and Soumik Mandal · 2018
Cited alongside, same era.
Sigmoid-weighted linear units for neural network function approximation in reinforcement learning
Stefan Elfwing, Eiji Uchibe, and Kenji Doya · 2018
Cited alongside, same era.
Neural tangent kernel: Convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel, and Clément Hongler · 2018
Cited alongside, same era.
Approximation by combinations of relu and squared relu ridge functions with l 1 l^{1} and l 0 l^{0} controls
Jason M Klusowski and Andrew R Barron · 2018
Cited alongside, same era.
On the approximation properties of random relu features
Yitong Sun, Anna Gilbert, and Ambuj Tewari · 2018
Cited alongside, same era.
Stochastic gradient descent optimizes over-parameterized deep relu networks
Difan Zou, Yuan Cao, Dongruo Zhou, and Quanquan Gu · 2018
Cited alongside, same era.
Approximation power of random neural networks
Bolton Bailey, Ziwei Ji, Matus Telgarsky, and Ruicheng Xian · 2019
Cited alongside, same era.
Later among the works it cites.
Barron spaces and the compositional function spaces for neural network models
Chao Ma, Lei Wu, et al · 2019
Later among the works it cites.
Samet Oymak and Mahdi Soltanolkotabi · 2019
Later among the works it cites.
Effect of activation functions on the training of overparametrized neural nets
Abhishek Panigrahi, Abhishek Shetty, and Navin Goyal · 2019
Later among the works it cites.
Quadratic suffices for over-parametrization via matrix chernoff bound
Zhao Song and Xin Yang · 2019
Later among the works it cites.
On the power and limitations of random features for understanding neural networks
Gilad Yehudai and Ohad Shamir · 2019
Later among the works it cites.
An improved analysis of training over-parameterized deep neural networks
Difan Zou and Quanquan Gu · 2019
Later among the works it cites.
Sharp representation theorems for relu networks with precise dependence on depth
Guy Bresler and Dheeraj Nagaraj · 2020
Closest in time.
Network size and weights size for memorization with two-layers neural networks
Sébastien Bubeck, Ronen Eldan, Yin Tat Lee, and Dan Mikulincer · 2020
Closest in time.