Fetching the paper…
Reading the bibliography…
It is known that $O(N)$ parameters are sufficient for neural networks to memorize arbitrary $N$ input-label pairs.
Some elementary inequalities relating to the gamma and incomplete gamma function
Walter Gautschi · 1959
Earlier work this paper cites.
On the capabilities of multilayer perceptrons
Eric B. Baum · 1988
Earlier work this paper cites.
What size net gives valid generalization?
Eric B. Baum and David Haussler · 1989
Earlier work this paper cites.
Approximation by superpositions of a sigmoidal function
George Cybenko · 1989
Earlier work this paper cites.
The lower bound of the capacity for a neural network with multiple hidden layers
Masami Yamasaki · 1993
Earlier work this paper cites.
Polynomial bounds for VC dimension of sigmoidal and general Pfaffian neural networks
Marek Karpinski and Angus Macintyre · 1997
Earlier work this paper cites.
Bounds for the computational power and learning complexity of analog neural nets
Wolfgang Maass · 1997
Earlier work this paper cites.
Shattering all sets of k points in “general position” requires (k—1)/2 parameters
Eduardo D. Sontag · 1997
Earlier work this paper cites.
Upper bounds on the number of hidden neurons in feedforward networks with arbitrary bounded nonlinear activation functions
Guang-Bin Huang and Haroon A Babri · 1998
Earlier work this paper cites.
Approximation theory of the MLP model in neural networks
Allan Pinkus · 1999
Earlier work this paper cites.
Learning capability and storage capacity of two-hidden-layer feedforward networks
Guang-Bin Huang · 2003
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky · 2009
Cited alongside, same era.
Reading digits in natural images with unsupervised feature learning
Yuval Netzer, T. Wang, Adam Coates, Alessandro Bissacco, Bo Wu, and Andrew Y. Ng · 2011
Cited alongside, same era.
Understanding machine learning: From theory to algorithms
Shai Shalev-Shwartz and Shai Ben-David · 2014
Cited alongside, same era.
The power of depth for feedforward neural networks
Ronen Eldan and Ohad Shamir · 2016
Cited alongside, same era.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Cited alongside, same era.
Benefits of depth in neural networks
Matus Telgarsky · 2016
Cited alongside, same era.
Optimization landscape and expressivity of deep CNNs
Quynh Nguyen and Matthias Hein · 2018
Later among the works it cites.
Optimal approximation of continuous functions by very deep relu networks
Dmitry Yarotsky · 2018
Later among the works it cites.
Nearly-tight VC-dimension and pseudodimension bounds for piecewise linear neural networks
Peter L. Bartlett, Nick Harvey, Christopher Liaw, and Abbas Mehrabian · 2019
Later among the works it cites.
Reconciling modern machine-learning practice and the classical bias–variance trade-off
Mikhail Belkin, Daniel Hsu, Siyuan Ma, and Soumik Mandal · 2019
Later among the works it cites.
Deep, skinny neural networks are not universal approximators
Jesse Johnson · 2019
Later among the works it cites.
Small ReLU networks are powerful memorizers: a tight analysis of memorization capacity
Chulhee Yun, Suvrit Sra, and Ali Jadbabaie · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Boris Hanin and Mark Sellke · 2017
Cited alongside, same era.
Identity matters in deep learning
Moritz Hardt and Tengyu Ma · 2017
Cited alongside, same era.
The expressive power of neural networks: A view from the width
Zhou Lu, Hongming Pu, Feicheng Wang, Zhiqiang Hu, and Liwei Wang · 2017
Cited alongside, same era.
Understanding deep learning requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2017
Cited alongside, same era.
Understanding deep neural networks with rectified linear units
Raman Arora, Amitabh Basu, Poorya Mianjy, and Anirbit Mukherjee · 2018
Cited alongside, same era.
Better depth-width trade-offs for neural networks through the lens of dynamical systems
Vaggos Chatziafratis, Sai Ganesh Nagarajan, and Ioannis Panageas
Cited in the paper.
Later among the works it cites.
Universal approximation with deep narrow networks
Patrick Kidger and Terry Lyons · 2020
Closest in time.
Deep double descent: Where bigger models and more data hurt
Preetum Nakkiran, Gal Kaplun, Yamini Bansal, Tristan Yang, Boaz Barak, and Ilya Sutskever · 2020
Closest in time.
Memory capacity of neural networks with threshold and ReLU activations
Roman Vershynin · 2020
Closest in time.
Are wider nets better given the same number of parameters?
Anna Golubeva, Guy Gur-Ari, and Behnam Neyshabur · 2021
Closest in time.
Minimum width for universal approximation
Sejun Park, Chulhee Yun, Jaeho Lee, and Jinwoo Shin · 2021
Closest in time.