Fetching the paper…
Reading the bibliography…
We study the memorization power of feedforward ReLU neural networks.
On the capabilities of multilayer perceptrons
E. B. Baum · 1988
Earlier work this paper cites.
Approximation by superpositions of a sigmoidal function
G. Cybenko · 1989
Earlier work this paper cites.
Bounds on the number of hidden neurons in multilayer perceptrons
S.-C. Huang, Y.-F. Huang, et al · 1991
Earlier work this paper cites.
A simple method to derive bounds on the size and to train multilayer neural networks
M. A. Sartori and P. J. Antsaklis · 1991
Earlier work this paper cites.
Multilayer feedforward networks with a nonpolynomial activation function can approximate any function
M. Leshno, V. Y. Lin, A. Pinkus, and S. Schocken · 1993
Earlier work this paper cites.
Bounding the vapnik-chervonenkis dimension of concept classes parameterized by real numbers
P. W. Goldberg and M. R. Jerrum · 1995
Earlier work this paper cites.
Shattering all sets of ‘k’points in “general position” requires (k—1)/2 parameters
E. D. Sontag · 1997
Earlier work this paper cites.
Almost linear vc-dimension bounds for piecewise polynomial networks
P. L. Bartlett, V. Maiorov, and R. Meir · 1998
Earlier work this paper cites.
Upper bounds on the number of hidden neurons in feedforward networks with arbitrary bounded nonlinear activation functions
G.-B. Huang and H. A. Babri · 1998
Earlier work this paper cites.
Learning capability and storage capacity of two-hidden-layer feedforward networks
G.-B. Huang · 2003
Earlier work this paper cites.
Neural network learning: Theoretical foundations
M. Anthony and P. L. Bartlett · 2009
Earlier work this paper cites.
On the representational efficiency of restricted boltzmann machines
J. Martens, A. Chattopadhya, T. Pitassi, and R. Zemel · 2013
Earlier work this paper cites.
Understanding machine learning: From theory to algorithms
S. Shalev-Shwartz and S. Ben-David · 2014
Earlier work this paper cites.
The power of depth for feedforward neural networks
R. Eldan and O. Shamir · 2016
Earlier work this paper cites.
Identity matters in deep learning
M. Hardt and T. Ma · 2016
Cited alongside, same era.
Why deep neural networks for function approximation?
S. Liang and R. Srikant · 2016
Cited alongside, same era.
Benefits of depth in neural networks
M. Telgarsky · 2016
Cited alongside, same era.
Depth separation for neural networks
A. Daniely · 2017
Cited alongside, same era.
Depth-width tradeoffs in approximating natural functions with neural networks
I. Safran and O. Shamir · 2017
Cited alongside, same era.
Error bounds for approximations with deep relu networks
D. Yarotsky · 2017
Cited alongside, same era.
Small relu networks are powerful memorizers: a tight analysis of memorization capacity
C. Yun, S. Sra, and A. Jadbabaie · 2019
Later among the works it cites.
Sharp representation theorems for relu networks with precise dependence on depth
G. Bresler and D. Nagaraj · 2020
Later among the works it cites.
Network size and weights size for memorization with two-layers neural networks
S. Bubeck, R. Eldan, Y. T. Lee, and D. Mikulincer · 2020
Later among the works it cites.
Memorizing gaussians with no over-parameterizaion via gradient decent on neural networks
A. Daniely · 2020
Later among the works it cites.
Provable memorization via deep neural networks using sub-linear parameters
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Optimization landscape and expressivity of deep cnns
Q. Nguyen and M. Hein · 2018
Cited alongside, same era.
Nearly-tight vc-dimension and pseudodimension bounds for piecewise linear neural networks
P. L. Bartlett, N. Harvey, C. Liaw, and A. Mehrabian · 2019
Cited alongside, same era.
Reconciling modern machine-learning practice and the classical bias–variance trade-off
M. Belkin, D. Hsu, S. Ma, and S. Mandal · 2019
Cited alongside, same era.
Depth-width trade-offs for relu networks via sharkovsky’s theorem
V. Chatziafratis, S. G. Nagarajan, I. Panageas, and X. Wang · 2019
Cited alongside, same era.
Neural networks learning and memorization with (almost) no over-parameterization
A. Daniely · 2019
Cited alongside, same era.
Deep double descent: Where bigger models and more data hurt
P. Nakkiran, G. Kaplun, Y. Bansal, T. Yang, B. Barak, and I. Sutskever · 2019
Cited alongside, same era.
S. Park, J. Lee, C. Yun, and J. Shin · 2020
Later among the works it cites.
Neural networks with small weights and depth-separation barriers
G. Vardi and O. Shamir · 2020
Later among the works it cites.
Memory capacity of neural networks with threshold and relu activations
R. Vershynin · 2020
Later among the works it cites.
A universal law of robustness via isoperimetry
S. Bubeck and M. Sellke · 2021
Closest in time.
A law of robustness for two-layers neural networks
S. Bubeck, Y. Li, and D. M. Nagaraj · 2021
Closest in time.
An exponential improvement on the memorization capacity of deep threshold networks
S. Rajput, K. Sreenivasan, D. Papailiopoulos, and A. Karbasi · 2021
Closest in time.
Size and depth separation in approximating benign functions with neural networks
G. Vardi, D. Reichman, T. Pitassi, and O. Shamir · 2021
Closest in time.
Depth separation beyond radial functions
L. Venturi, S. Jelassi, T. Ozuch, and J. Bruna · 2021
Closest in time.
Understanding deep learning (still) requires rethinking generalization
C. Zhang, S. Bengio, M. Hardt, B. Recht, and O. Vinyals · 2021
Closest in time.