Fetching the paper…
Reading the bibliography…
The minimal number of neurons required for a feedforward neural network to interpolate $n$ generic input-output pairs from $\mathbb{R}^d\times \mathbb{R}^{d'}$ is $\Theta(\sqrt{nd'})$.
Sard A (1942) The measure of the critical values of differentiable maps. Bulletin of the American Mathematical Society 48:883–890
1942
Earlier work this paper cites.
McCulloch WS, Pitts W (1943) A logical calculus of the ideas immanent in nervous activity. The bulletin of mathematical biophysics 5(4):115–133
1943
Earlier work this paper cites.
Rosenblatt F (1958) The perceptron: A probabilistic model for information storage and organization in the brain. Psychological Review 65(6):386–408
1958
Earlier work this paper cites.
Gantmacher FR (1960) The Theory of Matrices, Volume 1. Chelsea Publishing Company, New York
1960
Earlier work this paper cites.
Cover TM (1965) Geometrical and statistical properties of systems of linear inequalities with applications in pattern recognition. IEEE transactions on electronic computers 3:326–334
1965
Earlier work this paper cites.
Gunning RC, Rossi H (1965) Analytic functions of several complex variables. Prentice-Hall, Inc., Englewood Cliffs, N.J
1965
Earlier work this paper cites.
Rockafellar RT (1970) Convex Analysis. Princeton University Press, Princeton
1970
Earlier work this paper cites.
Guaraldo F, Macrì P, Tancredi A (1986) Topics on Real Analytic Spaces. Vieweg+Teubner Verlag
1986
Earlier work this paper cites.
Baum EB (1988) On the capabilities of multilayer perceptrons. Journal of Complexity 4(3):193–215
1988
Earlier work this paper cites.
Huang SC, Huang YF (1991) Bounds on the number of hidden neurons in multilayer perceptrons. IEEE Transactions on Neural Networks 2(1):47–55
1991
Cited alongside, same era.
Sakurai A (1992) n-h-1 networks store no less n*h+1 examples, but sometimes no more. In: International Joint Conference on Neural Networks (IJCNN), vol 3, pp 936–941
1992
Cited alongside, same era.
Yamasaki M (1993) The lower bound of the capacity for a neural network with multiple hidden layers. In: International Conference on Artificial Neural Networks (ICANN), pp 546–549
1993
Cited alongside, same era.
Huang GB (2003) Learning capability and storage capacity of two-hidden-layer feedforward networks. IEEE Transactions on Neural Networks 14(2):274–281
2003
Cited alongside, same era.
Heubach S, Mansour T (2004) Compositions of n with parts in a set. Congressus Numerantium 168:127
2004
Yun C, Sra S, Jadbabaie A (2019) Small relu networks are powerful memorizers: a tight analysis of memorization capacity. In: Neural Information Processing Systems (NeurIPS), vol 32
2019
Later among the works it cites.
Bubeck S, Eldan R, Lee YT, Mikulincer D (2020) Network size and size of the weights in memorization with two-layers neural networks. In: Neural Information Processing Systems (NeurIPS), vol 33, pp 4977–4986
2020
Later among the works it cites.
Vershynin R (2020) Memory capacity of neural networks with threshold and rectified linear unit activations. SIAM Journal on Mathematics of Data Science 2(4):1004–1033
2020
Later among the works it cites.
Nguyen Q, Mondelli M, Montufar GF (2021) Tight bounds on the smallest eigenvalue of the neural tangent kernel for deep relu networks. In: International Conference on Machine Learning (), vol 139, pp 8119–8129
2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Allman ES, Matias C, Rhodes JA (2009) Identifiability of parameters in latent structure models with many observed variables. The Annals of Statistics 37(6A):3099 – 3132
2009
Cited alongside, same era.
Lee JM (2013) Introduction to Smooth Manifolds. Springer
2013
Cited alongside, same era.
Axler S (2015) Linear Algebra Done Right, Third Edition. Springer, New York
2015
Cited alongside, same era.
Vaswani A, Shazeer N, Parmar N, Uszkoreit J, Jones L, Gomez AN, Kaiser Lu, Polosukhin I (2017) Attention is all you need. In: Neural Information Processing Systems (NeurIPS), vol 30
2017
Cited alongside, same era.
Park S, Lee J, Yun C, Shin J (2021) Provable memorization via deep neural networks using sub-linear parameters. In: Conference on Learning Theory (COLT), vol 134, pp 3627–3661
2021
Later among the works it cites.
Rajput S, Sreenivasan K, Papailiopoulos D, Karbasi A (2021) An exponential improvement on the memorization capacity of deep threshold networks. In: Neural Information Processing Systems (NeurIPS), vol 34, pp 12,674–12,685
2021
Later among the works it cites.
Bombari S, Amani MH, Mondelli M (2022) Memorization and optimization in deep neural networks with minimum over-parameterization. In: Neural Information Processing Systems (NeurIPS), vol 35, pp 7628–7640
2022
Later among the works it cites.
Madden L, Thrampoulidis C (2024) Memory capacity of two layer neural networks with smooth activations. SIAM Journal on Mathematics of Data Science 6(3):679–702
2024
Closest in time.