Fetching the paper…
Reading the bibliography…
Deep learning has been wildly successful in practice and most state-of-the-art machine learning methods are based on neural networks.
J. Radon, “Über die bestimmung von funktionen durch ihre integralwerte längs gewisser mannigfaltigkeiten,” Ber. Verh, Sachs Akad Wiss. , vol. 69, pp. 262–277, 1917
1917
Earlier work this paper cites.
W. S. McCulloch and W. Pitts, “A logical calculus of the ideas immanent in nervous activity,” The Bulletin of Mathematical Biophysics , vol. 5, no. 4, pp. 115–133, 1943
1943
Earlier work this paper cites.
S. D. Fisher and J. W. Jerome, “Spline solutions to L 1 L^{1} extremal problems in one and several variables,” Journal of Approximation Theory , vol. 13, no. 1, pp. 73–83, 1975
1975
Earlier work this paper cites.
G. Pisier, “Remarques sur un résultat non publié de B. Maurey,” Séminaire d’Analyse Fonctionnelle (dit ”Maurey-Schwartz”) , pp. 1–12, April 1981
1981
Earlier work this paper cites.
Y. LeCun, J. Denker, and S. Solla, “Optimal brain damage,” Advances in neural information processing systems , vol. 2, 1989
1989
Earlier work this paper cites.
L. K. Jones, “A simple lemma on greedy approximation in Hilbert space and convergence rates for projection pursuit regression and neural network training,” The Annals of Statistics , pp. 608–613, 1992
1992
Earlier work this paper cites.
L. I. Rudin, S. Osher, and E. Fatemi, “Nonlinear total variation based noise removal algorithms,” Physica D: nonlinear phenomena , vol. 60, no. 1-4, pp. 259–268, 1992
1992
Earlier work this paper cites.
A. R. Barron, “Universal approximation bounds for superpositions of a sigmoidal function,” IEEE Transactions on Information theory , vol. 39, no. 3, pp. 930–945, 1993
1993
Earlier work this paper cites.
R. Tibshirani, “Regression shrinkage and selection via the lasso,” Journal of the Royal Statistical Society: Series B (Methodological) , vol. 58, no. 1, pp. 267–288, 1996
1996
Earlier work this paper cites.
V. Kůrková, P. C. Kainen, and V. Kreinovich, “Estimates of the number of hidden units and variation with respect to half-spaces,” Neural Networks , vol. 10, no. 6, pp. 1061–1068, 1997
1997
Earlier work this paper cites.
E. Mammen and S. van de Geer, “Locally adaptive regression splines,” The Annals of Statistics , vol. 25, no. 1, pp. 387–413, 1997
1997
Earlier work this paper cites.
A. V. Oppenheim, A. S. Willsky, and S. H. Nawab, Signals & Systems , ser. Prentice-Hall Signal Processing Series. Prentice Hall, 1997
1997
Earlier work this paper cites.
E. J. Candès, “Ridgelets: Theory and applications,” Ph.D. dissertation, Stanford University, 1998
1998
Earlier work this paper cites.
D. L. Donoho and I. M. Johnstone, “Minimax estimation via wavelet shrinkage,” The Annals of Statistics , vol. 26, no. 3, pp. 879–921, 1998
1998
Earlier work this paper cites.
Y. Grandvalet, “Least absolute shrinkage is equivalent to quadratic penalization,” in International Conference on Artificial Neural Networks . Springer, 1998, pp. 201–206
1998
Earlier work this paper cites.
D. L. Donoho, “High-dimensional data analysis: The curses and blessings of dimensionality,” AMS Lectures , p. 32, 2000
2000
Earlier work this paper cites.
V. Kůrková and M. Sanguineti, “Bounds on rates of variable-basis and neural-network approximation,” IEEE Transactions on Information Theory , vol. 47, no. 6, pp. 2659–2665, 2001
2001
Earlier work this paper cites.
M. Vetterli, P. Marziliano, and T. Blu, “Sampling signals with finite rate of innovation,” IEEE Transactions on Signal Processing , vol. 50, no. 6, pp. 1417–1428, 2002
2002
Cited alongside, same era.
E. J. Candès, J. Romberg, and T. Tao, “Robust uncertainty principles: Exact signal reconstruction from highly incomplete frequency information,” IEEE Transactions on Information Theory , vol. 52, no. 2, pp. 489–509, 2006
2006
Cited alongside, same era.
D. L. Donoho, “Compressed sensing,” IEEE Transactions on Information Theory , vol. 52, no. 4, pp. 1289–1306, 2006
2006
Cited alongside, same era.
S. Mallat, A Wavelet Tour of Signal Processing , 3rd ed. Elsevier/Academic Press, Amsterdam, 2009
2009
Cited alongside, same era.
L. J. Ba and R. Caruana, “Do deep nets really need to be deep?” in Proceedings of the 27th International Conference on Neural Information Processing Systems-Volume 2 , 2014, pp. 2654–2662
T. Poggio, A. Banburski, and Q. Liao, “Theoretical issues in deep networks,” Proceedings of the National Academy of Sciences , vol. 117, no. 48, pp. 30 039–30 045, 2020
2020
Later among the works it cites.
J. Schmidt-Hieber, “Nonparametric regression using deep neural networks with ReLU activation function,” The Annals of Statistics , vol. 48, no. 4, pp. 1875–1897, 2020
2020
Later among the works it cites.
A. Golubeva, B. Neyshabur, and G. Gur-Ari, “Are wider nets better given the same number of parameters?” International Conference on Learning Representations , 2021
2021
Later among the works it cites.
R. Parhi and R. D. Nowak, “Banach space representer theorems for neural networks and ridge splines.” Journal of Machine Learning Research , vol. 22, no. 43, pp. 1–40, 2021
2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2014
Cited alongside, same era.
B. Neyshabur, R. Tomioka, and N. Srebro, “In search of the real inductive bias: On the role of implicit regularization in deep learning.” in International Conference on Learning Representations (Workshop) , 2015
2015
Cited alongside, same era.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2016, pp. 770–778
2016
Cited alongside, same era.
F. Bach, “Breaking the curse of dimensionality with convex neural networks,” The Journal of Machine Learning Research , vol. 18, no. 1, pp. 629–681, 2017
2017
Cited alongside, same era.
M. Unser, J. Fageot, and J. P. Ward, “Splines are universal solutions of linear inverse problems with generalized TV regularization,” SIAM Review , vol. 59, no. 4, pp. 769–793, 2017
2017
Cited alongside, same era.
A. Sanyal, P. H. Torr, and P. K. Dokania, “Stable rank normalization for improved generalization in neural networks and GANs,” International Conference on Learning Representations , 2019
2019
Cited alongside, same era.
P. Savarese, I. Evron, D. Soudry, and N. Srebro, “How do infinite width bounded norm networks look in function space?” in Conference on Learning Theory . PMLR, 2019, pp. 2667–2690
2019
Cited alongside, same era.
R. Balestriero and R. G. Baraniuk, “Mad max: Affine spline insights into deep learning,” Proceedings of the IEEE , vol. 109, no. 5, pp. 704–727, 2020
2020
Cited alongside, same era.
2021
Later among the works it cites.
T. Debarre, Q. Denoyelle, M. Unser, and J. Fageot, “Sparsest piecewise-linear regression of one-dimensional data,” Journal of Computational and Applied Mathematics , vol. 406, p. 114044, 2022
2022
Later among the works it cites.
R. Parhi, “On Ridge Splines, Neural Networks, and Variational Problems in Radon-Domain BV Spaces,” Ph.D. dissertation, The University of Wisconsin–Madison, 2022
2022
Later among the works it cites.
R. Parhi and R. D. Nowak, “What kinds of functions do deep neural networks learn? Insights from variational spline theory,” SIAM Journal on Mathematics of Data Science , vol. 4, no. 2, pp. 464–489, 2022
2022
Later among the works it cites.
J. W. Siegel and J. Xu, “Sharp bounds on the approximation rates, metric entropy, and n n -widths of shallow neural networks,” Foundations of Computational Mathematics , pp. 1–57, 2022
2022
Later among the works it cites.
2022
Later among the works it cites.
L. Yang, J. Zhang, J. Shenouda, D. Papailiopoulos, K. Lee, and R. D. Nowak, “A better way to decay: Proximal gradient training algorithms for neural nets,” in OPT 2022: Optimization for Machine Learning (NeurIPS Workshop) , 2022
2022
Later among the works it cites.
M. Yuan and Y. Lin, “Model selection and estimation in regression with grouped variables,” Journal of the Royal Statistical Society: Series B (Statistical Methodology) , vol. 68, no. 1, pp. 49–67, 2006
2022
Later among the works it cites.
F. Bartolucci, E. De Vito, L. Rosasco, and S. Vigogna, “Understanding neural networks with reproducing kernel Banach spaces,” Appl. Comput. Harmon. Anal. , vol. 62, pp. 194–236, 2023
2023
Closest in time.
R. Parhi and R. D. Nowak, “Near-minimax optimal estimation with shallow ReLU neural networks,” IEEE Transactions on Information Theory , vol. 69, no. 2, pp. 1125–1140, 2023
2023
Closest in time.
2023
Closest in time.
M. Unser, “Ridges, neural networks, and the Radon transform,” Journal of Machine Learning Research , vol. 24, no. 37, pp. 1–33, 2023
2023
Closest in time.