Fetching the paper…
Reading the bibliography…
We prove a universality theorem for learning with random features.
J. W. Lindeberg, “Eine neue herleitung des exponentialgesetzes in der wahrscheinlichkeitsrechnung,” Mathematische Zeitschrift , vol. 15, no. 1, pp. 211–225, 1922
1922
Earlier work this paper cites.
C. Stein, “A bound for the error in the normal approximation to the distribution of a sum of dependent random variables,” in Proceedings of the Sixth Berkeley Symposium on Mathematical Statistics and Probability, Volume 2: Probability Theory . The Regents of the University of California, 1972
1972
Earlier work this paper cites.
Y. Gordon, “Some inequalities for gaussian processes and applications,” Israel Journal of Mathematics , vol. 50, no. 4, pp. 265–289, 1985
1985
Earlier work this paper cites.
M. Mézard, G. Parisi, and M. Virasoro, Spin glass theory and beyond: An Introduction to the Replica Method and Its Applications . World Scientific Publishing Company, 1987, vol. 9
1987
Earlier work this paper cites.
J.-B. Hiriart-Urrut and C. Lemaréchal, Fundamentals of convex analysis . Berlin: Springer-Verlag, 2001
2001
Earlier work this paper cites.
A. D. Barbour and L. H. Y. Chen, An introduction to Stein’s method . World Scientific, 2005, vol. 4
2005
Earlier work this paper cites.
A. Rahimi and B. Recht, “Random features for large-scale kernel machines,” in Advances in neural information processing systems , 2008, pp. 1177–1184
2008
Earlier work this paper cites.
M. Talagrand, Mean Field Models for Spin Glasses . Springer, 2010, vol. 1
2010
Earlier work this paper cites.
S. B. Korada and A. Montanari, “Applications of the lindeberg principle in communications and statistical learning,” IEEE transactions on information theory , vol. 57, no. 4, pp. 2440–2450, 2011
2011
Earlier work this paper cites.
L. H. Y. Chen, L. Goldstein, and Q. Shao, Normal approximation by Stein’s method . New York: Springer, 2011
2011
Earlier work this paper cites.
V. Chandrasekaran, B. Recht, P. A. Parrilo, and A. S. Willsky, “The convex geometry of linear inverse problems,” Foundations of Computational Mathematics , vol. 12, no. 6, pp. 805–849, 2012
2012
Earlier work this paper cites.
X. Cheng and A. Singer, “The spectrum of random inner-product kernel matrices,” Random Matrices: Theory and Applications , vol. 2, no. 04, p. 1350010, 2013
2013
Earlier work this paper cites.
S. Boucheron, G. Lugosi, and P. Massart, Concentration inequalities: A nonasymptotic theory of independence . Oxford university press, 2013
2013
Earlier work this paper cites.
D. Amelunxen, M. Lotz, M. B. McCoy, and J. A. Tropp, “Living on the edge: Phase transitions in convex programs with random data,” Information and Inference: A Journal of the IMA , vol. 3, no. 3, pp. 224–294, 2014
2014
Earlier work this paper cites.
A. Daniely, R. Frostig, and Y. Singer, “Toward deeper understanding of neural networks: The power of initialization and a dual view on expressivity,” in Advances In Neural Information Processing Systems , 2016, pp. 2253–2261
2016
Earlier work this paper cites.
A. Daniely, “SGD learns the conjugate kernel class of the network,” in Advances in Neural Information Processing Systems , 2017, pp. 2422–2430
2017
Earlier work this paper cites.
F. Bach, “On the equivalence between kernel quadrature rules and random feature expansions,” The Journal of Machine Learning Research , vol. 18, no. 1, pp. 714–751, 2017
2017
Earlier work this paper cites.
J. Pennington and P. Worah, “Nonlinear random matrix theory for deep learning,” in Advances in Neural Information Processing Systems , 2017, pp. 2637–2646
2017
Cited alongside, same era.
A. Montanari and P.-M. Nguyen, “Universality of the elastic net error,” in 2017 IEEE International Symposium on Information Theory (ISIT) . IEEE, 2017, pp. 2338–2342
2017
Cited alongside, same era.
A. Panahi and B. Hassibi, “A universal analysis of large-scale regularized least squares solutions,” in Advances in Neural Information Processing Systems , 2017, pp. 3381–3390
2017
Cited alongside, same era.
A. Jacot, F. Gabriel, and C. Hongler, “Neural tangent kernel: Convergence and generalization in neural networks,” in Advances in neural information processing systems , 2018, pp. 8571–8580
2018
Cited alongside, same era.
E. Abbasi, F. Salehi, and B. Hassibi, “Universality in learning from linear measurements,” in Advances in Neural Information Processing Systems , 2019, pp. 12 372–12 382
2019
Later among the works it cites.
2020
Closest in time.
2020
Closest in time.
2020
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2018
Cited alongside, same era.
C. Louart, Z. Liao, and R. Couillet, “A random matrix approach to neural networks,” The Annals of Applied Probability , vol. 28, no. 2, pp. 1190–1248, 2018
2018
Cited alongside, same era.
C. Thrampoulidis, E. Abbasi, and B. Hassibi, “Precise error analysis of regularized M M -estimators in high-dimensions,” IEEE Trans. Inf. Theory , vol. 64, no. 8, pp. 5592–5628, 2018
2018
Cited alongside, same era.
N. El Karoui, “On the impact of predictor geometry on the performance on high-dimensional ridge-regularized generalized robust regression estimators,” Probability Theory and Related Fields , vol. 170, no. 1-2, pp. 95–175, 2018
2018
Cited alongside, same era.
S. Oymak and J. A. Tropp, “Universality laws for randomized dimension reduction, with applications,” Information and Inference: A Journal of the IMA , vol. 7, no. 3, pp. 337–446, 2018
2018
Cited alongside, same era.
R. Vershynin, High-dimensional Probability: An Introduction with Applications in Data Science . Cambridge University Press, 2018
2018
Cited alongside, same era.
2019
Cited alongside, same era.
2019
Cited alongside, same era.
2020
Closest in time.
2020
Closest in time.
2020
Closest in time.
2021
Closest in time.
2021
Closest in time.
2022
Closest in time.
2022
Closest in time.
2022
Closest in time.
2022
Closest in time.
2022
Closest in time.
2022
Closest in time.