Fetching the paper…
Reading the bibliography…
We study the problem of learning an unknown function using random feature models.
M. Sion, “On general minimax theorems.” Pacific J. Math. , vol. 8, no. 1, pp. 171–176, 1958
1958
Earlier work this paper cites.
T. Rockafellar, Convex Analysis . Princeton University Press, 1970
1970
Earlier work this paper cites.
P. K. Andersen and R. D. Gill, “Cox’s regression model for counting processes: A large sample study,” Ann. Statist. , vol. 10, no. 4, pp. 1100–1120, 12 1982
1982
Earlier work this paper cites.
M. Mezard, G. Parisi, and M. Virasoro, Spin Glass Theory and Beyond: An Introduction to the Replica Method and Its Applications , ser. World Scientific Lecture Notes in Physics. World Scientific, Nov. 1986, vol. 9
1986
Earlier work this paper cites.
W. K. Newey and D. McFadden, “Chapter 36 large sample estimation and hypothesis testing,” ser. Handbook of Econometrics. Elsevier, 1994, vol. 4, pp. 2111 – 2245
1994
Earlier work this paper cites.
R. K. Sundaram, A First Course in Optimization Theory . Cambridge University Press, 1996
1996
Earlier work this paper cites.
R. T. Rockafellar and R. J.-B. Wets, Variational Analysis . Springer-Verlag Berlin Heidelberg, 1998
1998
Earlier work this paper cites.
D. P. Bertsekas, A. Nedic, and A. E. Ozdaglar, Convex Analysis and Optimization . Athena Scientific, 2003
2003
Earlier work this paper cites.
M. Debbah, W. Hachem, P. Loubaton, and M. de Courville, “Mmse analysis of certain large isometric random precoded systems,” IEEE Transactions on Information Theory , vol. 49, no. 5, pp. 1293–1311, 2003
2003
Earlier work this paper cites.
S. Boyd and L. Vandenberghe, Convex Optimization . Cambridge University Press, 2004
2004
Earlier work this paper cites.
S. Boyd and L. Vandenberghe, Convex Optimization . Cambridge University Press, 2004
2004
Earlier work this paper cites.
R. L. Schilling, Measures, Integrals and Martingales . Cambridge University Press, 2005
2005
Cited alongside, same era.
A. Rahimi and B. Recht, “Random features for large-scale kernel machines,” in Advances in Neural Information Processing Systems 20 , 2008, pp. 1177–1184
2008
Cited alongside, same era.
A. Shapiro, D. Dentcheva, and A. Ruszczyński, Lectures on Stochastic Programming . Society for Industrial and Applied Mathematics, 2009
2009
Cited alongside, same era.
M. Rudelson and R. Vershynin, “Non-asymptotic theory of random matrices: extreme singular values,” 2010
2010
Cited alongside, same era.
R. Couillet and M. Debbah, Random Matrix Methods for Wireless Communications . Cambridge University Press, 2011
2011
Cited alongside, same era.
J. Pennington and P. Worah, “Nonlinear random matrix theory for deep learning,” in Advances in Neural Information Processing Systems 30 , I. Guyon, U. V. Luxburg, S. Bengio, H. Wallach, R. Fergus, S. Vishwanathan, and R. Garnett, Eds. Curran Associates, Inc., 2017, pp. 2637–2646
2017
Later among the works it cites.
S. Adachi, S. Iwata, Y. Nakatsukasa, and A. Takeda, “Solving the trust-region subproblem by a generalized eigenvalue problem,” SIAM Journal on Optimization , vol. 27, no. 1, pp. 269–291, 2017
2017
Later among the works it cites.
M. Belkin, S. Ma, and S. Mandal, “To understand deep learning we need to understand kernel learning,” in Proceedings of the 35th International Conference on Machine Learning , vol. 80, 10–15 Jul 2018, pp. 541–549
2018
Later among the works it cites.
S. Mei and A. Montanari, “The generalization error of random features regression: Precise asymptotics and double descent curve,” 2019
2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
X. Cheng and A. Singer, “The spectrum of random inner-product kernel matrices,” 2012
2012
Cited alongside, same era.
A. Klenke, Probability Theory: A Comprehensive Course . Springer-Verlag, 2014
2014
Cited alongside, same era.
2015
Cited alongside, same era.
C. Thrampoulidis, S. Oymak, and B. Hassibi, “Regularized linear regression: A precise analysis of the estimation error,” in Proceedings of The 28th Conference on Learning Theory , vol. 40. Paris, France: PMLR, 03–06 Jul 2015, pp. 1683–1709
2015
Cited alongside, same era.
A. Rudi and L. Rosasco, “Generalization properties of learning with random features,” 2016
2016
Cited alongside, same era.
2016
Cited alongside, same era.
A. Montanari, F. Ruan, Y. Sohn, and J. Yan, “The generalization error of max-margin linear classifiers: High-dimensional asymptotics in the overparametrized regime,” 2019
2019
Later among the works it cites.
M. Belkin, D. Hsu, S. Ma, and S. Mandal, “Reconciling modern machine-learning practice and the classical bias–variance trade-off,” Proceedings of the National Academy of Sciences , vol. 116, no. 32, pp. 15 849–15 854, 2019
2019
Later among the works it cites.
B. Ghorbani, S. Mei, T. Misiakiewicz, and A. Montanari, “Linearized two-layers neural networks in high dimension,” 2019
2019
Later among the works it cites.
S. Goldt, M. Mézard, F. Krzakala, and L. Zdeborová, “Modelling the influence of data structure on learning in neural networks: the hidden manifold model,” 2019
2019
Later among the works it cites.
F. Gerace, B. Loureiro, F. Krzakala, M. Mézard, and L. Zdeborová, “Generalisation error in learning with random features and the hidden manifold model,” 2020
2020
Closest in time.
J. Ba, M. Erdogdu, T. Suzuki, D. Wu, and T. Zhang, “Generalization of two-layer neural networks: An asymptotic viewpoint,” in International Conference on Learning Representations , 2020
2020
Closest in time.
P. Nakkiran, P. Venkat, S. Kakade, and T. Ma, “Optimal regularization can mitigate double descent,” 2020
2020
Closest in time.