Fetching the paper…
Reading the bibliography…
We consider the random-design least-squares regression problem within the reproducing kernel Hilbert space (RKHS) framework.
N. Aronszajn, “Theory of reproducing kernels,” Transactions of the American Mathematical Society
1950
Earlier work this paper cites.
H. Robbins and S. Monro, “A stochastic approxiation method,” The Annals of mathematical Statistics
1951
Earlier work this paper cites.
Dover publications, 1964
M. Abramowitz and I. Stegun, Handbook of mathematical functions · 1964
Earlier work this paper cites.
G. Kimeldorf and G. Wahba, “Some results on tchebycheffian spline functions,” Journal of Mathematical Analysis and Applications
1971
Earlier work this paper cites.
Masson, 1983
H. Brezis, Analyse fonctionnelle, Théorie et applications · 1983
Earlier work this paper cites.
SIAM, 1990
G. Wahba, Spline Models for observationnal data · 1990
Earlier work this paper cites.
I. M. Johnstone, “Minimax Bayes, asymptotic minimax and sparse wavelet priors,” Statistical Decision Theory and Related Topics
1994
Earlier work this paper cites.
J. Hopkins Univ. Press, 1996
G. H. Golub and C. F. V. Loan, Matrix Computations · 1996
Earlier work this paper cites.
H. W. Engl, M. Hanke, and N. A., “Regularization of inverse problems,” Klüwer Academic Publishers
1996
Earlier work this paper cites.
Courier Dover Publications, 1999
A. N. Kolmogorov and S. V. Fomin, Elements of the theory of functions and functional analysis · 1999
Earlier work this paper cites.
Pearson, 2000
B. Thomson, J. Bruckner, and A. M. Bruckner, Elementary real analysis · 2000
Earlier work this paper cites.
S. Smale and F. Cucker, “On the mathematical foundations of learning,” Bulletin of the American Mathematical Society
2001
Earlier work this paper cites.
C. Williams and M. Seeger, “Using the Nyström method to speed up kernel machines,” in Adv. NIPS
2001
Earlier work this paper cites.
MIT Press, 2002
B. Schölkopf and A. J. Smola, Learning with Kernels · 2002
Earlier work this paper cites.
F. Cucker and S. Smale, “Best choices for regularization parameters in learning theory: On the bias-variance problem,” Found. Comput. Math
2002
Earlier work this paper cites.
Cambridge University Press, 2004
J. Shawe-Taylor and N. Cristianini, Kernel Methods for Pattern Analysis · 2004
Earlier work this paper cites.
Springer, 2004
A. Berlinet and C. Thomas-Agnan, Reproducing kernel Hilbert spaces in probability and statistics · 2004
Earlier work this paper cites.
J. Kivinen, S. A.J., and R. C. Williamson, “Online learning with kernels,” IEEE transactions on signal processing
2004
Cited alongside, same era.
T. Zhang, “Solving large scale linear prediction problems using stochastic gradient descent algorithms,” ICML 2014 Proceedings of the twenty-first international conference on machine learning
2004
Cited alongside, same era.
E. De Vito, A. Caponetto, and L. Rosasco, “Model selection for regularized least-squares algorithm in learning theory,” Found. Comput. Math
2005
Cited alongside, same era.
T. Zhang and B. Yu, “Boosting with early stopping: convergence and consistency,” Annals of Statistics
2005
Cited alongside, same era.
O. Dekel, S. Shalev-Shwartz, and Y. Singer, “The Forgetron: A kernel-based perceptron on a fixed budget,” in Adv. NIPS
2005
Cited alongside, same era.
G. Raskutti, W. M.J., and Y. B., “Early stopping for non-parametric regression: An optimal data-dependent stopping rule,” 49th Annual Allerton Conference on Communication, Control, and Computing
2011
Later among the works it cites.
P. Tarrès and Y. Yao, “Online learning as stochastic approximation of regularization paths,” ArXiv e-prints 1103.5538, 2011
2011
Later among the works it cites.
S. Shalev-Shwartz, “Online learning and online convex optimization,” Foundations and Trends in Machine Learning
2011
Later among the works it cites.
F. Bach and E. Moulines, “Non-asymptotic analysis of stochastic approximation algorithms for machine learning,” in Adv. NIPS
2011
Later among the works it cites.
B. K. Sriperumbudur, K. Fukumizu, and G. R. Lanckriet, “Universality, characteristic kernels and rkhs embedding of measures,” The Journal of Machine Learning Research
2011
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A. Bordes, S. Ertekin, J. Weston, and L. Bottou, “Fast kernel classifiers with online and active learning,” Journal of Machine Learning Research
2005
Cited alongside, same era.
C. A. Micchelli, Y. Xu, and H. Zhang, “Universal kernels,” The Journal of Machine Learning Research
2006
Cited alongside, same era.
PhD thesis, University of California at Berkeley, 2006
Y. Yao, A dynamic Theory of Learning · 2006
Cited alongside, same era.
S. Smale and D.-X. Zhou, “Learning theory estimates via integral operators and their approximations,” Constructive Approximation
2007
Cited alongside, same era.
A. Caponnetto and E. De Vito, “Optimal Rates for the Regularized Least-Squares Algorithm,” Foundations of Computational Mathematics
2007
Cited alongside, same era.
Y. Yao, L. Rosasco, and A. Caponnetto, “On early stopping in gradient descent learning,” Constructive Approximation
2007
Cited alongside, same era.
Springer Publishing Company, Incorporated, 1st ed., 2008
A. B. Tsybakov, Introduction to Nonparametric Estimation · 2008
Cited alongside, same era.
Later among the works it cites.
M. W. Mahoney, “Randomized algorithms for matrices and data,” Foundations and Trends in Machine Learning
2011
Later among the works it cites.
E. Hazan and S. Kale, “Beyond the regret minimization barrier: an optimal algorithm for stochastic strongly-convex optimization,” Proceedings of the International Conference on Learning Theory (COLT)
2011
Later among the works it cites.
F. Bach, “Sharp analysis of low-rank kernel matrix approximations,” Proceedings of the International Conference on Learning Theory (COLT)
2012
Later among the works it cites.
S. Lacoste-Julien, M. Schmidt, and F. Bach, “A simpler approach to obtaining an O(1/t) rate for the stochastic projected subgradient method,” ArXiv e-prints 1212.2002, 2012
2012
Later among the works it cites.
F. Bach and E. Moulines, “Non-strongly-convex smooth stochastic approximation with convergence rate O(1/n),” Advances in Neural Information Processing Systems (NIPS)
2013
Later among the works it cites.
F. Bach, “Sharp analysis of low-rank kernel matrix approximations,” in Proceedings of the International Conference on Learning Theory (COLT)
2013
Later among the works it cites.
L. Rosasco, A. Tacchetti, and S. Villa, “Regularization by Early Stopping for Online Learning Algorithms,” ArXiv e-prints
2014
Closest in time.
P. Mikusinski and E. Weiss, “The Bochner Integral,” ArXiv e-prints
2014
Closest in time.
J. Vert, “Kernel Methods,” 2014
2014
Closest in time.
D. Hsu, S. M. Kakade, and T. Zhang, “Random design analysis of ridge regression,” Foundations of Computational Mathematics
2014
Closest in time.
N. Flammarion and F. Bach, “From averaging to acceleration, there is only a step-size,” in Proceedings of the International Conference on Learning Theory (COLT)
2015
Closest in time.