Fetching the paper…
Reading the bibliography…
The strategy of early stopping is a regularization technique based on choosing a stopping time for an iterative algorithm.
Functions of positive and negative type and their connection with the theory of integral equations
J. Mercer · 1909
Earlier work this paper cites.
Theory of reproducing kernels
N. Aronszajn · 1950
Earlier work this paper cites.
Piecewise-polynomial approximations of functions of the classes W p α W_{p}^{\alpha}
M. S. Birman and M. Z. Solomjak · 1967
Earlier work this paper cites.
Some results on Tchebycheffian spline functions
G. Kimeldorf and G. Wahba · 1971
Earlier work this paper cites.
A bound on tail probabilities for quadratic forms in independent random variables whose distributions are not necessarily symmetric
F. T. Wright · 1973
Earlier work this paper cites.
Theory and methods related to the singular value expansion and Landweber’s iteration for integral equations of the first kind
O. N. Strand · 1974
Earlier work this paper cites.
Distribution-free inequalities for the deleted and holdout error estimates
L. Devroye and T. J. Wagner · 1979
Earlier work this paper cites.
A formal comparison of methods proposed for the numerical solution of first kind integral equations
R. S. Anderssen and P. M. Prenter · 1981
Earlier work this paper cites.
Estimation of the mean of a multivariate normal distribution
C. M. Stein · 1981
Earlier work this paper cites.
Reproducing Kernel Hilbert Spaces : Applications in Statistical Signal Processing
H. L. Weinert (ed.), editor · 1982
Earlier work this paper cites.
Additive regression and other nonparametric models
C. J. Stone · 1985
Earlier work this paper cites.
Three topics in ill-posed problems
G. Wahba · 1987
Earlier work this paper cites.
Theory of Reproducing Kernels and its Applications
S. Saitoh · 1988
Earlier work this paper cites.
On variance estimation in nonparametric regression
P. Hall and J.S. Marron · 1990
Earlier work this paper cites.
Generalization and parameter estimation in feedforward nets: Some experiments
N. Morgan and H. Bourlard · 1990
Earlier work this paper cites.
Spline models for observational data
G. Wahba · 1990
Cited alongside, same era.
A decision-theoretic generalization of on-line learning and an application to boosting
Y. Freund and R. Schapire · 1997
Cited alongside, same era.
Boosting algorithms as gradient descent
L. Mason, J. Baxter, P. Bartlett, and M. Frean · 1999
Cited alongside, same era.
Information-theoretic determination of minimax rates of convergence
Y. Yang and A. Barron · 1999
Cited alongside, same era.
Empirical Processes in M-Estimation
S. van de Geer · 2000
Cited alongside, same era.
Maximum likelihood estimation by markov chain monte carlo approximation
M. G. Gu and H. T. Zhu · 2001
Cited alongside, same era.
On the Nyström method for approximating a Gram matrix for improved kernel-based learning
P. Drineas and M. W. Mahoney · 2005
Later among the works it cites.
Boosting with early stopping: Convergence and consistency
T. Zhang and B. Yu · 2005
Later among the works it cites.
Convexity, classification, and risk bounds
P. L. Bartlett, M. I. Jordan, and J. D. McAuliffe · 2006
Later among the works it cites.
Adaptation for regularization operators in learning theory
A. Caponetto and Y. Yao · 2006
Later among the works it cites.
Optimal rates for regularization operators in learning theory
A. Caponneto · 2006
Later among the works it cites.
Adaboost is consistent
P. Bartlett and M. Traskin · 2007
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
M. Ledoux · 2001
Cited alongside, same era.
Gaussian and Rademacher complexities: Risk bounds and structural results
P. Bartlett and S. Mendelson · 2002
Cited alongside, same era.
Smoothing spline ANOVA models
C. Gu · 2002
Cited alongside, same era.
Geometric parameters of kernel machines
S. Mendelson · 2002
Cited alongside, same era.
Learning with Kernels
B. Schölkopf and A. Smola · 2002
Cited alongside, same era.
Boosting with L 2 L^{2} loss: Regression and classification
P. Buhlmann and B.Yu · 2003
Cited alongside, same era.
On regularization algorithms in learning theory
F. Bauer, S. Pereverzev, and L. Rosasco · 2007
Later among the works it cites.
On early stopping in gradient descent learning
Y. Yao, L. Rosasco, and A. Caponnetto · 2007
Later among the works it cites.
Approximation and learning by greedy algorithms
A. R. Barron, A. Cohen, W. Dahmen, and R. A. DeVore · 2008
Later among the works it cites.
Optimal learning rates for kernel conjugate gradient regression
G. Blanchard and M. Kramer · 2010
Later among the works it cites.
Cross-validation based adaptation for regularization operators in learning theory
A. Caponetto and Y. Yao · 2010
Later among the works it cites.
Adaptive kernel methods using the balancing principle
E. De Vito, S. Pereverzyev, and L. Rosasco · 2010
Later among the works it cites.
Implementing regularization implicitly via approximate eigenvector computation
L. Orecchia and M. W. Mahoney · 2011
Later among the works it cites.
Minimax-optimal rates for sparse additive models over kernel classes via convex programming
G. Raskutti, M. J. Wainwright, and B. Yu · 2012
Later among the works it cites.