Fetching the paper…
Reading the bibliography…
We characterize the power-law asymptotics of learning curves for Gaussian process regression (GPR) under the assumption that the eigenspectrum of the prior and the eigenexpansion coefficients of the target function follow a power law.
Asymptotic behavior of the eigenvalues of certain integral equations
H. Widom · 1963
Earlier work this paper cites.
Four types of learning curves
S. Amari, N. Fujita, and S. Shinomoto · 1992
Earlier work this paper cites.
Statistical theory of learning curves under entropic loss criterion
S. Amari and N. Murata · 1993
Earlier work this paper cites.
Multivariate integration and approximation for random fields satisfying Sacks-Ylvisaker conditions
K. Ritter, G. W. Wasilkowski, and H. Woźniakowski · 1995
Earlier work this paper cites.
Rigorous learning curve bounds from statistical mechanics
D. Haussler, M. Kearns, H. S. Seung, and N. Tishby · 1996
Earlier work this paper cites.
Bayesian Learning for Neural Networks
R. M. Neal · 1996
Earlier work this paper cites.
Mutual information, metric entropy and cumulative relative entropy risk
D. Haussler and M. Opper · 1997
Earlier work this paper cites.
Computing with infinite networks
C. K. Williams · 1997
Earlier work this paper cites.
Information-theoretic characterization of Bayes performance and the choice of priors in parametric and nonparametric problems
A. R. Barron · 1998
Earlier work this paper cites.
General bounds on Bayes errors for regression with Gaussian processes
M. Opper and F. Vivarelli · 1999
Earlier work this paper cites.
Learning curves for Gaussian processes
P. Sollich · 1999
Earlier work this paper cites.
Upper and lower bounds on the learning curve for Gaussian processes
C. K. Williams and F. Vivarelli · 2000
Earlier work this paper cites.
Gaussian process regression with mismatched models
P. Sollich · 2001
Earlier work this paper cites.
A variational approach to learning curves
M. Opper and D. Malzahn · 2002
Earlier work this paper cites.
Learning curves for Gaussian process regression: Approximations and bounds
P. Sollich and A. Halees · 2002
Earlier work this paper cites.
Accurate error bounds for the eigenvalues of the kernel matrix
M. L. Braun · 2006
Earlier work this paper cites.
Gaussian processes for machine learning
C. K. Williams and C. E. Rasmussen · 2006
Earlier work this paper cites.
Optimal rates for the regularized least-squares algorithm
A. Caponnetto and E. De Vito · 2007
Earlier work this paper cites.
Information consistency of nonparametric Gaussian process methods
M. W. Seeger, S. M. Kakade, and D. P. Foster · 2008
Earlier work this paper cites.
Kernel methods for deep learning
Y. Cho and L. K. Saul · 2009
Earlier work this paper cites.
Optimal rates for regularized least squares regression
I. Steinwart, D. R. Hush, C. Scovel, et al · 2009
Earlier work this paper cites.
Algebraic Geometry and Statistical Learning Theory
S. Watanabe · 2009
Cited alongside, same era.
Bayesian nonparametric models
P. Orbanz and Y. W. Teh · 2010
Cited alongside, same era.
Information rates of nonparametric Gaussian process methods
A. Van Der Vaart and H. Van Zanten · 2011
Cited alongside, same era.
Interpolation of spatial data: Some theory for kriging
M. L. Stein · 2012
Cited alongside, same era.
User-friendly tail bounds for sums of random matrices
J. A. Tropp · 2012
Cited alongside, same era.
Asymptotic analysis of the learning curve for Gaussian process regression
L. Le Gratiet and J. Garnier · 2015
Cited alongside, same era.
Minimizers of the empirical risk and risk monotonicity
M. Loog, T. Viering, and A. Mey · 2019
Later among the works it cites.
Bayesian deep convolutional networks with many channels are gaussian processes
R. Novak, L. Xiao, Y. Bahri, J. Lee, G. Yang, D. A. Abolafia, J. Pennington, and J. Sohl-Dickstein · 2019
Later among the works it cites.
The convergence rate of neural networks for learned functions of different frequencies
B. Ronen, D. Jacobs, Y. Kasten, and S. Kritchman · 2019
Later among the works it cites.
Open problem: Monotonicity of learning
T. Viering, A. Mey, and M. Loog · 2019
Later among the works it cites.
Wide feedforward or recurrent neural networks of any architecture are gaussian processes
G. Yang · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Toward deeper understanding of neural networks: The power of initialization and a dual view on expressivity
A. Daniely, R. Frostig, and Y. Singer · 2016
Cited alongside, same era.
Deep learning scaling is predictable, empirically
J. Hestness, S. Narang, N. Ardalani, G. Diamos, H. Jun, H. Kianinejad, M. Patwary, M. Ali, Y. Yang, and Y. Zhou · 2017
Cited alongside, same era.
To understand deep learning we need to understand kernel learning
M. Belkin, S. Ma, and S. Mandal · 2018
Cited alongside, same era.
Optimal rates for regularization of statistical inverse learning problems
G. Blanchard and N. Mücke · 2018
Cited alongside, same era.
Gaussian process behaviour in wide deep neural networks
A. G. de G. Matthews, J. Hron, M. Rowland, R. E. Turner, and Z. Ghahramani · 2018
Cited alongside, same era.
Neural tangent kernel: Convergence and generalization in neural networks
A. Jacot, F. Gabriel, and C. Hongler · 2018
Cited alongside, same era.
G. Yang and H. Salman · 2019
Later among the works it cites.
Spectrum dependent learning curves in kernel regression and wide neural networks
B. Bordelon, A. Canatar, and C. Pehlevan · 2020
Later among the works it cites.
Every model learned by gradient descent is approximately a kernel machine
P. Domingos · 2020
Later among the works it cites.
Sobolev norm learning rates for regularized least-squares algorithms
S. Fischer and I. Steinwart · 2020
Later among the works it cites.
Finite versus infinite neural networks: an empirical study
J. Lee, S. Schoenholz, J. Pennington, B. Adlam, L. Xiao, R. Novak, and J. Sohl-Dickstein · 2020
Later among the works it cites.
Asymptotic learning curves of kernel methods: empirical data versus teacher–student paradigm
S. Spigler, M. Geiger, and M. Wyart · 2020
Later among the works it cites.
Explaining neural scaling laws
Y. Bahri, E. Dyer, J. Kaplan, J. Lee, and U. Sharma · 2021
Closest in time.
On the sample complexity of learning with geometric stability
A. Bietti, L. Venturi, and J. Bruna · 2021
Closest in time.
A theory of universal learning
O. Bousquet, S. Hanneke, S. Moran, R. van Handel, and A. Yehudayoff · 2021
Closest in time.
Spectral bias and task-model alignment explain generalization in kernel regression and infinitely wide neural networks
A. Canatar, B. Bordelon, and C. Pehlevan · 2021
Closest in time.
Generalization error rates in kernel regression: The crossover from the noiseless to noisy regime
H. Cui, B. Loureiro, F. Krzakala, and L. Zdeborová · 2021
Closest in time.
Optimal rates for averaged stochastic gradient descent under neural tangent kernel regime
A. Nitanda and T. Suzuki · 2021
Closest in time.
On information gain and regret bounds in Gaussian process bandits
S. Vakili, K. Khezeli, and V. Picheny · 2021
Closest in time.
Universal scaling laws in the gradient descent training of neural networks
M. Velikanov and D. Yarotsky · 2021
Closest in time.
The shape of learning curves: A review
T. Viering and M. Loog · 2021
Closest in time.