Fantastic generalization measures and where to find them
Jiang, Y., Neyshabur, B., Mobahi, H., Krishnan, D., and Bengio, S · 2020
Later among the works it cites.
Scaling laws for neural language models
Original
Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D · 2020
Later among the works it cites.
Finite versus infinite neural networks: an empirical study
Lee, J., Schoenholz, S. S., Pennington, J., Adlam, B., Xiao, L., Novak, R., and Sohl-Dickstein, J · 2020
Later among the works it cites.
The generalization error of random features regression: Precise asymptotics and double descent curve
Original
Mei, S. and Montanari, A · 2020
Later among the works it cites.
From fixed-X to random-X regression: Bias-variance decompositions, covariance penalties, and prediction error estimation
Rosset, S. and Tibshirani, R. J · 2020
Later among the works it cites.
On the optimal weighted ℓ 2 \ell_{2} regularization in overparameterized linear regression
Wu, D. and Xu, J · 2020
Later among the works it cites.
Explaining neural scaling laws
Original
Bahri, Y., Dyer, E., Kaplan, J., Lee, J., and Sharma, U · 2021
Later among the works it cites.
Cross-validation: what does it estimate and how well does it do it?
Original
Bates, S., Hastie, T., and Tibshirani, R · 2021
Later among the works it cites.
Spectral bias and task-model alignment explain generalization in kernel regression and infinitely wide neural networks
Canatar, A., Bordelon, B., and Pehlevan, C · 2021
Later among the works it cites.
Generalization error rates in kernel regression: The crossover from the noiseless to noisy regime
Cui, H., Loureiro, B., Krzakala, F., and Zdeborova, L · 2021
Later among the works it cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., Uszkoreit, J., and Houlsby, N · 2021
Later among the works it cites.
Learning curves of generic features maps for realistic datasets with a teacher-student model
Loureiro, B., Gerbelot, C., Cui, H., Goldt, S., Krzakala, F., Mezard, M., and Zdeborova, L · 2021
Later among the works it cites.
A theory of high dimensional regression with arbitrary correlations between input features and target functions: sample complexity, multiple descent curves and a hierarchy of phase transitions
Mel, G. and Ganguli, S · 2021
Later among the works it cites.
Pervasive label errors in test sets destabilize machine learning benchmarks
Northcutt, C. G., Athalye, A., and Mueller, J · 2021
Later among the works it cites.
Uniform consistency of cross-validation estimators for high-dimensional ridge regression
Patil, P., Wei, Y., Rinaldo, A., and Tibshirani, R. J · 2021
Later among the works it cites.
Asymptotics of ridge(less) regression under general source condition
Richards, D., Mourtada, J., and Rosasco, L · 2021
Later among the works it cites.
A first-principles theory of neural network generalization, October 2021
Simon, J. B · 2021
Later among the works it cites.
Neural tangent kernel eigenvalues accurately predict generalization
Original
Simon, J. B., Dickens, M., and DeWeese, M. R · 2021
Later among the works it cites.
Robust and nonparametric statistics, April 2021
Steinhardt, J · 2021
Later among the works it cites.
Learning Theory from First Principles
Bach, F · 2023
Closest in time.