Statistical dynamics of classical systems
Martin, P. C., Siggia, E., and Rose, H · 1973
Earlier work this paper cites.
Dynamic theory of the spin-glass phase
Sompolinsky, H. and Zippelius, A · 1981
Earlier work this paper cites.
Fast rates for regularized least-squares algorithm
Caponnetto, A. and Vito, E. D · 2005
Earlier work this paper cites.
Algorithms for learning kernels based on centered alignment
Cortes, C., Mohri, M., and Rostamizadeh, A · 2012
Earlier work this paper cites.
Wide residual networks
Original
Zagoruyko, S. and Komodakis, N · 2016
Earlier work this paper cites.
Deep learning scaling is predictable, empirically
Original
Hestness, J., Narang, S., Ardalani, N., Diamos, G., Jun, H., Kianinejad, H., Patwary, M. M. A., Yang, Y., and Zhou, Y · 2017
Earlier work this paper cites.
Path integral approach to random neural networks
Crisanti, A. and Sompolinsky, H · 2018
Earlier work this paper cites.
Neural tangent kernel: Convergence and generalization in neural networks
Jacot, A., Gabriel, F., and Hongler, C · 2018
Earlier work this paper cites.
On lazy training in differentiable programming
Chizat, L., Oyallon, E., and Bach, F · 2019
Earlier work this paper cites.
Passed & spurious: Descent algorithms and local minima in spiked matrix-tensor models
Mannelli, S. S., Krzakala, F., Urbani, P., and Zdeborova, L · 2019
Earlier work this paper cites.
High-dimensional dynamics of generalization error in neural networks
Advani, M. S., Saxe, A. M., and Sompolinsky, H · 2020
Earlier work this paper cites.
Spectrum dependent learning curves in kernel regression and wide neural networks
Bordelon, B., Canatar, A., and Pehlevan, C · 2020
Earlier work this paper cites.
Asymptotics of wide networks from feynman diagrams
Dyer, E. and Gur-Ari, G · 2020
Earlier work this paper cites.
Double trouble in double descent: Bias and variance (s) in the lazy regime
d’Ascoli, S., Refinetti, M., Biroli, G., and Krzakala, F · 2020
Earlier work this paper cites.
Deep learning versus kernel learning: an empirical study of loss landscape geometry and the time evolution of the neural tangent kernel
Fort, S., Dziugaite, G. K., Paul, M., Kharaghani, S., Roy, D. M., and Ganguli, S · 2020
Earlier work this paper cites.
Scaling description of generalization with number of parameters in deep learning
Geiger, M., Jacot, A., Spigler, S., Gabriel, F., Sagun, L., d’Ascoli, S., Biroli, G., Hongler, C., and Wyart, M · 2020
Earlier work this paper cites.
Statistical field theory for neural networks , volume 970
Helias, M. and Dahmen, D · 2020
Earlier work this paper cites.
Scaling laws for neural language models
Original
Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D · 2020
Earlier work this paper cites.
Dynamical mean-field theory for stochastic gradient descent in gaussian mixture classification
Mignacco, F., Krzakala, F., Urbani, P., and Zdeborová, L · 2020
Earlier work this paper cites.