Fetching the paper…
Reading the bibliography…
A key property of neural networks driving their success is their ability to learn features from data.
A useful theorem for nonlinear devices having gaussian inputs
Price, R · 1958
Earlier work this paper cites.
A simple weight decay can improve generalization
Krogh, A. and Hertz, J · 1991
Earlier work this paper cites.
Suppressing chaos in neural networks by noise
Molgedey, L., Schuchhardt, J., and Schuster, H · 1992
Earlier work this paper cites.
Bayesian Learning for Neural Networks
Neal, R. M · 1996
Earlier work this paper cites.
The Fokker-Planck Equation
Risken, H · 1996
Earlier work this paper cites.
Computing with infinite networks
Williams, C · 1996
Earlier work this paper cites.
The mnist database of handwritten digits, 1998
LeCun, Y., Cortes, C., and Burges, C. J · 1998
Earlier work this paper cites.
Probability, Random Variables, and Stochastic Processes
Papoulis, A. and Pillai, S. U · 2002
Earlier work this paper cites.
On the asymptotics of wide networks with polynomial activations, 2020
Aitken, K. and Gur-Ari, G · 2006
Earlier work this paper cites.
The Gaussian equivalence of generative models for learning with two-layer neural networks
Goldt, S., Reeves, G., Mézard, M., Krzakala, F., and Zdeborová, L · 2006
Earlier work this paper cites.
The large deviation approach to statistical mechanics
Touchette, H · 2009
Earlier work this paper cites.
Feature Learning in Infinite-Width Neural Networks
Yang, G. and Hu, E. J · 2011
Earlier work this paper cites.
Algorithms for learning kernels based on centered alignment
Cortes, C., Mohri, M., and Rostamizadeh, A · 2012
Earlier work this paper cites.
Exact solutions to the nonlinear dynamics of learning in deep linear neural networks
Saxe, A., Mcclelland, J., and Ganguli, S · 2014
Earlier work this paper cites.
Exponential expressivity in deep neural networks through transient chaos
Poole, B., Lahiri, S., Raghu, M., Sohl-Dickstein, J., and Ganguli, S · 2016
Earlier work this paper cites.
Deep neural networks as gaussian processes
Lee, J., Bahri, Y., Novak, R., Schoenholz, S. S., Pennington, J., and Sohl-Dickstein, J · 2017
Earlier work this paper cites.
Resurrecting the sigmoid in deep learning through dynamical isometry: theory and practice
Pennington, J., Schoenholz, S. S., and Ganguli, S · 2017
Earlier work this paper cites.
Deep information propagation
Schoenholz, S. S., Gilmer, J., Ganguli, S., and Sohl-Dickstein, J · 2017
Earlier work this paper cites.
Neural tangent kernel: Convergence and generalization in neural networks
Jacot, A., Gabriel, F., and Hongler, C · 2018
Earlier work this paper cites.
Deep neural networks as gaussian processes
Lee, J., Sohl-Dickstein, J., Pennington, J., Novak, R., Schoenholz, S., and Bahri, Y · 2018
Earlier work this paper cites.
Finite size corrections for neural network gaussian processes
Antognini, J. M · 2019
Cited alongside, same era.
Initialization of relus for dynamical isometry
Burkholz, R. and Dubatovka, A · 2019
Cited alongside, same era.
On lazy training in differentiable programming
Chizat, L., Oyallon, E., and Bach, F · 2019
Cited alongside, same era.
Deep convolutional networks as shallow gaussian processes
Garriga-Alonso, A., Rasmussen, C. E., and Aitchison, L · 2019
Cited alongside, same era.
Wide neural networks of any depth evolve as linear models under gradient descent
Lee, J., Xiao, L., Schoenholz, S., Bahri, Y., Novak, R., Sohl-Dickstein, J., and Pennington, J · 2019
Cited alongside, same era.
Bayesian deep convolutional networks with many channels are gaussian processes
Novak, R., Xiao, L., Bahri, Y., Lee, J., Yang, G., Abolafia, D. A., Pennington, J., and Sohl-dickstein, J · 2019
Classifying high-dimensional gaussian mixtures: Where kernel methods fail and neural networks succeed
Refinetti, M., Goldt, S., Krzakala, F., and Zdeborova, L · 2021
Later among the works it cites.
Tuning large neural networks via zero-shot hyperparameter transfer
Yang, G., Hu, E. J., Babuschkin, I., Sidor, S., Liu, X., Farhi, D., Ryder, N., Pachocki, J., Chen, W., and Gao, J · 2021
Later among the works it cites.
Asymptotics of representation learning in finite bayesian neural networks
Zavatone-Veth, J. A., Canatar, A., Ruben, B., and Pehlevan, C · 2021
Later among the works it cites.
On the opportunities and risks of foundation models
Bommasani, R., Hudson, D. A., Adeli, E., Altman, R., Arora, S., von Arx, S., Bernstein, M. S., Bohg, J., Bosselut, A., Brunskill, E., Brynjolfsson, E., Buch, S., Card, D., Castellon, R., Chatterji, N., Chen, A., Creel, K., Davis, J. Q., Demszky, D., Donahue, C., Doumbouya, M., Durmus, E., Ermon, S., Etchemendy, J., Ethayarajh, K., Fei-Fei, L., Finn, C., Gale, T., Gillespie, L., Goel, K., Goodman, N., Grossman, S., Guha, N., Hashimoto, T., Henderson, P., Hewitt, J., Ho, D. E., Hong, J., Hsu, K., Huang, J., Icard, T., Jain, S., Jurafsky, D., Kalluri, P., Karamcheti, S., Keeling, G., Khani, F., Khattab, O., Koh, P. W., Krass, M., Krishna, R., Kuditipudi, R., Kumar, A., Ladhak, F., Lee, M., Lee, T., Leskovec, J., Levent, I., Li, X. L., Li, X., Ma, T., Malik, A., Manning, C. D., Mirchandani, S., Mitchell, E., Munyikwa, Z., Nair, S., Narayan, A., Narayanan, D., Newman, B., Nie, A., Niebles, J. C., Nilforoshan, H., Nyarko, J., Ogut, G., Orr, L., Papadimitriou, I., Park, J. S., Piech, C., Portelance, E., Potts, C., Raghunathan, A., Reich, R., Ren, H., Rong, F., Roohani, Y., Ruiz, C., Ryan, J., Ré, C., Sadigh, D., Sagawa, S., Santhanam, K., Shih, A., Srinivasan, K., Tamkin, A., Taori, R., Thomas, A. W., Tramér, F., Wang, R. E., Wang, W., Wu, B., Wu, J., Wu, Y., Xie, S. M., Yasunaga, M., You, J., Zaharia, M., Zhang, M., Zhang, T., Zhang, X., Zhang, Y., Zheng, L., Zhou, K., and Liang, P · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Wide feedforward or recurrent neural networks of any architecture are gaussian processes
Yang, G · 2019
Cited alongside, same era.
Asymptotics of wide networks from feynman diagrams
Dyer, E. and Gur-Ari, G · 2020
Cited alongside, same era.
Disentangling feature and lazy training in deep neural networks
Geiger, M., Spigler, S., Jacot, A., and Wyart, M · 2020
Cited alongside, same era.
Infinite attention: NNGP and NTK for deep attention networks
Hron, J., Bahri, Y., Sohl-Dickstein, J., and Novak, R · 2020
Cited alongside, same era.
Dynamics of deep neural networks and neural tangent hierarchy
Huang, J. and Yau, H.-T · 2020
Cited alongside, same era.
Finite versus infinite neural networks: an empirical study
Lee, J., Schoenholz, S., Pennington, J., Adlam, B., Xiao, L., Novak, R., and Sohl-Dickstein, J · 2020
Cited alongside, same era.
Later among the works it cites.
A kernel analysis of feature learning in deep neural networks
Canatar, A. and Pehlevan, C · 2022
Later among the works it cites.
The generalization error of random features regression: Precise asymptotics and the double descent curve
Mei, S. and Montanari, A · 2022
Later among the works it cites.
Learning sparse features can lead to overfitting in neural networks
Petrini, L., Cagnetta, F., Vanden-Eijnden, E., and Wyart, M · 2022
Later among the works it cites.
The Principles of Deep Learning Theory
Roberts, D. A., Yaida, S., and Hanin, B · 2022
Later among the works it cites.
Unified field theoretical approach to deep and recurrent neuronal networks
Segadlo, K., Epping, B., van Meegen, A., Dahmen, D., Krämer, M., and Helias, M · 2022
Later among the works it cites.
Contrasting random and learned features in deep bayesian linear regression
Zavatone-Veth, J. A., Tong, W. L., and Pehlevan, C · 2022
Later among the works it cites.
Self-consistent dynamical field theory of kernel evolution in wide neural networks*
Bordelon, B. and Pehlevan, C · 2023
Later among the works it cites.
Criticality versus uniformity in deep neural networks
Bukva, A., de Gier, J., Grosvenor, K. T., Jefferson, R., Schalm, K., and Schwander, E · 2023
Later among the works it cites.
Bayes-optimal learning of deep random networks of extensive-width
Cui, H., Krzakala, F., and Zdeborova, L · 2023
Later among the works it cites.
Optimal signal propagation in resnets through residual scaling
Fischer, K., Dahmen, D., and Helias, M · 2023
Later among the works it cites.
Bayesian interpolation with deep linear networks
Hanin, B. and Zlokapa, A · 2023
Later among the works it cites.
A theory of data variability in neural network bayesian inference
Lindner, J., Dahmen, D., Krämer, M., and Helias, M · 2023
Later among the works it cites.
A statistical mechanics framework for bayesian deep neural networks beyond the infinite-width limit
Pacelli, R., Ariosto, S., Pastore, M., Ginelli, F., Gherardi, M., and Rotondo, P · 2023
Later among the works it cites.
A theory of representation learning gives a deep generalisation of kernel methods
Yang, A. X., Robeyns, M., Milsom, E., Anson, B., Schoots, N., and Aitchison, L · 2023
Later among the works it cites.
Baglioni, P., Pacelli, R., Aiudi, R., Di Renzo, F., Vezzani, A., Burioni, R., and Rotondo, P · 2024
Closest in time.
Separation of scales and a thermodynamic description of feature learning in some cnns
Seroussi, I., Naveh, G., and Ringel, Z · 2041
Closest in time.