Fetching the paper…
Reading the bibliography…
Deep neural networks (DNNs) in the infinite width/channel limit have received much attention recently, as they provide a clear analytical window to deep learning via mappings to Gaussian Processes (GPs).
Statistical field theory for neural networks
Helias, M. and Dahmen, D. (2019) · 1901
Earlier work this paper cites.
Wide neural networks of any depth evolve as linear models under gradient descent
Lee, J., Xiao, L., Schoenholz, S. S., Bahri, Y., Novak, R., Sohl-Dickstein, J., and Pennington, J. (2019) · 1902
Earlier work this paper cites.
On Exact Computation with an Infinitely Wide Neural Net
Arora, S., Du, S. S., Hu, W., Li, Z., Salakhutdinov, R., and Wang, R. (2019) · 1904
Earlier work this paper cites.
Learning Curves for Deep Neural Networks: A Gaussian Field Theory Perspective
Cohen, O., Malka, O., and Ringel, Z. (2019) · 1906
Earlier work this paper cites.
Saddlepoint Approximations in Statistics
Daniels, H. E. (1954) · 1954
Earlier work this paper cites.
Neural networks and principal component analysis: Learning from examples without local minima
Baldi, P. and Hornik, K. (1989) · 1989
Earlier work this paper cites.
Priors for infinite networks
Neal, R. M. (1996) · 1996
Earlier work this paper cites.
Predicting the outputs of finite networks trained with noisy gradients
Naveh, G., Ben-David, O., Sompolinsky, H., and Ringel, Z. (2020) · 2004
Earlier work this paper cites.
Understanding gaussian process regression using the equivalent kernel
Sollich, P. and Williams, C. K. (2004) · 2004
Earlier work this paper cites.
Gaussian Processes for Machine Learning (Adaptive Computation and Machine Learning)
Rasmussen, C. E. and Williams, C. K. I. (2005) · 2005
Earlier work this paper cites.
The recurrent neural tangent kernel
Alemohammad, S., Wang, Z., Balestriero, R., and Baraniuk, R. (2020) · 2006
Earlier work this paper cites.
When do neural networks outperform kernel methods?
Ghorbani, B., Mei, S., Misiakiewicz, T., and Montanari, A. (2020) · 2006
Earlier work this paper cites.
Optimization and generalization of shallow neural networks with quadratic activation functions
Mannelli, S. S., Vanden-Eijnden, E., and Zdeborová, L. (2020) · 2006
Earlier work this paper cites.
Finite versus infinite neural networks: an empirical study
Lee, J., Schoenholz, S. S., Pennington, J., Adlam, B., Xiao, L., Novak, R., and Sohl-Dickstein, J. (2020) · 2007
Earlier work this paper cites.
Asymptotics of wide convolutional neural networks
Andreassen, A. and Dyer, E. (2020) · 2008
Cited alongside, same era.
The eigenvalues and eigenvectors of finite, low rank perturbations of large random matrices
Benaych-Georges, F. and Nadakuditi, R. R. (2011) · 2011
Cited alongside, same era.
Feature learning in infinite-width neural networks
Yang, G. and Hu, E. J. (2020) · 2011
Cited alongside, same era.
The singular values and vectors of low rank perturbations of large rectangular random matrices
Benaych-Georges, F. and Nadakuditi, R. R. (2012) · 2012
Cited alongside, same era.
Statistical mechanics of deep linear neural networks: The back-propagating renormalization group
Li, Q. and Sompolinsky, H. (2020) · 2012
Gaussian process behaviour in wide deep neural networks
Matthews, A. G. d. G., Rowland, M., Hron, J., Turner, R. E., and Ghahramani, Z. (2018) · 2018
Later among the works it cites.
Bayesian Deep Convolutional Networks with Many Channels are Gaussian Processes
Novak, R., Xiao, L., Lee, J., Bahri, Y., Yang, G., Abolafia, D. A., Pennington, J., and Sohl-Dickstein, J. (2018) · 2018
Later among the works it cites.
On lazy training in differentiable programming
Chizat, L., Oyallon, E., and Bach, F. (2019) · 2019
Later among the works it cites.
Spectrum dependent learning curves in kernel regression and wide neural networks
Bordelon, B., Canatar, A., and Pehlevan, C. (2020) · 2020
Later among the works it cites.
Asymptotics of wide networks from feynman diagrams
Dyer, E. and Gur-Ari, G. (2020) · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Exact solutions to the nonlinear dynamics of learning in deep linear neural networks
Saxe, A. M., McClelland, J. L., and Ganguli, S. (2013) · 2013
Cited alongside, same era.
Feature learning in deep neural networks - studies on speech recognition
Yu, D., Seltzer, M., Li, J., Huang, J.-T., and Seide, F. (2013) · 2013
Cited alongside, same era.
How transferable are features in deep neural networks?
Yosinski, J., Clune, J., Bengio, Y., and Lipson, H. (2014) · 2014
Cited alongside, same era.
Tensor Methods in Statistics
Mccullagh, P. (2017) · 2017
Cited alongside, same era.
Implicit bias of gradient descent on linear convolutional networks
Gunasekar, S., Lee, J., Soudry, D., and Srebro, N. (2018) · 2018
Cited alongside, same era.
Neural Tangent Kernel: Convergence and Generalization in Neural Networks
Jacot, A., Gabriel, F., and Hongler, C. (2018) · 2018
Cited alongside, same era.
An analytic theory of generalization dynamics and transfer learning in deep linear networks
Lampinen, A. K. and Ganguli, S. (2018) · 2018
Cited alongside, same era.
Disentangling feature and lazy training in deep neural networks
Geiger, M., Spigler, S., Jacot, A., and Wyart, M. (2020) · 2020
Later among the works it cites.
Infinite attention: Nngp and ntk for deep attention networks
Hron, J., Bahri, Y., Sohl-Dickstein, J., and Novak, R. (2020) · 2020
Later among the works it cites.
Non-gaussian processes and neural networks at finite widths
Yaida, S. (2020) · 2020
Later among the works it cites.
Spectral bias and task-model alignment explain generalization in kernel regression and infinitely wide neural networks
Canatar, A., Bordelon, B., and Pehlevan, C. (2021) · 2021
Closest in time.
Landscape and training regimes in deep learning
Geiger, M., Petrini, L., and Wyart, M. (2021) · 2021
Closest in time.
Quantifying the benefit of using differentiable learning over tangent kernels
Malach, E., Kamath, P., Abbe, E., and Srebro, N. (2021) · 2021
Closest in time.
Refinetti, M., Goldt, S., Krzakala, F., and Zdeborová, L. (2021) · 2021
Closest in time.
Exact priors of finite neural networks
Zavatone-Veth, J. A. and Pehlevan, C. (2021) · 2021
Closest in time.