Fetching the paper…
Reading the bibliography…
A recent line of works studied wide deep neural networks (DNNs) by approximating them as Gaussian Processes (GPs).
A Simple Baseline for Bayesian Uncertainty in Deep Learning
Maddox, W., Garipov, T., Izmailov, P., Vetrov, D., and Wilson, A. G. (2019) · 1902
Earlier work this paper cites.
Mean field limit of the learning dynamics of multilayer neural networks
Nguyen, P.-M. (2019) · 1902
Earlier work this paper cites.
On Exact Computation with an Infinitely Wide Neural Net
Arora, S., Du, S. S., Hu, W., Li, Z., Salakhutdinov, R., and Wang, R. (2019) · 1904
Earlier work this paper cites.
A mean-field limit for certain deep neural networks
Araújo, D., Oliveira, R. I., and Yukimura, D. (2019) · 1906
Earlier work this paper cites.
The Convergence Rate of Neural Networks for Learned Functions of Different Frequencies
Basri, R., Jacobs, D., Kasten, Y., and Kritchman, S. (2019) · 1906
Earlier work this paper cites.
Learning Curves for Deep Neural Networks: A Gaussian Field Theory Perspective
Cohen, O., Malka, O., and Ringel, Z. (2019) · 1906
Earlier work this paper cites.
Finite depth and width corrections to the neural tangent kernel
Hanin, B. and Nica, M. (2019) · 1909
Earlier work this paper cites.
Dynamics of deep neural networks and neural tangent hierarchy
Huang, J. and Yau, H.-T. (2019) · 1909
Earlier work this paper cites.
Rank Correlation and Product-Moment Correlation
Moran, P. A. P. (1948) · 1948
Earlier work this paper cites.
Optimal storage properties of neural network models
Gardner, E. and Derrida, B. (1988) · 1988
Earlier work this paper cites.
The Fokker-Planck Equation: Methods of Solution and Applications
Risken, H. and Frank, T. (1996) · 1996
Earlier work this paper cites.
Mean-field analysis of two-layer neural networks: Non-asymptotic rates and generalization bounds
Chen, Z., Cao, Y., Gu, Q., and Zhang, T. (2020) · 2002
Earlier work this paper cites.
Tzen, B. and Raginsky, M. (2020) · 2002
Earlier work this paper cites.
Quantum Field Theory in a Nutshell
Zee, A. (2003) · 2003
Earlier work this paper cites.
Understanding gaussian process regression using the equivalent kernel
Sollich, P. and Williams, C. K. (2004) · 2004
Earlier work this paper cites.
Gaussian Processes for Machine Learning (Adaptive Computation and Machine Learning)
Rasmussen, C. E. and Williams, C. K. I. (2005) · 2005
Earlier work this paper cites.
Finite versus infinite neural networks: an empirical study
Lee, J., Schoenholz, S. S., Pennington, J., Adlam, B., Xiao, L., Novak, R., and Sohl-Dickstein, J. (2020) · 2007
Cited alongside, same era.
Kernel methods for deep learning
Cho, Y. and Saul, L. K. (2009) · 2009
Cited alongside, same era.
Mcmc using hamiltonian dynamics
Neal, R. M. et al · 2011
Cited alongside, same era.
Bayesian learning via stochastic gradient langevin dynamics
Welling, M. and Teh, Y. W. (2011) · 2011
Cited alongside, same era.
Stochastic gradient descent tricks
Bottou, L. (2012) · 2012
Cited alongside, same era.
Techniques and applications of path integration
Schulman, L. S. (2012) · 2012
Cited alongside, same era.
Deep neural networks as gaussian processes
Lee, J., Sohl-dickstein, J., Pennington, J., Novak, R., Schoenholz, S., and Bahri, Y. (2018) · 2018
Later among the works it cites.
Gaussian process behaviour in wide deep neural networks
Matthews, A. G. d. G., Rowland, M., Hron, J., Turner, R. E., and Ghahramani, Z. (2018) · 2018
Later among the works it cites.
A mean field view of the landscape of two-layer neural networks
Mei, S., Montanari, A., and Nguyen, P.-M. (2018) · 2018
Later among the works it cites.
Towards Understanding the Role of Over-Parametrization in Generalization of Neural Networks
Neyshabur, B., Li, Z., Bhojanapalli, S., LeCun, Y., and Srebro, N. (2018) · 2018
Later among the works it cites.
Bayesian Deep Convolutional Networks with Many Channels are Gaussian Processes
Novak, R., Xiao, L., Lee, J., Bahri, Y., Yang, G., Abolafia, D. A., Pennington, J., and Sohl-Dickstein, J. (2018) · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Toward Deeper Understanding of Neural Networks: The Power of Initialization and a Dual View on Expressivity
Daniely, A., Frostig, R., and Singer, Y. (2016) · 2016
Cited alongside, same era.
Consistency and fluctuations for stochastic gradient langevin dynamics
Teh, Y. W., Thiery, A. H., and Vollmer, S. J. (2016) · 2016
Cited alongside, same era.
Hoffer, E., Hubara, I., and Soudry, D. (2017) · 2017
Cited alongside, same era.
Stochastic Gradient Descent as Approximate Bayesian Inference
Mandt, S., Hoffman, M. D., and Blei, D. M. (2017) · 2017
Cited alongside, same era.
Tensor Methods in Statistics
Mccullagh, P. (2017) · 2017
Cited alongside, same era.
On the use of the edgeworth expansion in cosmology i: how to foresee and evade its pitfalls
Sellentin, E., Jaffe, A. H., and Heavens, A. F. (2017) · 2017
Cited alongside, same era.
Later among the works it cites.
On the Spectral Bias of Neural Networks
Rahaman, N., Baratin, A., Arpit, D., Draxler, F., Lin, M., Hamprecht, F. A., Bengio, Y., and Courville, A. (2018) · 2018
Later among the works it cites.
On lazy training in differentiable programming
Chizat, L., Oyallon, E., and Bach, F. (2019) · 2019
Later among the works it cites.
Asymptotics of wide networks from feynman diagrams
Dyer, E. and Gur-Ari, G. (2020) · 2020
Closest in time.
Double trouble in double descent: Bias and variance (s) in the lazy regime
d’Ascoli, S., Refinetti, M., Biroli, G., and Krzakala, F. (2020) · 2020
Closest in time.
Scaling description of generalization with number of parameters in deep learning
Geiger, M., Jacot, A., Spigler, S., Gabriel, F., Sagun, L., d’Ascoli, S., Biroli, G., Hongler, C., and Wyart, M. (2020) · 2020
Closest in time.
The large learning rate phase of deep learning: the catapult mechanism
Lewkowycz, A., Bahri, Y., Dyer, E., Sohl-Dickstein, J., and Gur-Ari, G. (2020) · 2020
Closest in time.
Non-gaussian processes and neural networks at finite widths
Yaida, S. (2020) · 2020
Closest in time.
Landscape and training regimes in deep learning
Geiger, M., Petrini, L., and Wyart, M. (2021) · 2021
Closest in time.
Is sgd a bayesian sampler? well, almost
Mingard, C., Valle-Pérez, G., Skalse, J., and Louis, A. A. (2021) · 2021
Closest in time.
On the origin of implicit regularization in stochastic gradient descent
Smith, S. L., Dherin, B., Barrett, D. G., and De, S. (2021) · 2021
Closest in time.