Fetching the paper…
Reading the bibliography…
We show at a physics level of rigor that Bayesian inference with a fully connected neural network and a shaped nonlinearity of the form $\phi(t) = t + \psi t^3/L$ is (perturbatively) solvable in the regime where the number of training datapoints $P$ , the input dimension $N_0$, the network layer widths $N$, and the network depth $L$ are simultaneously large.
Bayesian interpolation
David JC MacKay · 1992
Earlier work this paper cites.
Priors for infinite networks
Radford M Neal · 1996
Earlier work this paper cites.
Gaussian processes for machine learning (gpml) toolbox
Carl Edward Rasmussen and Hannes Nickisch · 2010
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton · 2012
Earlier work this paper cites.
Dropout as a bayesian approximation: Representing model uncertainty in deep learning
Yarin Gal and Zoubin Ghahramani · 2016
Earlier work this paper cites.
Deep learning scaling is predictable, empirically
Joel Hestness, Sharan Narang, Newsha Ardalani, Gregory Diamos, Heewoo Jun, Hassan Kianinejad, Md Patwary, Mostofa Ali, Yang Yang, and Yanqi Zhou · 2017
Earlier work this paper cites.
A correspondence between random neural networks and statistical field theory
Samuel S Schoenholz, Jeffrey Pennington, and Jascha Sohl-Dickstein · 2017
Earlier work this paper cites.
Deep convolutional networks as shallow gaussian processes
Adrià Garriga-Alonso, Carl Edward Rasmussen, and Laurence Aitchison · 2018
Earlier work this paper cites.
Which neural net architectures give rise to exploding and vanishing gradients?
Boris Hanin · 2018
Earlier work this paper cites.
Deep neural networks as gaussian processes
Jaehoon Lee, Yasaman Bahri, Roman Novak, Samuel S Schoenholz, Jeffrey Pennington, and Jascha Sohl-Dickstein · 2018
Earlier work this paper cites.
Bayesian deep convolutional networks with many channels are gaussian processes
Roman Novak, Lechao Xiao, Jaehoon Lee, Yasaman Bahri, Greg Yang, Jiri Hron, Daniel A Abolafia, Jeffrey Pennington, and Jascha Sohl-Dickstein · 2018
Earlier work this paper cites.
A simple baseline for bayesian uncertainty in deep learning
Wesley J Maddox, Pavel Izmailov, Timur Garipov, Dmitry P Vetrov, and Andrew Gordon Wilson · 2019
Earlier work this paper cites.
Greg Yang · 2019
Earlier work this paper cites.
Greg Yang · 2019
Earlier work this paper cites.
Language models are few-shot learners
Tom B Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Earlier work this paper cites.
Bayesian deep ensembles via the neural tangent kernel
Bobby He, Balaji Lakshminarayanan, and Yee Whye Teh · 2020
Cited alongside, same era.
Products of many large random matrices and gradients in deep neural networks
Boris Hanin and Mihai Nica · 2020
Cited alongside, same era.
Scaling laws for neural language models
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei · 2020
Cited alongside, same era.
Bayesian deep learning and a probabilistic perspective of generalization
Andrew G Wilson and Pavel Izmailov · 2020
Cited alongside, same era.
Non-gaussian processes and neural networks at finite widths
Sho Yaida · 2020
Cited alongside, same era.
What are bayesian neural network posteriors really like?
Wide bayesian neural networks have a simple weight posterior: theory and accelerated sampling
Jiri Hron, Roman Novak, Jeffrey Pennington, and Jascha Sohl-Dickstein · 2022
Later among the works it cites.
The neural covariance sde: Shaped infinite depth-and-width networks at initialization
Mufan Bill Li, Mihai Nica, and Daniel M Roy · 2022
Later among the works it cites.
The Principles of Deep Learning Theory: An Effective Theory Approach to Understanding Neural Networks
Daniel A Roberts, Sho Yaida, and Boris Hanin · 2022
Later among the works it cites.
Deep learning without shortcuts: Shaping the kernel with tailored rectifiers
Guodong Zhang, Aleksandar Botev, and James Martens · 2022
Later among the works it cites.
Contrasting random and learned features in deep bayesian linear regression
Jacob A. Zavatone-Veth, William L. Tong, and Cengiz Pehlevan · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Pavel Izmailov, Sharad Vikram, Matthew D Hoffman, and Andrew Gordon Gordon Wilson · 2021
Cited alongside, same era.
Highly accurate protein structure prediction with alphafold
John Jumper, Richard Evans, Alexander Pritzel, Tim Green, Michael Figurnov, Olaf Ronneberger, Kathryn Tunyasuvunakool, Russ Bates, Augustin Žídek, Anna Potapenko, et al · 2021
Cited alongside, same era.
Statistical mechanics of deep linear neural networks: The backpropagating kernel renormalization
Qianyi Li and Haim Sompolinsky · 2021
Cited alongside, same era.
Precise characterization of the prior predictive distribution of deep relu networks
Lorenzo Noci, Gregor Bachmann, Kevin Roth, Sebastian Nowozin, and Thomas Hofmann · 2021
Cited alongside, same era.
A self consistent theory of gaussian processes captures feature learning effects in finite cnns
Gadi Naveh and Zohar Ringel · 2021
Cited alongside, same era.
Tensor programs ii: Neural tangent kernel for any architecture
Greg Yang · 2021
Cited alongside, same era.
Exact marginal prior distributions of finite bayesian neural networks
Jacob Zavatone-Veth and Cengiz Pehlevan · 2021
Cited alongside, same era.
R Aiudi, R Pacelli, A Vezzani, R Burioni, and P Rotondo · 2023
Later among the works it cites.
Structures of neural network effective theories
Ian Banta, Tianji Cai, Nathaniel Craig, and Zhengkang Zhang · 2023
Later among the works it cites.
Optimal learning of deep random networks of extensive-width
Hugo Cui, Florent Krzakala, and Lenka Zdeborová · 2023
Later among the works it cites.
Quantitative clts in deep neural networks
Stefano Favaro, Boris Hanin, Domenico Marinucci, Ivan Nourdin, and Giovanni Peccati · 2023
Later among the works it cites.
Random neural networks in the infinite width limit as gaussian processes
Boris Hanin · 2023
Later among the works it cites.
Bayesian interpolation with deep linear networks
Boris Hanin and Alexander Zlokapa · 2023
Later among the works it cites.
Separation of scales and a thermodynamic description of feature learning in some cnns
Inbar Seroussi, Gadi Naveh, and Zohar Ringel · 2023
Later among the works it cites.
Critical feature learning in deep neural networks
Kristen Fischer, Javed Lindner, David Dahmen, Zohar Ringel, Michael Kramer, and Mortiz Helias · 2024
Closest in time.
The shaped transformer: Attention models in the infinite depth-and-width limit
Lorenzo Noci, Chuning Li, Mufan Li, Bobby He, Thomas Hofmann, Chris J Maddison, and Dan Roy · 2024
Closest in time.
Deep neural network initialization with sparsity inducing activations
Ilan Price, Nicholas Daultry Ball, Samuel CH Lam, Adam C Jones, and Jared Tanner · 2024
Closest in time.