Fetching the paper…
Reading the bibliography…
We analyze feature learning in infinite-width neural networks trained with gradient flow through a self-consistent dynamical field theory.
Calculation of partition functions
John Hubbard · 1959
Earlier work this paper cites.
A bound for the error in the normal approximation to the distribution of a sum of dependent random variables
Charles Stein · 1972
Earlier work this paper cites.
Statistical dynamics of classical systems
Paul Cecil Martin, ED Siggia, and HA Rose · 1973
Earlier work this paper cites.
Dynamics as a substitute for replicas in systems with quenched random impurities
C De Dominicis · 1978
Earlier work this paper cites.
Dynamic theory of the spin-glass phase
Haim Sompolinsky and Annette Zippelius · 1981
Earlier work this paper cites.
Relaxational dynamics of the edwards-anderson model and the mean-field theory of spin-glasses
Haim Sompolinsky and Annette Zippelius · 1982
Earlier work this paper cites.
A method for solving the convex programming problem with convergence rate o (1/kˆ 2)
Yurii E Nesterov · 1983
Earlier work this paper cites.
Handbook of stochastic methods
Crispin W Gardiner et al · 1985
Earlier work this paper cites.
Chaos in random neural networks
Haim Sompolinsky, Andrea Crisanti, and Hans-Jurgen Sommers · 1988
Earlier work this paper cites.
Suppressing chaos in neural networks by noise
Lutz Molgedey, J Schuchhardt, and Heinz G Schuster · 1992
Earlier work this paper cites.
Large deviations for langevin spin glass dynamics
G Ben Arous and Alice Guionnet · 1995
Earlier work this paper cites.
Symmetric langevin spin glass dynamics
G Ben Arous and Alice Guionnet · 1997
Earlier work this paper cites.
Dynamics of batch learning in multilayer neural networks
Kenji Fukumizu · 1998
Earlier work this paper cites.
Advanced mathematical methods for scientists and engineers I: Asymptotic methods and perturbation theory
Carl M Bender and Steven Orszag · 1999
Earlier work this paper cites.
Cugliandolo-kurchan equations for dynamics of spin-glasses
Gérard Ben Arous, Amir Dembo, and Alice Guionnet · 2006
Earlier work this paper cites.
Cubic regularization of newton method and its global performance
Yurii Nesterov and Boris T Polyak · 2006
Earlier work this paper cites.
Random recurrent neural networks dynamics
M Samuelides and Bruno Cessac · 2007
Earlier work this paper cites.
Statistical physics of fields
Mehran Kardar · 2007
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky, Geoffrey Hinton, et al · 2009
Earlier work this paper cites.
Stimulus-dependent suppression of chaos in recurrent neural networks
Kanaka Rajan, LF Abbott, and Haim Sompolinsky · 2010
Earlier work this paper cites.
Ito and stratonovich calculuses in stochastic field theory
Juha Honkonen · 2011
Earlier work this paper cites.
Bayesian learning for neural networks
Radford M Neal · 2012
Earlier work this paper cites.
Exact solutions to the nonlinear dynamics of learning in deep linear neural networks
Andrew M Saxe, James L McClelland, and Surya Ganguli · 2013
Earlier work this paper cites.
Deep learning
Yann LeCun, Yoshua Bengio, and Geoffrey Hinton · 2015
Earlier work this paper cites.
Deep learning
Ian Goodfellow, Yoshua Bengio, and Aaron Courville · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Exponential expressivity in deep neural networks through transient chaos
Ben Poole, Subhaneil Lahiri, Maithra Raghu, Jascha Sohl-Dickstein, and Surya Ganguli · 2016
Earlier work this paper cites.
Deep information propagation
Samuel S Schoenholz, Justin Gilmer, Surya Ganguli, and Jascha Sohl-Dickstein · 2017
Earlier work this paper cites.
Mean field residual networks: On the edge of chaos
Greg Yang and Samuel Schoenholz · 2017
Earlier work this paper cites.
Why momentum really works
Gabriel Goh · 2017
Earlier work this paper cites.
Deep neural networks as gaussian processes
Jaehoon Lee, Jascha Sohl-dickstein, Jeffrey Pennington, Roman Novak, Sam Schoenholz, and Yasaman Bahri · 2018
Earlier work this paper cites.
Gaussian process behaviour in wide deep neural networks
Alexander G. de G. Matthews, Jiri Hron, Mark Rowland, Richard E. Turner, and Zoubin Ghahramani · 2018
Earlier work this paper cites.
Neural tangent kernel: Convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel, and Clement Hongler · 2018
Cited alongside, same era.
A mean field view of the landscape of two-layer neural networks
Song Mei, Andrea Montanari, and Phan-Minh Nguyen · 2018
Cited alongside, same era.
On the global convergence of gradient descent for over-parameterized models using optimal transport
Lenaic Chizat and Francis Bach · 2018
Cited alongside, same era.
Trainability and accuracy of neural networks: An interacting particle system approach
Grant M Rotskoff and Eric Vanden-Eijnden · 2018
Cited alongside, same era.
Path integral approach to random neural networks
A Crisanti and H Sompolinsky · 2018
Cited alongside, same era.
Out-of-equilibrium dynamical mean-field equations for the perceptron model
Dynamical mean-field theory for stochastic gradient descent in gaussian mixture classification
Francesca Mignacco, Florent Krzakala, Pierfrancesco Urbani, and Lenka Zdeborová · 2020
Later among the works it cites.
Numerical solution of the dynamical mean field theory of infinite-dimensional equilibrium liquids
Alessandro Manacorda, Grégory Schehr, and Francesco Zamponi · 2020
Later among the works it cites.
High-dimensional dynamics of generalization error in neural networks
Madhu S Advani, Andrew M Saxe, and Haim Sompolinsky · 2020
Later among the works it cites.
On the training dynamics of deep networks with l _ 2 l\_2 regularization
Aitor Lewkowycz and Guy Gur-Ari · 2020
Later among the works it cites.
Tensor programs iv: Feature learning in infinite-width neural networks
Greg Yang and Edward J Hu · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Elisabeth Agoritsas, Giulio Biroli, Pierfrancesco Urbani, and Francesco Zamponi · 2018
Cited alongside, same era.
JAX: composable transformations of Python+NumPy programs, 2018
James Bradbury, Roy Frostig, Peter Hawkins, Matthew James Johnson, Chris Leary, Dougal Maclaurin, George Necula, Adam Paszke, Jake VanderPlas, Skye Wanderman-Milne, and Qiao Zhang · 2018
Cited alongside, same era.
Bayesian deep convolutional networks with many channels are gaussian processes
Roman Novak, Lechao Xiao, Jaehoon Lee, Yasaman Bahri, Greg Yang, Jiri Hron, Daniel A Abolafia, Jeffrey Pennington, and Jascha Sohl-Dickstein · 2018
Cited alongside, same era.
Wide neural networks of any depth evolve as linear models under gradient descent
Jaehoon Lee, Lechao Xiao, Samuel Schoenholz, Yasaman Bahri, Roman Novak, Jascha Sohl-Dickstein, and Jeffrey Pennington · 2019
Cited alongside, same era.
On exact computation with an infinitely wide neural net
Sanjeev Arora, Simon S Du, Wei Hu, Zhiyuan Li, Russ R Salakhutdinov, and Ruosong Wang · 2019
Cited alongside, same era.
Gradient descent finds global minima of deep neural networks
Simon Du, Jason Lee, Haochuan Li, Liwei Wang, and Xiyu Zhai · 2019
Cited alongside, same era.
A convergence theory for deep learning via over-parameterization
Zeyuan Allen-Zhu, Yuanzhi Li, and Zhao Song · 2019
Cited alongside, same era.
Abdulkadir Canatar, Blake Bordelon, and Cengiz Pehlevan · 2021
Later among the works it cites.
Learning curves for overparametrized deep neural networks: A field theory perspective
Omry Cohen, Or Malka, and Zohar Ringel · 2021
Later among the works it cites.
Learning curves of generic features maps for realistic datasets with a teacher-student model
Bruno Loureiro, Cedric Gerbelot, Hugo Cui, Sebastian Goldt, Florent Krzakala, Marc Mezard, and Lenka Zdeborova · 2021
Later among the works it cites.
Neural tangent kernel eigenvalues accurately predict generalization
James B Simon, Madeline Dickens, and Michael R DeWeese · 2021
Later among the works it cites.
Asymptotics of representation learning in finite bayesian neural networks
Jacob Zavatone-Veth, Abdulkadir Canatar, Ben Ruben, and Cengiz Pehlevan · 2021
Later among the works it cites.
Predicting the outputs of finite deep neural networks trained with noisy gradients
Gadi Naveh, Oded Ben David, Haim Sompolinsky, and Zohar Ringel · 2021
Later among the works it cites.
The principles of deep learning theory
Daniel A Roberts, Sho Yaida, and Boris Hanin · 2021
Later among the works it cites.
Unified field theory for deep and recurrent neural networks, 2021
Kai Segadlo, Bastian Epping, Alexander van Meegen, David Dahmen, Michael Krämer, and Moritz Helias · 2021
Later among the works it cites.
A self consistent theory of gaussian processes captures feature learning effects in finite cnns
Gadi Naveh and Zohar Ringel · 2021
Later among the works it cites.
Separation of scales and a thermodynamic description of feature learning in some cnns
Inbar Seroussi and Zohar Ringel · 2021
Later among the works it cites.
Statistical mechanics of deep linear neural networks: The backpropagating kernel renormalization
Qianyi Li and Haim Sompolinsky · 2021
Later among the works it cites.
Depth induces scale-averaging in overparameterized linear bayesian neural networks
Jacob A Zavatone-Veth and Cengiz Pehlevan · 2021
Later among the works it cites.
Modeling from features: a mean-field framework for over-parameterized deep neural networks
Cong Fang, Jason Lee, Pengkun Yang, and Tong Zhang · 2021
Later among the works it cites.
Tuning large neural networks via zero-shot hyperparameter transfer
Greg Yang, Edward Hu, Igor Babuschkin, Szymon Sidor, Xiaodong Liu, David Farhi, Nick Ryder, Jakub Pachocki, Weizhu Chen, and Jianfeng Gao · 2021
Later among the works it cites.
Stochasticity helps to navigate rough landscapes: comparing gradient-descent-based algorithms in the phase retrieval problem
Francesca Mignacco, Pierfrancesco Urbani, and Lenka Zdeborová · 2021
Later among the works it cites.
The high-dimensional asymptotics of first order methods with random data
Michael Celentano, Chen Cheng, and Andrea Montanari · 2021
Later among the works it cites.
The effective noise of stochastic gradient descent
Francesca Mignacco and Pierfrancesco Urbani · 2021
Later among the works it cites.
Deep linear networks dynamics: Low-rank biases induced by initialization scale and l2 regularization
Arthur Jacot, François Ged, Franck Gabriel, Berfin Şimşek, and Clément Hongler · 2021
Later among the works it cites.
Tensor programs iib: Architectural universality of neural tangent kernel training dynamics
Greg Yang and Etai Littwin · 2021
Later among the works it cites.
A theory of neural tangent kernel alignment and its influence on training, 2021
Haozhe Shan and Blake Bordelon · 2021
Later among the works it cites.
A theory of representation learning in deep neural networks gives a deep generalisation of kernel methods, 2021
Adam X. Yang, Maxime Robeyns, Edward Milsom, Nandi Schoots, and Laurence Aitchison · 2021
Later among the works it cites.
Optimization with momentum: Dynamical, control-theoretic, and symplectic perspectives
Michael Muehlebach and Michael I Jordan · 2021
Later among the works it cites.
Correlation functions in random fully connected neural networks at finite width
Boris Hanin · 2022
Closest in time.
Contrasting random and learned features in deep bayesian linear regression
Jacob A Zavatone-Veth, William L Tong, and Cengiz Pehlevan · 2022
Closest in time.
Efficient computation of deep nonlinear infinite-width neural networks that learn features
Greg Yang, Michael Santacroce, and Edward J Hu · 2022
Closest in time.
Neural networks as kernel learners: The silent alignment effect
Alexander Atanasov, Blake Bordelon, and Cengiz Pehlevan · 2022
Closest in time.
Optimal learning rate schedules in high-dimensional non-convex optimization problems, 2022
Stéphane d’Ascoli, Maria Refinetti, and Giulio Biroli · 2022
Closest in time.