Dynamical Isometry and a Mean Field Theory of LSTMs and GRUs
Original
Dar Gilboa, Bo Chang, Minmin Chen, Greg Yang, Samuel S. Schoenholz, Ed H. Chi, and Jeffrey Pennington · 1901
Earlier work this paper cites.
Mean-field theory of two-layers neural networks: dimension-free bounds and kernel limit
Original
Song Mei, Theodor Misiakiewicz, and Andrea Montanari · 1902
Earlier work this paper cites.
Mean Field Limit of the Learning Dynamics of Multilayer Neural Networks
Original
Phan-Minh Nguyen · 1902
Earlier work this paper cites.
Scaling Limits of Wide Neural Networks with Weight Sharing: Gaussian Process Behavior, Gradient Independence, and Neural Tangent Kernel Derivation
Original
Greg Yang · 1902
Earlier work this paper cites.
A Mean Field Theory of Batch Normalization
Original
Greg Yang, Jeffrey Pennington, Vinay Rao, Jascha Sohl-Dickstein, and Samuel S. Schoenholz · 1902
Earlier work this paper cites.
Mean Field Analysis of Deep Neural Networks
Original
Justin Sirignano and Konstantinos Spiliopoulos · 1903
Earlier work this paper cites.
A mean-field limit for certain deep neural networks
Original
Dyego Araújo, Roberto I. Oliveira, and Daniel Yukimura · 1906
Earlier work this paper cites.
Wider Networks Learn Better Features
Original
Dar Gilboa and Guy Gur-Ari · 1909
Earlier work this paper cites.
Why bigger is not always better: on finite and infinite neural networks
Original
Laurence Aitchison · 1910
Earlier work this paper cites.
Tensor Programs I: Wide Feedforward or Recurrent Neural Networks of Any Architecture are Gaussian Processes
Original
Greg Yang · 1910
Earlier work this paper cites.
A Rigorous Framework for the Mean Field Limit of Multilayer Neural Networks
Original
Phan-Minh Nguyen and Huy Tuan Pham · 2001
Earlier work this paper cites.
On the infinite width limit of neural networks with a standard parameterization
Original
Jascha Sohl-Dickstein, Roman Novak, Samuel S. Schoenholz, and Jaehoon Lee · 2001
Earlier work this paper cites.
Implicit Bias of Gradient Descent for Wide Two-layer Neural Networks Trained with the Logistic Loss
Original
Lenaic Chizat and Francis Bach · 2002
Earlier work this paper cites.
Train Large, Then Compress: Rethinking Model Size for Efficient Training and Inference of Transformers
Original
Zhuohan Li, Eric Wallace, Sheng Shen, Kevin Lin, Kurt Keutzer, Dan Klein, and Joseph E. Gonzalez · 2002
Earlier work this paper cites.
Kernel and Rich Regimes in Overparametrized Models
Original
Blake Woodworth, Suriya Gunasekar, Jason D. Lee, Edward Moroshko, Pedro Savarese, Itay Golan, Daniel Soudry, and Nathan Srebro · 2002
Earlier work this paper cites.
The large learning rate phase of deep learning: the catapult mechanism
Original
Aitor Lewkowycz, Yasaman Bahri, Ethan Dyer, Jascha Sohl-Dickstein, and Guy Gur-Ari · 2003
Earlier work this paper cites.
Language Models are Few-Shot Learners
Original
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei · 2005
Earlier work this paper cites.
Language Models are Few-Shot Learners
Original
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei · 2005
Earlier work this paper cites.