Fetching the paper…
Reading the bibliography…
In this work, we analyze various scaling limits of the training dynamics of transformer models in the feature learning regime.
Statistical dynamics of classical systems
Paul Cecil Martin, ED Siggia, and HA Rose · 1973
Earlier work this paper cites.
Dynamic theory of the spin-glass phase
Haim Sompolinsky and Annette Zippelius · 1981
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
On the global convergence of gradient descent for over-parameterized models using optimal transport
Lenaic Chizat and Francis Bach · 2018
Earlier work this paper cites.
Neural tangent kernel: Convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel, and Clément Hongler · 2018
Earlier work this paper cites.
Mean-field theory of two-layers neural networks: dimension-free bounds and kernel limit
Song Mei, Theodor Misiakiewicz, and Andrea Montanari · 2019
Earlier work this paper cites.
On lazy training in differentiable programming
Lenaic Chizat, Edouard Oyallon, and Francis Bach · 2019
Earlier work this paper cites.
Passed & spurious: Descent algorithms and local minima in spiked matrix-tensor models
Stefano Sarao Mannelli, Florent Krzakala, Pierfrancesco Urbani, and Lenka Zdeborova · 2019
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al · 2020
Earlier work this paper cites.
Scaling laws for neural language models
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei · 2020
Earlier work this paper cites.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Earlier work this paper cites.
Infinite attention: Nngp and ntk for deep attention networks
Jiri Hron, Yasaman Bahri, Jascha Sohl-Dickstein, and Roman Novak · 2020
Earlier work this paper cites.
On the distance between two neural networks and the stability of learning
Jeremy Bernstein, Arash Vahdat, Yisong Yue, and Ming-Yu Liu · 2020
Earlier work this paper cites.
Statistical field theory for neural networks , volume 970
Moritz Helias and David Dahmen · 2020
Earlier work this paper cites.
Dynamical mean-field theory for stochastic gradient descent in gaussian mixture classification
Francesca Mignacco, Florent Krzakala, Pierfrancesco Urbani, and Lenka Zdeborová · 2020
Cited alongside, same era.
The deep bootstrap framework: Good online learners are good offline generalizers
Preetum Nakkiran, Behnam Neyshabur, and Hanie Sedghi · 2020
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu · 2020
Cited alongside, same era.
Swin transformer: Hierarchical vision transformer using shifted windows
Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo · 2021
Cited alongside, same era.
Tuning large neural networks via zero-shot hyperparameter transfer
Ge Yang, Edward Hu, Igor Babuschkin, Szymon Sidor, Xiaodong Liu, David Farhi, Nick Ryder, Jakub Pachocki, Weizhu Chen, and Jianfeng Gao · 2021
The neural covariance sde: Shaped infinite depth-and-width networks at initialization
Mufan Li, Mihai Nica, and Dan Roy · 2022
Later among the works it cites.
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al · 2023
Later among the works it cites.
Feature-learning networks are consistent across widths at realistic scales, 2023
Nikhil Vyas, Alexander Atanasov, Blake Bordelon, Depen Morwani, Sabarish Sainathan, and Cengiz Pehlevan · 2023
Later among the works it cites.
Dynamics of finite width kernel and prediction fluctuations in mean field neural networks
Blake Bordelon and Cengiz Pehlevan · 2023
Later among the works it cites.
Effective theory of transformers at initialization, 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Tensor programs iv: Feature learning in infinite-width neural networks
Greg Yang and Edward J Hu · 2021
Cited alongside, same era.
Attention is not all you need: Pure attention loses rank doubly exponentially with depth
Yihe Dong, Jean-Baptiste Cordonnier, and Andreas Loukas · 2021
Cited alongside, same era.
A survey on vision transformer
Kai Han, Yunhe Wang, Hanting Chen, Xinghao Chen, Jianyuan Guo, Zhenhua Liu, Yehui Tang, An Xiao, Chunjing Xu, Yixing Xu, et al · 2022
Cited alongside, same era.
Scenic: A jax library for computer vision research and beyond
Mostafa Dehghani, Alexey Gritsenko, Anurag Arnab, Matthias Minderer, and Yi Tay · 2022
Cited alongside, same era.
Training compute-optimal large language models
Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, et al · 2022
Cited alongside, same era.
Signal propagation in transformers: Theoretical perspectives and the role of rank collapse
Lorenzo Noci, Sotiris Anagnostidis, Luca Biggio, Antonio Orvieto, Sidak Pal Singh, and Aurelien Lucchi · 2022
Cited alongside, same era.
Rigorous dynamical mean field theory for stochastic gradient descent methods
Cedric Gerbelot, Emanuele Troiani, Francesca Mignacco, Florent Krzakala, and Lenka Zdeborova · 2022
Cited alongside, same era.
Emily Dinan, Sho Yaida, and Susan Zhang · 2023
Later among the works it cites.
Simplifying transformer blocks
Bobby He and Thomas Hofmann · 2023
Later among the works it cites.
On the infinite-depth limit of finite-width neural networks
Soufiane Hayou · 2023
Later among the works it cites.
Neural signature kernels as infinite-width-depth-limits of controlled resnets
Nicola Muca Cirone, Maud Lemercier, and Cristopher Salvi · 2023
Later among the works it cites.
Automatic gradient descent: Deep learning without hyperparameters
Jeremy Bernstein, Chris Mingard, Kevin Huang, Navid Azizan, and Yisong Yue · 2023
Later among the works it cites.
Geometric dynamics of signal propagation predict trainability of transformers, 2024
Aditya Cowsik, Tamra Nebabu, Xiao-Liang Qi, and Surya Ganguli · 2024
Closest in time.
The shaped transformer: Attention models in the infinite depth-and-width limit
Lorenzo Noci, Chuning Li, Mufan Li, Bobby He, Thomas Hofmann, Chris J Maddison, and Dan Roy · 2024
Closest in time.
Lénaïc Chizat and Praneeth Netrapalli · 2024
Closest in time.
Getting vit in shape: Scaling laws for compute-optimal model design
Ibrahim M Alabdulmohsin, Xiaohua Zhai, Alexander Kolesnikov, and Lucas Beyer · 2024
Closest in time.