Fetching the paper…
Reading the bibliography…
Going beyond stochastic gradient descent (SGD), what new phenomena emerge in wide neural networks trained by adaptive optimizers like Adam? Here we show: The same dichotomy between feature learning and kernel behaviors (as in SGD) holds for general optimizers as well, including Adam -- albeit with a nonlinear notion of "kernel." We derive the corresponding "neural tangent" and "maximal update" limits for any architecture.
Greg Yang · 1902
Earlier work this paper cites.
Large Batch Optimization for Deep Learning: Training BERT in 76 minutes
Yang You, Jing Li, Sashank Reddi, Jonathan Hseu, Sanjiv Kumar, Srinadh Bhojanapalli, Xiaodan Song, James Demmel, Kurt Keutzer, and Cho-Jui Hsieh · 1904
Earlier work this paper cites.
Greg Yang · 1910
Earlier work this paper cites.
Bayesian learning for neural networks
Geoffrey E. Hinton and R. Neal · 1995
Earlier work this paper cites.
On the distance between two neural networks and the stability of learning
Jeremy Bernstein, Arash Vahdat, Yisong Yue, and Ming-Yu Liu · 2002
Earlier work this paper cites.
Tensor programs ii: Neural tangent kernel for any architecture
Greg Yang · 2006
Earlier work this paper cites.
Tensor programs iii: Neural matrix laws
Greg Yang · 2009
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
John C. Duchi, Elad Hazan, and Yoram Singer · 2010
Earlier work this paper cites.
Topics in random matrix theory
Terence Tao · 2012
Earlier work this paper cites.
Upper and lower bounds for stochastic processes: modern methods and classical problems
Michel Talagrand · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Incorporating nesterov momentum into adam
Timothy Dozat · 2016
Earlier work this paper cites.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter · 2017
Earlier work this paper cites.
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, L. Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Large Batch Training of Convolutional Networks
Yang You, Igor Gitman, and Boris Ginsburg · 2017
Earlier work this paper cites.
signSGD: Compressed Optimisation for Non-Convex Problems, February 2018
Jeremy Bernstein, Yu-Xiang Wang, Kamyar Azizzadenesheli, and Anima Anandkumar · 2018
Earlier work this paper cites.
On the global convergence of gradient descent for over-parameterized models using optimal transport
Lénaïc Chizat and Francis R. Bach · 2018
Cited alongside, same era.
Neural tangent kernel: convergence and generalization in neural networks (invited paper)
Arthur Jacot, Franck Gabriel, and Clément Hongler · 2018
Cited alongside, same era.
Deep neural networks as gaussian processes
Jaehoon Lee, Y. Bahri, Roman Novak, S. Schoenholz, Jeffrey Pennington, and Jascha Sohl-Dickstein · 2018
Cited alongside, same era.
Gaussian process behaviour in wide deep neural networks
A. Matthews, M. Rowland, J. Hron, R. Turner, and Zoubin Ghahramani · 2018
Cited alongside, same era.
Infinite attention: Nngp and ntk for deep attention networks
Jiri Hron, Yasaman Bahri, Jascha Narain Sohl-Dickstein, and Roman Novak · 2020
Later among the works it cites.
Dynamics of deep neural networks and neural tangent hierarchy
Jiaoyang Huang and H. T. Yau · 2020
Later among the works it cites.
Improving transformer optimization through better initialization
Xiaoshan Huang, Felipe Pérez, Jimmy Ba, and Maksims Volkovs · 2020
Later among the works it cites.
Understanding the difficulty of training transformers
Liyuan Liu, Xiaodong Liu, Jianfeng Gao, Weizhu Chen, and Jiawei Han · 2020
Later among the works it cites.
A rigorous framework for the mean field limit of multilayer neural networks
Phan-Minh Nguyen and Huy-Tuan Pham · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Grant M. Rotskoff and Eric Vanden-Eijnden · 2018
Cited alongside, same era.
Adafactor: Adaptive Learning Rates with Sublinear Memory Cost
Noam Shazeer and Mitchell Stern · 2018
Cited alongside, same era.
Adaptive learning rate methods
Zhiming Zhou, Qingru Zhang, Guansong Lu, Hongwei Wang, Weinan Zhang, and Yong Yu · 2018
Cited alongside, same era.
A convergence theory for deep learning via over-parameterization
Zeyuan Allen-Zhu, Yuanzhi Li, and Zhao Song · 2019
Cited alongside, same era.
On exact computation with an infinitely wide neural net
Sanjeev Arora, S. Du, Wei Hu, Zhiyuan Li, R. Salakhutdinov, and Ruosong Wang · 2019
Cited alongside, same era.
Layer rotation: a surprisingly powerful indicator of generalization in deep networks?
Simon Carbonnelle and Christophe De Vleeschouwer · 2019
Cited alongside, same era.
On lazy training in differentiable programming
Lénaïc Chizat, Edouard Oyallon, and Francis R. Bach · 2019
Cited alongside, same era.
Graph neural tangent kernel: Fusing graph neural networks with graph kernels
Simon Shaolei Du, Kangcheng Hou, Barnabás Póczos, Ruslan Salakhutdinov, Ruosong Wang, and Keyulu Xu · 2019
Cited alongside, same era.
Mean field analysis of neural networks: A law of large numbers
Justin A. Sirignano and Konstantinos Spiliopoulos · 2020
Later among the works it cites.
Feature learning in infinite-width neural networks
Greg Yang and Edward J. Hu · 2020
Later among the works it cites.
The recurrent neural tangent kernel
Sina Alemohammad, Zichao Wang, Randall Balestriero, and Richard Baraniuk · 2021
Later among the works it cites.
Etai Littwin, Omid Saremi, Shuangfei Zhai, Vimal Thilak, Hanlin Goh, Joshua M. Susskind, and Greg Yang · 2021
Later among the works it cites.
Learning by Turning: Neural Architecture Aware Optimisation
Yang Liu, Jeremy Bernstein, Markus Meister, and Yisong Yue · 2021
Later among the works it cites.
Tensor programs iib: Architectural universality of neural tangent kernel training dynamics
Greg Yang and Etai Littwin · 2021
Later among the works it cites.
Non-Gaussian Tensor Programs
Eugene Golikov and Greg Yang · 2022
Later among the works it cites.
A Kernel-Based View of Language Model Fine-Tuning, October 2022
Sadhika Malladi, Alexander Wettig, Dingli Yu, Danqi Chen, and Sanjeev Arora · 2022
Later among the works it cites.
Meta-Principled Family of Hyperparameter Scaling Strategies
Sho Yaida · 2022
Later among the works it cites.
Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer
Greg Yang, Edward J. Hu, Igor Babuschkin, Szymon Sidor, Xiaodong Liu, David Farhi, Nick Ryder, Jakub Pachocki, Weizhu Chen, and Jianfeng Gao · 2022
Later among the works it cites.
On the limit of the largest eigenvalue of the large dimensional sample covariance matrix
Y. Q. Yin, Z. D. Bai, and P. R. Krishnaiah · 2064
Closest in time.