Fetching the paper…
Reading the bibliography…
A recent goal in the theory of deep learning is to identify how neural networks can escape the "lazy training," or Neural Tangent Kernel (NTK) regime, where the network is coupled with its first order Taylor expansion at initialization.
Spectral methods for data science: A statistical perspective
Yuxin Chen, Yuejie Chi, Jianqing Fan, and Cong Ma · 1935
Earlier work this paper cites.
The rotation of eigenvectors by a perturbation. iii
Chandler Davis and W. M. Kahan · 1970
Earlier work this paper cites.
Backward feature correction: How deep learning performs deep learning
Zeyuan Allen-Zhu and Yuanzhi Li · 2001
Earlier work this paper cites.
Taylorized training: Towards better approximation of neural network training at finite width, 2020
Yu Bai, Ben Krause, Huan Wang, Caiming Xiong, and Richard Socher · 2002
Earlier work this paper cites.
Andrea Montanari and Yiqiao Zhong · 2007
Earlier work this paper cites.
Matrix completion has no spurious local minimum
Rong Ge, Jason D. Lee, and Tengyu Ma · 2016
Earlier work this paper cites.
In 29th Annual Conference on Learning Theory , Proceedings of Machine Learning Research, pages 1246–1257, 2016
Jason D. Lee, Max Simchowitz, Michael I. Jordan, and Benjamin Recht · 2016
Earlier work this paper cites.
Learning one-hidden-layer neural networks with landscape design
Rong Ge, Jason D. Lee, and Tengyu Ma · 2017
Earlier work this paper cites.
How to escape saddle points efficiently
Chi Jin, Rong Ge, Praneeth Netrapalli, Sham M. Kakade, and Michael I. Jordan · 2017
Earlier work this paper cites.
JAX: composable transformations of Python+NumPy programs, 2018
James Bradbury, Roy Frostig, Peter Hawkins, Matthew James Johnson, Chris Leary, Dougal Maclaurin, George Necula, Adam Paszke, Jake VanderPlas, Skye Wanderman-Milne, and Qiao Zhang · 2018
Earlier work this paper cites.
On the power of over-parametrization in neural networks with quadratic activation
Simon S. Du and Jason D. Lee · 2018
Earlier work this paper cites.
Neural tangent kernel: Convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel, and Clément Hongler · 2018
Earlier work this paper cites.
Learning overparameterized neural networks via stochastic gradient descent on structured data
Yuanzhi Li and Yingyu Liang · 2018
Earlier work this paper cites.
Theoretical insights into the optimization landscape of over-parameterized shallow neural networks
Mahdi Soltanolkotabi, Adel Javanmard, and Jason D. Lee · 2018
Cited alongside, same era.
High-dimensional probability: An introduction with applications in data science , volume 47
Roman Vershynin · 2018
Cited alongside, same era.
Colin Wei, Jason D. Lee, Qiang Liu, and Tengyu Ma · 2018
Cited alongside, same era.
Stochastic gradient descent optimizes over-parameterized deep relu networks, 2018
Difan Zou, Yuan Cao, Dongruo Zhou, and Quanquan Gu · 2018
Cited alongside, same era.
What can resnet learn efficiently, going beyond kernels?
Zeyuan Allen-Zhu and Yuanzhi Li · 2019
Cited alongside, same era.
When do neural networks outperform kernel methods?
Behrooz Ghorbani, Song Mei, Theodor Misiakiewicz, and Andrea Montanari · 2020
Later among the works it cites.
The surprising simplicity of the early-time learning dynamics of neural networks
Wei Hu, Lechao Xiao, Ben Adlam, and Jeffrey Pennington · 2020
Later among the works it cites.
Dynamics of deep neural networks and neural tangent hierarchy
Jiaoyang Huang and Horng-Tzer Yau · 2020
Later among the works it cites.
Finite versus infinite neural networks: an empirical study
Jaehoon Lee, Samuel Schoenholz, Jeffrey Pennington, Ben Adlam, Lechao Xiao, Roman Novak, and Jascha Sohl-Dickstein · 2020
Later among the works it cites.
Learning over-parametrized two-layer neural networks beyond ntk
Yuanzhi Li, Tengyu Ma, and Hongyang R. Zhang · 2020
Later among the works it cites.
Neural tangents: Fast and easy infinite neural networks in python
Roman Novak, Lechao Xiao, Jiri Hron, Jaehoon Lee, Alexander A. Alemi, Jascha Sohl-Dickstein, and Samuel S. Schoenholz · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Generalization bounds of stochastic gradient descent for wide and deep neural networks
Yuan Cao and Quanquan Gu · 2019
Cited alongside, same era.
On lazy training in differentiable programming
Lénaïc Chizat, Edouard Oyallon, and Francis Bach · 2019
Cited alongside, same era.
Limitations of lazy training of two-layers neural network
Behrooz Ghorbani, Song Mei, Theodor Misiakiewicz, and Andrea Montanari · 2019
Cited alongside, same era.
Wide neural networks of any depth evolve as linear models under gradient descent
Jaehoon Lee, Lechao Xiao, Samuel Schoenholz, Yasaman Bahri, Roman Novak, Jascha Sohl-Dickstein, and Jeffrey Pennington · 2019
Cited alongside, same era.
High-dimensional statistics: A non-asymptotic viewpoint , volume 48
Martin J. Wainwright · 2019
Cited alongside, same era.
On the power and limitations of random features for understanding neural networks
Gilad Yehudai and Ohad Shamir · 2019
Cited alongside, same era.
Beyond linearization: On quadratic and higher-order approximation of wide neural networks
Yu Bai and Jason D. Lee · 2020
Cited alongside, same era.
Later among the works it cites.
Kernel and rich regimes in overparametrized models
Blake Woodworth, Suriya Gunasekar, Jason D. Lee, Edward Moroshko, Pedro Savarese, Itay Golan, Daniel Soudry, and Nathan Srebro · 2020
Later among the works it cites.
Deep learning: A statistical viewpoint
Peter L. Bartlett, Andrea Montanari, and Alexander Rakhlin · 2021
Later among the works it cites.
Linearized two-layers neural networks in high dimension
Behrooz Ghorbani, Song Mei, Theodor Misiakiewicz, and Andrea Montanari · 2021
Later among the works it cites.
On nonconvex optimization for machine learning: Gradients, stochasticity, and saddle points
Chi Jin, Praneeth Netrapalli, Rong Ge, Sham M. Kakade, and Michael I. Jordan · 2021
Later among the works it cites.
Quantifying the benefit of using differentiable learning over tangent kernels
Eran Malach, Pritish Kamath, Emmanuel Abbe, and Nathan Srebro · 2021
Later among the works it cites.
The three stages of learning dynamics in high-dimensional kernel methods
Nikhil Ghosh, Song Mei, and Bin Yu · 2022
Closest in time.
Generalization error of random features and kernel methods: hypercontractivity and kernel matrix concentration
Song Mei, Theodor Misiakiewicz, and Andrea Montanari · 2022
Closest in time.