Fetching the paper…
Reading the bibliography…
Current theoretical results on optimization trajectories of neural networks trained by gradient descent typically have the form of rigorous but potentially loose bounds on the loss values.
Asymptotic behavior of the spectrum of weakly polar integral operators
Birman, M. Š. and Solomjak, M. Z · 1970
Earlier work this paper cites.
Dynamically stable infinite-width limits of neural classifiers
Golikov, E. A · 2006
Earlier work this paper cites.
Kernel methods for deep learning
Cho, Y. and Saul, L · 2009
Earlier work this paper cites.
The loss surfaces of multilayer networks
Choromanska, A., Henaff, M., Mathieu, M., Arous, G. B., and LeCun, Y · 2015
Earlier work this paper cites.
Exponential expressivity in deep neural networks through transient chaos
Poole, B., Lahiri, S., Raghu, M., Sohl-Dickstein, J., and Ganguli, S · 2016
Earlier work this paper cites.
Geometry of neural network loss surfaces via random matrix theory
Pennington, J. and Bahri, Y · 2017
Earlier work this paper cites.
On the global convergence of gradient descent for over-parameterized models using optimal transport
Chizat, L. and Bach, F · 2018
Earlier work this paper cites.
Neural tangent kernel: Convergence and generalization in neural networks
Jacot, A., Gabriel, F., and Hongler, C · 2018
Earlier work this paper cites.
A mean field view of the landscape of two-layer neural networks
Mei, S., Montanari, A., and Nguyen, P.-M · 2018
Earlier work this paper cites.
Trainability and accuracy of neural networks: An interacting particle system approach
Rotskoff, G. M. and Vanden-Eijnden, E · 2018
Earlier work this paper cites.
Collective evolution of weights in wide neural networks
Yarotsky, D · 2018
Cited alongside, same era.
Fine-grained analysis of optimization and generalization for overparameterized two-layer neural networks
Arora, S., Du, S., Hu, W., Li, Z., and Wang, R · 2019
Cited alongside, same era.
On lazy training in differentiable programming
Chizat, L., Oyallon, E., and Bach, F · 2019
Cited alongside, same era.
On the impact of the activation function on deep neural networks training
Hayou, S., Doucet, A., and Rousseau, J · 2019
Cited alongside, same era.
Wide neural networks of any depth evolve as linear models under gradient descent
Lee, J., Xiao, L., Schoenholz, S., Bahri, Y., Novak, R., Sohl-Dickstein, J., and Pennington, J · 2019
Cited alongside, same era.
The surprising simplicity of the early-time learning dynamics of neural networks
Hu, W., Xiao, L., Adlam, B., and Pennington, J · 2020
Later among the works it cites.
The large learning rate phase of deep learning: the catapult mechanism
Lewkowycz, A., Bahri, Y., Dyer, E., Sohl-Dickstein, J., and Gur-Ari, G · 2020
Later among the works it cites.
On the linearity of large non-linear models: when and why the tangent kernel is constant
Liu, C., Zhu, L., and Belkin, M · 2020
Later among the works it cites.
Disentangling trainability and generalization in deep neural networks
Xiao, L., Pennington, J., and Schoenholz, S · 2020
Later among the works it cites.
A fine-grained spectral perspective on neural networks
Yang, G. and Salman, H · 2020
Later among the works it cites.
A deep conditioning treatment of neural networks
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Su, L. and Yang, P · 2019
Cited alongside, same era.
Gradient dynamics of shallow univariate relu networks
Williams, F., Trager, M., Silva, C., Panozzo, D., Zorin, D., and Bruna, J · 2019
Cited alongside, same era.
The neural tangent kernel in high dimensions: Triple descent and a multi-scale theory of generalization
Adlam, B. and Pennington, J · 2020
Cited alongside, same era.
Towards understanding the spectral bias of deep learning
Cao, Y., Fang, Z., Wu, Y., Zhou, D.-X., and Gu, Q · 2020
Cited alongside, same era.
Spectra of the conjugate kernel and neural tangent kernel for linear-width neural networks
Fan, Z. and Wang, Z · 2020
Cited alongside, same era.
Towards a general theory of infinite-width limits of neural classifiers
Golikov, E
Cited in the paper.
Agarwal, N., Awasthi, P., and Kale, S · 2021
Closest in time.
Explaining neural scaling laws
Bahri, Y., Dyer, E., Kaplan, J., Lee, J., and Sharma, U · 2021
Closest in time.
Canatar, A., Bordelon, B., and Pehlevan, C · 2021
Closest in time.
Mean-field behaviour of neural tangent kernel for deep neural networks
Hayou, S., Doucet, A., and Rousseau, J · 2021
Closest in time.
Optimal rates for averaged stochastic gradient descent under neural tangent kernel regime
Nitanda, A. and Suzuki, T · 2021
Closest in time.