Fetching the paper…
Reading the bibliography…
We study the optimization of wide neural networks (NNs) via gradient flow (GF) in setups that allow feature learning while admitting non-asymptotic global convergence guarantees.
A topological property of real analytic subsets
Stanislaw Lojasiewicz · 1963
Earlier work this paper cites.
Gradient methods for the minimisation of functionals
Boris T. Polyak · 1963
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
Xavier Glorot and Yoshua Bengio · 2010
Earlier work this paper cites.
Topology and geometry of half-rectified network optimization
C Daniel Freeman and Joan Bruna · 2016
Earlier work this paper cites.
Deep learning without poor local minima
Kenji Kawaguchi · 2016
Earlier work this paper cites.
No bad local minima: Data independent training error guarantees for multilayer neural networks
Daniel Soudry and Yair Carmon · 2016
Earlier work this paper cites.
Breaking the curse of dimensionality with convex neural networks
Francis Bach · 2017
Earlier work this paper cites.
Deep neural networks as gaussian processes
Jaehoon Lee, Yasaman Bahri, Roman Novak, Samuel S Schoenholz, Jeffrey Pennington, and Jascha Sohl-Dickstein · 2017
Earlier work this paper cites.
An analytical formula of population gradient for two-layered ReLU network and its applications in convergence and critical point analysis
Yuandong Tian · 2017
Earlier work this paper cites.
Gradient descent with identity initialization efficiently learns positive definite linear transformations by deep residual networks
Peter Bartlett, Dave Helmbold, and Philip Long · 2018
Earlier work this paper cites.
On the global convergence of gradient descent for over-parameterized models using optimal transport
Lenaic Chizat and Francis Bach · 2018
Earlier work this paper cites.
Gradient descent finds global minima of deep neural networks
Simon S Du, Jason D Lee, Haochuan Li, Liwei Wang, and Xiyu Zhai · 2018
Earlier work this paper cites.
Neural tangent kernel: Convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel, and Clément Hongler · 2018
Earlier work this paper cites.
Learning overparameterized neural networks via stochastic gradient descent on structured data
Yuanzhi Li and Yingyu Liang · 2018
Earlier work this paper cites.
A mean field view of the landscape of two-layer neural networks
Song Mei, Andrea Montanari, and Phan-Minh Nguyen · 2018
Earlier work this paper cites.
Parameters as interacting particles: long time convergence and asymptotic error scaling of neural networks
Grant Rotskoff and Eric Vanden-Eijnden · 2018
Earlier work this paper cites.
Spurious local minima are common in two-layer ReLU neural networks
Itay Safran and Ohad Shamir · 2018
Earlier work this paper cites.
Theoretical insights into the optimization landscape of over-parameterized shallow neural networks
Mahdi Soltanolkotabi, Adel Javanmard, and Jason D Lee · 2018
Earlier work this paper cites.
High-Dimensional Probability: An Introduction with Applications in Data Science
Roman Vershynin · 2018
Earlier work this paper cites.
Learning and generalization in overparameterized neural networks, going beyond two layers
Zeyuan Allen-Zhu, Yuanzhi Li, and Yingyu Liang · 2019
Earlier work this paper cites.
A convergence theory for deep learning via over-parameterization
Zeyuan Allen-Zhu, Yuanzhi Li, and Zhao Song · 2019
Earlier work this paper cites.
A mean-field limit for certain deep neural networks
Dyego Araújo, Roberto I Oliveira, and Daniel Yukimura · 2019
Earlier work this paper cites.
Sanjeev Arora, Simon S Du, Wei Hu, Zhiyuan Li, and Ruosong Wang · 2019
Earlier work this paper cites.
Sparse optimization on measures with over-parameterized gradient descent
Lenaic Chizat · 2019
Earlier work this paper cites.
On lazy training in differentiable programming
Lenaic Chizat, Edouard Oyallon, and Francis Bach · 2019
Earlier work this paper cites.
Width provably matters in optimization for deep linear neural networks
Simon Du and Wei Hu · 2019
Earlier work this paper cites.
Gradient descent finds global minima of deep neural networks
Simon Du, Jason Lee, Haochuan Li, Liwei Wang, and Xiyu Zhai · 2019
Cited alongside, same era.
Disentangling feature and lazy learning in deep neural networks: an empirical study
Mario Geiger, Stefano Spigler, Arthur Jacot, and Matthieu Wyart · 2019
Cited alongside, same era.
Limitations of lazy training of two-layers neural network
Behrooz Ghorbani, Song Mei, Theodor Misiakiewicz, and Andrea Montanari · 2019
Cited alongside, same era.
Dynamics of stochastic gradient descent for two-layer neural networks in the teacher-student setup
Sebastian Goldt, Madhu Advani, Andrew M Saxe, Florent Krzakala, and Lenka Zdeborová · 2019
Cited alongside, same era.
Mean-field langevin dynamics and energy landscape of neural networks
Kaitong Hu, Zhenjie Ren, David Siska, and Lukasz Szpruch · 2019
Cited alongside, same era.
Finite depth and width corrections to the neural tangent kernel
Boris Hanin and Mihai Nica · 2020
Later among the works it cites.
Analysis of a two-layer neural network via displacement convexity
Adel Javanmard, Marco Mondelli, and Andrea Montanari · 2020
Later among the works it cites.
Finite versus infinite neural networks: an empirical study
Jaehoon Lee, Samuel S. Schoenholz, Jeffrey Pennington, Ben Adlam, Lechao Xiao, Roman Novak, and Jascha Sohl-Dickstein · 2020
Later among the works it cites.
Learning over-parametrized two-layer neural networks beyond ntk
Yuanzhi Li, Tengyu Ma, and Hongyang R. Zhang · 2020
Later among the works it cites.
Loss landscapes and optimization in over-parameterized non-linear systems and neural networks
Chaoyue Liu, Libin Zhu, and Mikhail Belkin · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Dynamics of deep neural networks and neural tangent hierarchy
Jiaoyang Huang and Horng-Tzer Yau · 2019
Cited alongside, same era.
Mean-field neural odes via relaxed optimal control
Jean-François Jabir, David Šiška, and Łukasz Szpruch · 2019
Cited alongside, same era.
Mean-field theory of two-layers neural networks: dimension-free bounds and kernel limit
Song Mei, Theodor Misiakiewicz, and Andrea Montanari · 2019
Cited alongside, same era.
Mean field limit of the learning dynamics of multilayer neural networks
Phan-Minh Nguyen · 2019
Cited alongside, same era.
Overparameterized nonlinear learning: Gradient descent takes the shortest path?
Samet Oymak and Mahdi Soltanolkotabi · 2019
Cited alongside, same era.
Pytorch: An imperative style, high-performance deep learning library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al · 2019
Cited alongside, same era.
Mean field analysis of deep neural networks
Justin Sirignano and Konstantinos Spiliopoulos · 2019
Cited alongside, same era.
Yiping Lu, Chao Ma, Yulong Lu, Jianfeng Lu, and Lexing Ying · 2020
Later among the works it cites.
A rigorous framework for the mean field limit of multilayer neural networks
Phan-Minh Nguyen and Huy Tuan Pham · 2020
Later among the works it cites.
Atsushi Nitanda, Denny Wu, and Taiji Suzuki · 2020
Later among the works it cites.
Toward moderate overparameterization: Global convergence guarantees for training shallow neural networks
Samet Oymak and Mahdi Soltanolkotabi · 2020
Later among the works it cites.
Optimization and generalization of shallow neural networks with quadratic activation functions
Stefano Sarao Mannelli, Eric Vanden-Eijnden, and Lenka Zdeborová · 2020
Later among the works it cites.
Mean field analysis of neural networks: A law of large numbers
Justin Sirignano and Konstantinos Spiliopoulos · 2020
Later among the works it cites.
On the convergence of gradient descent training for two-layer relu-networks in the mean field regime
Stephan Wojtowytsch · 2020
Later among the works it cites.
Can shallow neural networks beat the curse of dimensionality? a mean field training perspective
Stephan Wojtowytsch and E Weinan · 2020
Later among the works it cites.
Kernel and rich regimes in overparametrized models
Blake Woodworth, Suriya Gunasekar, Jason D Lee, Edward Moroshko, Pedro Savarese, Itay Golan, Daniel Soudry, and Nathan Srebro · 2020
Later among the works it cites.
The global optimization geometry of shallow linear neural networks
Zhihui Zhu, Daniel Soudry, Yonina C Eldar, and Michael B Wakin · 2020
Later among the works it cites.
Gradient descent optimizes over-parameterized deep relu networks
Difan Zou, Yuan Cao, Dongruo Zhou, and Quanquan Gu · 2020
Later among the works it cites.
Niladri S Chatterji, Philip M Long, and Peter L Bartlett · 2021
Later among the works it cites.
Proxy convexity: A unified framework for the analysis of neural networks trained by gradient descent
Spencer Frei and Quanquan Gu · 2021
Later among the works it cites.
Landscape and training regimes in deep learning
Mario Geiger, Leonardo Petrini, and Matthieu Wyart · 2021
Later among the works it cites.
Are wider nets better given the same number of parameters?
Anna Golubeva, Guy Gur-Ari, and Behnam Neyshabur · 2021
Later among the works it cites.
Phase diagram for two-layer relu neural networks at infinite-width limit
Tao Luo, Zhi-Qin John Xu, Zheng Ma, and Yaoyu Zhang · 2021
Later among the works it cites.
The barron space and the flow-induced function spaces for neural network models
Chao Ma, Lei Wu, et al · 2021
Later among the works it cites.
Global convergence of three-layer neural networks in the mean field regime
Huy Tuan Pham and Phan-Minh Nguyen · 2021
Later among the works it cites.
Tensor programs iv: Feature learning in infinite-width neural networks
Greg Yang and Edward J. Hu · 2021
Later among the works it cites.
A local convergence theory for mildly over-parameterized two-layer neural network
Mo Zhou, Rong Ge, and Chi Jin · 2021
Later among the works it cites.