Fetching the paper…
Reading the bibliography…
Finding the optimal configuration of parameters in ResNet is a nonconvex minimization problem, but first-order methods nevertheless find the global optimum in the overparameterized regime.
Gradient flows: in metric spaces and in the space of probability measures
L. Ambrosio, N. Gigli, and G. Savaré · 2008
Earlier work this paper cites.
Towards a mathematical understanding of neural network-based machine learning: what we know and what we don’t
W. E, C. Ma, S. Wojtowytsch, and L. Wu · 2009
Earlier work this paper cites.
Escaping from saddle points — online stochastic gradient for tensor decomposition
R. Ge, F. Huang, C. Jin, and Y. Yuan · 2015
Earlier work this paper cites.
Deep learning without poor local minima
K. Kawaguchi · 2016
Earlier work this paper cites.
Gradient descent can take exponential time to escape saddle points
S. Du, C. Jin, J. Lee, M. Jordan, A. Singh, and B. Póczos · 2017
Earlier work this paper cites.
How to escape saddle points efficiently
C. Jin, R. Ge, P. Netrapalli, S. Kakade, and M. Jordan · 2017
Earlier work this paper cites.
The loss surface of deep and wide neural networks
Q. Nguyen and M. Hein · 2017
Earlier work this paper cites.
On the global convergence of gradient descent for over-parameterized models using optimal transport
L. Chizat and F. Bach · 2018
Earlier work this paper cites.
On the power of over-parametrization in neural networks with quadratic activation
S. Du and J. Lee · 2018
Earlier work this paper cites.
Learning one-hidden-layer neural networks with landscape design
R. Ge, J. Lee, and T. Ma · 2018
Earlier work this paper cites.
Neural tangent kernel: Convergence and generalization in neural networks
A. Jacot, F. Gabriel, and C. Hongler · 2018
Earlier work this paper cites.
Learning overparameterized neural networks via stochastic gradient descent on structured data
Y. Li and Y. Liang · 2018
Earlier work this paper cites.
A mean field view of the landscape of two-layer neural networks
S. Mei, A. Montanari, and P. M. Nguyen · 2018
Earlier work this paper cites.
Optimization landscape and expressivity of deep cnns
Q. Nguyen and M. Hein · 2018
Cited alongside, same era.
Global optimality conditions for deep neural networks
C. Yun, S. Sra, and A. Jadbabaie · 2018
Cited alongside, same era.
What can ResNet learn efficiently, going beyond kernels?
Z. Allen-Zhu and Y. Li · 2019
Cited alongside, same era.
A convergence theory for deep learning via over-parameterization
Z. Allen-Zhu, Y. Li, and Z. Song · 2019
Cited alongside, same era.
A mean-field limit for certain deep neural networks
D. Araújo, R. Oliveira, and D. Yukimura · 2019
Cited alongside, same era.
On exact computation with an infinitely wide neural net
S. Arora, S. Du, W. Hu, Z. Li, R. Salakhutdinov, and R. Wang · 2019
Cited alongside, same era.
Gradient descent optimizes over-parameterized deep relu networks
D. Zou, Y. Cao, D. Zhou, and Q. Gu · 2019
Later among the works it cites.
Generalization of two-layer neural networks: An asymptotic viewpoint
J. Ba, M. Erdogdu, T. Suzuki, D. Wu, and T. Zhang · 2020
Later among the works it cites.
A generalized neural tangent kernel analysis for two-layer neural networks
Z. Chen, Y. Cao, Q. Gu, and T. Zhang · 2020
Later among the works it cites.
On the linearity of large non-linear models: when and why the tangent kernel is constant
C. Liu, L. Zhu, and M. Belkin · 2020
Later among the works it cites.
A mean field analysis of deep ResNet and beyond: Towards provably optimization via overparameterization from depth
Y. Lu, C. Ma, Y. Lu, J. Lu, and L. Ying · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Convex formulation of overparameterized deep neural networks
C. Fang, Y. Gu, W. Zhang, and T. Zhang · 2019
Cited alongside, same era.
Algorithm-dependent generalization bounds for overparameterized deep residual networks
S. Frei, Y. Cao, and Q. Gu · 2019
Cited alongside, same era.
Mean field limit of the learning dynamics of multilayer neural networks
P. M. Nguyen · 2019
Cited alongside, same era.
Theoretical insights into the optimization landscape of over-parameterized shallow neural networks
M. Soltanolkotabi, A. Javanmard, and J. Lee · 2019
Cited alongside, same era.
Regularization matters: Generalization and optimization of neural nets v.s. their induced kernel
C. Wei, J. Lee, Q. Liu, and T. Ma · 2019
Cited alongside, same era.
Convergence theory of learning over-parameterized resnet: A full characterization
H. Zhang, D. Yu, M. Yi, W. Chen, and T. Liu · 2019
Cited alongside, same era.
Mean field analysis of neural networks: A law of large numbers
J. Sirignano and K. Spiliopoulos · 2020
Later among the works it cites.
Mean field analysis of deep neural networks
J. Sirignano and K. Spiliopoulos · 2020
Later among the works it cites.
On the convergence of gradient descent training for two-layer relu-networks in the mean field regime
S. Wojtowytsch · 2020
Later among the works it cites.
When does gradient descent with logistic loss interpolate using deep networks with smoothed relu activations?
N. Chatterji, P. Long, and P. Bartlett · 2021
Closest in time.
Overparameterization of deep resnet: zero loss and mean-field analysis
Z. Ding, S. Chen, Q. Li, and S. Wright · 2021
Closest in time.
Mean-field neural odes via relaxed optimal control
J. Jabir, D. Šiška, and Ł. Szpruch · 2021
Closest in time.
A rigorous framework for the mean field limit of multilayer neural networks
P. M. Nguyen and H. Pham · 2021
Closest in time.