Fetching the paper…
Reading the bibliography…
The behavior of the gradient descent (GD) algorithm is analyzed for a deep neural network model with skip-connections.
Accurate error bounds for the eigenvalues of the kernel matrix
Mikio L Braun · 2006
Earlier work this paper cites.
Random features for large-scale kernel machines
Ali Rahimi and Benjamin Recht · 2008
Earlier work this paper cites.
Uniform approximation of functions with random bases
Ali Rahimi and Benjamin Recht · 2008
Earlier work this paper cites.
Weighted sums of random kitchen sinks: Replacing minimization with randomization in learning
Ali Rahimi and Benjamin Recht · 2009
Earlier work this paper cites.
In search of the real inductive bias: On the role of implicit regularization in deep learning
Behnam Neyshabur, Ryota Tomioka, and Nathan Srebro · 2014
Earlier work this paper cites.
Understanding machine learning: From theory to algorithms
Shai Shalev-Shwartz and Shai Ben-David · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Norm-based capacity control in neural networks
Behnam Neyshabur, Ryota Tomioka, and Nathan Srebro · 2015
Earlier work this paper cites.
SGD learns the conjugate kernel class of the network
Amit Daniely · 2017
Earlier work this paper cites.
Diverse neural network learns true target functions
Bo Xie, Yingyu Liang, and Le Song · 2017
Earlier work this paper cites.
Understanding deep learning requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2017
Earlier work this paper cites.
Learning and generalization in overparameterized neural networks, going beyond two layers
Zeyuan Allen-Zhu, Yuanzhi Li, and Yingyu Liang · 2018
Cited alongside, same era.
A convergence theory for deep learning via over-parameterization
Zeyuan Allen-Zhu, Yuanzhi Li, and Zhao Song · 2018
Cited alongside, same era.
Gradient descent with identity initialization efficiently learns positive definite linear transformations
Peter L. Bartlett, David P. Helmbold, and Philip M. Long · 2018
Cited alongside, same era.
Learning with SGD and random features
Luigi Carratino, Alessandro Rudi, and Lorenzo Rosasco · 2018
Cited alongside, same era.
On the global convergence of gradient descent for over-parameterized models using optimal transport
Lénaïc Chizat and Francis Bach · 2018
Cited alongside, same era.
Mean field analysis of neural networks: A central limit theorem
Justin Sirignano and Konstantinos Spiliopoulos · 2018
Later among the works it cites.
A mean field view of the landscape of two-layers neural networks
Mei Song, Andrea Montanari, and Phan-Minh Nguyen · 2018
Later among the works it cites.
Stochastic gradient descent optimizes over-parameterized deep ReLU networks
Difan Zou, Yuan Cao, Dongruo Zhou, and Quanquan Gu · 2018
Later among the works it cites.
A convergence analysis of gradient descent for deep linear neural networks
Sanjeev Arora, Nadav Cohen, Noah Golowich, and Wei Hu · 2019
Closest in time.
Sanjeev Arora, Simon S Du, Wei Hu, Zhiyuan Li, and Ruosong Wang · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Gradient descent finds global minima of deep neural networks
Simon S Du, Jason D Lee, Haochuan Li, Liwei Wang, and Xiyu Zhai · 2018
Cited alongside, same era.
A priori estimates for two-layer neural networks
Weinan E, Chao Ma, and Lei Wu · 2018
Cited alongside, same era.
Neural tangent kernel: Convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel, and Clément Hongler · 2018
Cited alongside, same era.
Learning overparameterized neural networks via stochastic gradient descent on structured data
Yuanzhi Li and Yingyu Liang · 2018
Cited alongside, same era.
Foundations of machine learning
Mehryar Mohri, Afshin Rostamizadeh, and Ameet Talwalkar · 2018
Cited alongside, same era.
Parameters as interacting particles: long time convergence and asymptotic error scaling of neural networks
Grant Rotskoff and Eric Vanden-Eijnden · 2018
Cited alongside, same era.
Closest in time.
A generalization theory of gradient descent for learning over-parameterized deep ReLU networks
Yuan Cao and Quanquan Gu · 2019
Closest in time.
Width provably matters in optimization for deep linear neural networks
Simon S Du and Wei Hu · 2019
Closest in time.
Gradient descent provably optimizes over-parameterized neural networks
Simon S. Du, Xiyu Zhai, Barnabas Poczos, and Aarti Singh · 2019
Closest in time.
A priori estimates of the population risk for residual networks
Weinan E, Chao Ma, and Qingcan Wang · 2019
Closest in time.
Weinan E, Chao Ma, and Lei Wu · 2019
Closest in time.