Fetching the paper…
Reading the bibliography…
ReLU neural-networks have been in the focus of many recent theoretical works, trying to explain their empirical success.
Samet Oymak and Mahdi Soltanolkotabi · 1902
Earlier work this paper cites.
Random features for large-scale kernel machines
Ali Rahimi and Benjamin Recht · 2008
Earlier work this paper cites.
Optimistic Rates for Learning with a Smooth Loss
Nathan Srebro, Karthik Sridharan, and Ambuj Tewari · 2010
Earlier work this paper cites.
Exact solutions to the nonlinear dynamics of learning in deep linear neural networks
Andrew M. Saxe, James L. McClelland, and Surya Ganguli · 2013
Earlier work this paper cites.
On the computational efficiency of training neural networks
Roi Livni, Shai Shalev-Shwartz, and Ohad Shamir · 2014
Earlier work this paper cites.
Understanding machine learning: From theory to algorithms
Shai Shalev-Shwartz and Shai Ben-David · 2014
Earlier work this paper cites.
Toward deeper understanding of neural networks: The power of initialization and a dual view on expressivity
Amit Daniely, Roy Frostig, and Yoram Singer · 2016
Earlier work this paper cites.
Deep learning without poor local minima
Kenji Kawaguchi · 2016
Earlier work this paper cites.
Diverse neural network learns true target functions
Bo Xie, Yingyu Liang, and Le Song · 2016
Earlier work this paper cites.
Sgd learns over-parameterized networks that provably generalize on linearly separable data
Alon Brutzkus, Amir Globerson, Eran Malach, and Shai Shalev-Shwartz · 2017
Earlier work this paper cites.
Sgd learns the conjugate kernel class of the network
Amit Daniely · 2017
Cited alongside, same era.
Depth creates no bad local minima
Haihao Lu and Kenji Kawaguchi · 2017
Cited alongside, same era.
Failures of gradient-based deep learning
Shai Shalev-Shwartz, Ohad Shamir, and Shaked Shammah · 2017
Cited alongside, same era.
Learning and generalization in overparameterized neural networks, going beyond two layers
Zeyuan Allen-Zhu, Yuanzhi Li, and Yingyu Liang · 2018
Cited alongside, same era.
A convergence theory for deep learning via over-parameterization
Zeyuan Allen-Zhu, Yuanzhi Li, and Zhao Song · 2018
Cited alongside, same era.
On the connection between learning two-layers neural networks and tensor decomposition
Marco Mondelli and Andrea Montanari · 2018
Later among the works it cites.
Overparameterized nonlinear learning: Gradient descent takes the shortest path?
Samet Oymak and Mahdi Soltanolkotabi · 2018
Later among the works it cites.
Distribution-specific hardness of learning neural networks
Ohad Shamir · 2018
Later among the works it cites.
Matrix Chernoff bound, November 2018
Wikipedia · 2018
Later among the works it cites.
The Global Optimization Geometry of Shallow Linear Neural Networks
Zhihui Zhu, Daniel Soudry, Yonina C. Eldar, and Michael B. Wakin · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Sanjeev Arora, Nadav Cohen, Noah Golowich, and Wei Hu · 2018
Cited alongside, same era.
On the optimization of deep networks: Implicit acceleration by overparameterization
Sanjeev Arora, Nadav Cohen, and Elad Hazan · 2018
Cited alongside, same era.
Gradient descent provably optimizes over-parameterized neural networks
Simon S Du, Xiyu Zhai, Barnabas Poczos, and Aarti Singh · 2018
Cited alongside, same era.
EE227c: Convex Optimization and Approximation
Moritz Hardt · 2018
Cited alongside, same era.
Sanjeev Arora, Simon S Du, Wei Hu, Zhiyuan Li, and Ruosong Wang · 2019
Closest in time.
Wide neural networks of any depth evolve as linear models under gradient descent
Jaehoon Lee, Lechao Xiao, Samuel S Schoenholz, Yasaman Bahri, Jascha Sohl-Dickstein, and Jeffrey Pennington · 2019
Closest in time.
Chao Ma, Lei Wu, et al · 2019
Closest in time.
Theoretical insights into the optimization landscape of over-parameterized shallow neural networks
Mahdi Soltanolkotabi, Adel Javanmard, and Jason D Lee · 2019
Closest in time.