Fetching the paper…
Reading the bibliography…
The generalization mystery of overparametrized deep nets has motivated efforts to understand how gradient descent (GD) converges to low-loss solutions that generalize well.
Generalized gradients and applications
Frank H. Clarke · 1975
Earlier work this paper cites.
An introduction to o-minimal geometry
Michel Coste · 2000
Earlier work this paper cites.
Nonsmooth analysis and control theory , volume 178
Francis H. Clarke, Yuri S. Ledyaev, Ronald J. Stern, and Peter R. Wolenski · 2008
Earlier work this paper cites.
Characterizations of łojasiewicz inequalities: subgradient flows, talweg, convexity
Jérôme Bolte, Aris Daniilidis, Olivier Ley, and Laurent Mazet · 2010
Earlier work this paper cites.
Approximate KKT points and a proximity measure for termination
Joydeep Dutta, Kalyanmoy Deb, Rupesh Tulshyan, and Ramnik Arora · 2013
Earlier work this paper cites.
Rectifier nonlinearities improve neural network acoustic models
Andrew L. Maas, Awni Y. Hannun, and Andrew Y. Ng · 2013
Earlier work this paper cites.
Delving deep into rectifiers: Surpassing human-level performance on ImageNet classification
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2015
Earlier work this paper cites.
Spectrally-normalized margin bounds for neural networks
Peter L Bartlett, Dylan J Foster, and Matus J Telgarsky · 2017
Earlier work this paper cites.
Understanding deep learning requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2017
Earlier work this paper cites.
Learning and generalization in overparameterized neural networks, going beyond two layers, 2018
Zeyuan Allen-Zhu, Yuanzhi Li, and Yingyu Liang · 2018
Earlier work this paper cites.
SGD learns over-parameterized networks that provably generalize on linearly separable data
Alon Brutzkus, Amir Globerson, Eran Malach, and Shai Shalev-Shwartz · 2018
Earlier work this paper cites.
Algorithmic regularization in learning deep homogeneous models: Layers are automatically balanced
Simon S. Du, Wei Hu, and Jason D. Lee · 2018
Earlier work this paper cites.
Neural tangent kernel: Convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel, and Clement Hongler · 2018
Earlier work this paper cites.
Risk and parameter convergence of logistic regression
Ziwei Ji and Matus Telgarsky · 2018
Earlier work this paper cites.
Algorithmic regularization in over-parameterized matrix sensing and neural networks with quadratic activations
Yuanzhi Li, Tengyu Ma, and Hongyang Zhang · 2018
Earlier work this paper cites.
Gradient descent quantizes relu network features
Hartmut Maennel, Olivier Bousquet, and Sylvain Gelly · 2018
Cited alongside, same era.
A PAC-bayesian approach to spectrally-normalized margin bounds for neural networks
Behnam Neyshabur, Srinadh Bhojanapalli, and Nathan Srebro · 2018
Cited alongside, same era.
Stochastic gradient descent optimizes over-parameterized deep relu networks
Difan Zou, Yuan Cao, Dongruo Zhou, and Quanquan Gu · 2018
Cited alongside, same era.
A convergence theory for deep learning via over-parameterization
Zeyuan Allen-Zhu, Yuanzhi Li, and Zhao Song · 2019
Cited alongside, same era.
Why do larger models generalize better? A theoretical perspective via the XOR problem
Alon Brutzkus and Amir Globerson · 2019
Cited alongside, same era.
Gradient descent maximizes the margin of homogeneous neural networks
Kaifeng Lyu and Jian Li · 2020
Later among the works it cites.
Implicit bias in deep linear classification: Initialization scale vs training accuracy
Edward Moroshko, Blake E Woodworth, Suriya Gunasekar, Jason D Lee, Nati Srebro, and Daniel Soudry · 2020
Later among the works it cites.
An analytic theory of shallow networks dynamics for hinge loss classification
Franco Pellegrini and Giulio Biroli · 2020
Later among the works it cites.
The pitfalls of simplicity bias in neural networks
Harshay Shah, Kaustav Tamuly, Aditi Raghunathan, Prateek Jain, and Praneeth Netrapalli · 2020
Later among the works it cites.
Kernel and rich regimes in overparametrized models
Blake Woodworth, Suriya Gunasekar, Jason D. Lee, Edward Moroshko, Pedro Savarese, Itay Golan, Daniel Soudry, and Nathan Srebro · 2020
Later among the works it cites.
How much over-parameterization is sufficient to learn deep ReLU networks?
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
On lazy training in differentiable programming
Lénaïc Chizat, Edouard Oyallon, and Francis Bach · 2019
Cited alongside, same era.
Implicit regularization of discrete gradient dynamics in linear neural networks
Gauthier Gidel, Francis Bach, and Simon Lacoste-Julien · 2019
Cited alongside, same era.
SGD on neural networks learns functions of increasing complexity
Dimitris Kalimeris, Gal Kaplun, Preetum Nakkiran, Benjamin Edelman, Tristan Yang, Boaz Barak, and Haofeng Zhang · 2019
Cited alongside, same era.
Gradient dynamics of shallow univariate relu networks
Francis Williams, Matthew Trager, Daniele Panozzo, Claudio Silva, Denis Zorin, and Joan Bruna · 2019
Cited alongside, same era.
Implicit bias of gradient descent for wide two-layer neural networks trained with the logistic loss
Lénaïc Chizat and Francis Bach · 2020
Cited alongside, same era.
Stochastic subgradient method converges on tame functions
Damek Davis, Dmitriy Drusvyatskiy, Sham Kakade, and Jason D. Lee · 2020
Cited alongside, same era.
The implicit bias of depth: How incremental learning drives generalization
Daniel Gissin, Shai Shalev-Shwartz, and Amit Daniely · 2020
Cited alongside, same era.
Zixiang Chen, Yuan Cao, Difan Zou, and Quanquan Gu · 2021
Closest in time.
Provable generalization of sgd-trained neural networks of any width in the presence of adversarial label noise
Spencer Frei, Yuan Cao, and Quanquan Gu · 2021
Closest in time.
Towards resolving the implicit bias of gradient descent for matrix factorization: Greedy low-rank learning
Zhiyuan Li, Yuping Luo, and Kaifeng Lyu · 2021
Closest in time.
Phase diagram for two-layer relu neural networks at infinite-width limit
Tao Luo, Zhi-Qin John Xu, Zheng Ma, and Yaoyu Zhang · 2021
Closest in time.
Extreme memorization via scale of initialization
Harsh Mehta, Ashok Cutkosky, and Behnam Neyshabur · 2021
Closest in time.
The inductive bias of ReLU networks on orthogonally separable data
Mary Phuong and Christoph H Lampert · 2021
Closest in time.
Implicit regularization in tensor factorization
Noam Razin, Asaf Maman, and Nadav Cohen · 2021
Closest in time.
Towards understanding learning in neural networks with linear teachers
Roei Sarussi, Alon Brutzkus, and Amir Globerson · 2021
Closest in time.
Towards understanding the condensation of two-layer neural networks at initial training
Zhi-Qin John Xu, Hanxu Zhou, Tao Luo, and Yaoyu Zhang · 2021
Closest in time.