Generalized gradients and applications
Frank H Clarke · 1975
Earlier work this paper cites.
Characterization of the subdifferential of some matrix norms
G Alistair Watson · 1992
Earlier work this paper cites.
Nonsmooth analysis in control theory: a survey
Francis Clarke · 2001
Earlier work this paper cites.
Empirical Margin Distributions and Bounding the Generalization Error of Combined Classifiers
V. Koltchinskii and D. Panchenko · 2002
Earlier work this paper cites.
Matrix completion with noise
Emmanuel J Candes and Yaniv Plan · 2010
Earlier work this paper cites.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2015
Earlier work this paper cites.
Implicit regularization in matrix factorization
Suriya Gunasekar, Blake E Woodworth, Srinadh Bhojanapalli, Behnam Neyshabur, and Nati Srebro · 2017
Earlier work this paper cites.
Implicit bias of gradient descent on linear convolutional networks
Suriya Gunasekar, Jason D Lee, Daniel Soudry, and Nati Srebro · 2018
Earlier work this paper cites.
Neural tangent kernel: Convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel, and Clement Hongler · 2018
Earlier work this paper cites.
Foundations of Machine Learning
Mehryar Mohri, Afshin Rostamizadeh, and Ameet Talwalkar · 2018
Earlier work this paper cites.
The implicit bias of gradient descent on separable data
Daniel Soudry, Elad Hoffer, Mor Shpigel Nacson, Suriya Gunasekar, and Nathan Srebro · 2018
Earlier work this paper cites.
A convergence theory for deep learning via over-parameterization
Zeyuan Allen-Zhu, Yuanzhi Li, and Zhao Song · 2019
Earlier work this paper cites.
Implicit regularization in deep matrix factorization
Sanjeev Arora, Nadav Cohen, Wei Hu, and Yuping Luo · 2019
Earlier work this paper cites.
On exact computation with an infinitely wide neural net
Sanjeev Arora, Simon S Du, Wei Hu, Zhiyuan Li, Russ R Salakhutdinov, and Ruosong Wang · 2019
Earlier work this paper cites.
Generalization bounds of stochastic gradient descent for wide and deep neural networks
Yuan Cao and Quanquan Gu · 2019
Earlier work this paper cites.
On lazy training in differentiable programming
Lénaïc Chizat, Edouard Oyallon, and Francis Bach · 2019
Earlier work this paper cites.
The Goldilocks zone: Towards better understanding of neural network loss landscapes
Stanislav Fort and Adam Scherlis · 2019
Earlier work this paper cites.
On the Rademacher complexity of linear hypothesis sets
Original
Pranjal Awasthi, Natalie Frank, and Mehryar Mohri · 2020
Earlier work this paper cites.
Implicit regularization for deep neural networks driven by an Ornstein-Uhlenbeck like process
Guy Blanc, Neha Gupta, Gregory Valiant, and Paul Valiant · 2020
Earlier work this paper cites.
Implicit bias of gradient descent for wide two-layer neural networks trained with the logistic loss
Lenaic Chizat and Francis Bach · 2020
Earlier work this paper cites.
Simple and effective regularization methods for training on noisily labeled data with generalization guarantee
Wei Hu, Zhiyuan Li, and Dingli Yu · 2020
Earlier work this paper cites.