Fetching the paper…
Reading the bibliography…
Despite extensive studies, the underlying reason as to why overparameterized neural networks can generalize remains elusive.
Improved sample complexities for deep networks and robust classification via an all-layer margin
Colin Wei and Tengyu Ma · 1910
Earlier work this paper cites.
Flat minima
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
On variants of the johnson–lindenstrauss lemma
Jiří Matoušek · 2008
Earlier work this paper cites.
Smoothness, low noise and fast rates
Nathan Srebro, Karthik Sridharan, and Ambuj Tewari · 2010
Earlier work this paper cites.
Tight bounds for the expected risk of linear classifiers and pac-bayes finite-sample guarantees
Jean Honorio and Tommi Jaakkola · 2014
Earlier work this paper cites.
Understanding machine learning: From theory to algorithms
Shai Shalev-Shwartz and Shai Ben-David · 2014
Earlier work this paper cites.
On large-batch training for deep learning: Generalization gap and sharp minima
Nitish Shirish Keskar, Dheevatsa Mudigere, Jorge Nocedal, Mikhail Smelyanskiy, and Ping Tak Peter Tang · 2016
Earlier work this paper cites.
Understanding deep learning requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2016
Earlier work this paper cites.
Sharp minima can generalize for deep nets
Laurent Dinh, Razvan Pascanu, Samy Bengio, and Yoshua Bengio · 2017
Earlier work this paper cites.
Gintare Karolina Dziugaite and Daniel M Roy · 2017
Earlier work this paper cites.
Implicit regularization in matrix factorization
Suriya Gunasekar, Blake E Woodworth, Srinadh Bhojanapalli, Behnam Neyshabur, and Nati Srebro · 2017
Earlier work this paper cites.
Three factors influencing minima in sgd
Stanisław Jastrzebski, Zachary Kenton, Devansh Arpit, Nicolas Ballas, Asja Fischer, Yoshua Bengio, and Amos Storkey · 2017
Earlier work this paper cites.
Yuanzhi Li, Tengyu Ma, and Hongyang Zhang · 2017
Earlier work this paper cites.
Exploring generalization in deep learning
Behnam Neyshabur, Srinadh Bhojanapalli, David McAllester, and Nati Srebro · 2017
Earlier work this paper cites.
The loss landscape of overparameterized neural networks
Yaim Cooper · 2018
Cited alongside, same era.
Neural tangent kernel: Convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel, and Clément Hongler · 2018
Cited alongside, same era.
The implicit bias of gradient descent on separable data
Daniel Soudry, Elad Hoffer, Mor Shpigel Nacson, Suriya Gunasekar, and Nathan Srebro · 2018
Cited alongside, same era.
How sgd selects the global minima in over-parameterized learning: A dynamical stability perspective
Lei Wu, Chao Ma, et al · 2018
Cited alongside, same era.
Implicit regularization for deep neural networks driven by an ornstein-uhlenbeck like process
Guy Blanc, Neha Gupta, Gregory Valiant, and Paul Valiant · 2019
Cited alongside, same era.
On linear stability of sgd and input-smoothness of neural networks
Chao Ma and Lexing Ying · 2021
Later among the works it cites.
Diametrical risk minimization: Theory and computations
Matthew D Norton and Johannes O Royset · 2021
Later among the works it cites.
Understanding gradient descent on edge of stability in deep learning
Sanjeev Arora, Zhiyuan Li, and Abhishek Panigrahi · 2022
Later among the works it cites.
Peter L Bartlett, Philip M Long, and Olivier Bousquet · 2022
Later among the works it cites.
Flat minima generalize for low-rank matrix recovery
Lijun Ding, Dmitriy Drusvyatskiy, and Maryam Fazel · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yiding Jiang, Behnam Neyshabur, Hossein Mobahi, Dilip Krishnan, and Samy Bengio · 2019
Cited alongside, same era.
36-709: Advanced probability theory
Alessandro Rinaldo · 2019
Cited alongside, same era.
Regularization matters: Generalization and optimization of neural nets vs their induced kernel
Colin Wei, Jason D Lee, Qiang Liu, and Tengyu Ma · 2019
Cited alongside, same era.
Convergence rates for the stochastic gradient descent method for non-convex objective functions
Benjamin Fehrman, Benjamin Gess, and Arnulf Jentzen · 2020
Cited alongside, same era.
Shape matters: Understanding the implicit bias of the noise covariance
Jeff Z HaoChen, Colin Wei, Jason D Lee, and Tengyu Ma · 2020
Cited alongside, same era.
Kernel and rich regimes in overparametrized models
Blake Woodworth, Suriya Gunasekar, Jason D Lee, Edward Moroshko, Pedro Savarese, Itay Golan, Daniel Soudry, and Nathan Srebro · 2020
Cited alongside, same era.
Label noise sgd provably prefers flat global minimizers, 2021
Alex Damian, Tengyu Ma, and Jason Lee · 2021
Cited alongside, same era.
Fast mixing of stochastic gradient descent with normalization and weight decay
Zhiyuan Li, Tianhao Wang, and Dingli Yu · 2022
Later among the works it cites.
Understanding the generalization benefit of normalization layers: Sharpness reduction
Kaifeng Lyu, Zhiyuan Li, and Sanjeev Arora · 2022
Later among the works it cites.
Implicit bias of the step size in linear diagonal neural networks
Mor Shpigel Nacson, Kavya Ravichandran, Nathan Srebro, and Daniel Soudry · 2022
Later among the works it cites.
Explicit regularization in overparametrized models via noise injection
Antonio Orvieto, Anant Raj, Hans Kersting, and Francis Bach · 2022
Later among the works it cites.
Statistically meaningful approximation: a case study on approximating turing machines with transformers
Colin Wei, Yining Chen, and Tengyu Ma · 2022
Later among the works it cites.
How does sharpness-aware minimization minimize sharpness?
Kaiyue Wen, Tengyu Ma, and Zhiyuan Li · 2022
Later among the works it cites.
The inductive bias of flatness regularization for deep matrix factorization
Khashayar Gatmiry, Zhiyuan Li, Ching-Yao Chuang, Sashank Reddi, Tengyu Ma, and Stefanie Jegelka · 2023
Closest in time.
The implicit regularization of dynamical stability in stochastic gradient descent
Lei Wu and Weijie J Su · 2023
Closest in time.