Fetching the paper…
Reading the bibliography…
Recent works on over-parameterized neural networks have shown that the stochasticity in optimizers has the implicit regularization effect of minimizing the sharpness of the loss function (in particular, the trace of its Hessian) over the family zero-loss solutions.
Improved sample complexities for deep networks and robust classification via an all-layer margin
Colin Wei and Tengyu Ma · 1910
Earlier work this paper cites.
Flat minima
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Guaranteed minimum-rank solutions of linear matrix equations via nuclear norm minimization
Benjamin Recht, Maryam Fazel, and Pablo A Parrilo · 2010
Earlier work this paper cites.
Non-asymptotic theory of random matrices: extreme singular values
Mark Rudelson and Roman Vershynin · 2010
Earlier work this paper cites.
Smoothness, low noise and fast rates
Nathan Srebro, Karthik Sridharan, and Ambuj Tewari · 2010
Earlier work this paper cites.
Tight oracle inequalities for low-rank matrix recovery from a minimal number of noisy random measurements
Emmanuel J Candes and Yaniv Plan · 2011
Earlier work this paper cites.
On large-batch training for deep learning: Generalization gap and sharp minima
Nitish Shirish Keskar, Dheevatsa Mudigere, Jorge Nocedal, Mikhail Smelyanskiy, and Ping Tak Peter Tang · 2016
Earlier work this paper cites.
Sharp minima can generalize for deep nets
Laurent Dinh, Razvan Pascanu, Samy Bengio, and Yoshua Bengio · 2017
Earlier work this paper cites.
Gintare Karolina Dziugaite and Daniel M Roy · 2017
Earlier work this paper cites.
Implicit regularization in matrix factorization
Suriya Gunasekar, Blake E Woodworth, Srinadh Bhojanapalli, Behnam Neyshabur, and Nati Srebro · 2017
Earlier work this paper cites.
Three factors influencing minima in sgd
Stanisław Jastrzebski, Zachary Kenton, Devansh Arpit, Nicolas Ballas, Asja Fischer, Yoshua Bengio, and Amos Storkey · 2017
Earlier work this paper cites.
Yuanzhi Li, Tengyu Ma, and Hongyang Zhang · 2017
Earlier work this paper cites.
Exploring generalization in deep learning
Behnam Neyshabur, Srinadh Bhojanapalli, David McAllester, and Nati Srebro · 2017
Earlier work this paper cites.
Understanding deep learning requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2017
Earlier work this paper cites.
Essentially no barriers in neural network energy landscape
Felix Draxler, Kambis Veschgini, Manfred Salmhofer, and Fred Hamprecht · 2018
Earlier work this paper cites.
Loss surfaces, mode connectivity, and fast ensembling of dnns
Timur Garipov, Pavel Izmailov, Dmitrii Podoprikhin, Dmitry P Vetrov, and Andrew G Wilson · 2018
Earlier work this paper cites.
Implicit regularization in nonconvex statistical estimation: Gradient descent converges linearly for phase retrieval and matrix completion
Cong Ma, Kaizheng Wang, Yuejie Chi, and Yuxin Chen · 2018
Earlier work this paper cites.
How sgd selects the global minima in over-parameterized learning: A dynamical stability perspective
Lei Wu, Chao Ma, and Weinan E · 2018
Cited alongside, same era.
Implicit regularization in deep matrix factorization
Sanjeev Arora, Nadav Cohen, Wei Hu, and Yuping Luo · 2019
Cited alongside, same era.
Implicit regularization for deep neural networks driven by an ornstein-uhlenbeck like process
Guy Blanc, Neha Gupta, Gregory Valiant, and Paul Valiant · 2019
Cited alongside, same era.
The implicit bias of depth: How incremental learning drives generalization
Daniel Gissin, Shai Shalev-Shwartz, and Amit Daniely · 2019
Cited alongside, same era.
Fantastic generalization measures and where to find them
Yiding Jiang, Behnam Neyshabur, Hossein Mobahi, Dilip Krishnan, and Samy Bengio · 2019
Diametrical risk minimization: Theory and computations
Matthew D Norton and Johannes O Royset · 2021
Later among the works it cites.
Implicit regularization in tensor factorization
Noam Razin, Asaf Maman, and Nadav Cohen · 2021
Later among the works it cites.
Small random initialization is akin to spectral learning: Optimization and generalization guarantees for overparameterized low-rank matrix reconstruction
Dominik Stöger and Mahdi Soltanolkotabi · 2021
Later among the works it cites.
Regularizing neural networks via adversarial model perturbation
Yaowei Zheng, Richong Zhang, and Yongyi Mao · 2021
Later among the works it cites.
Towards understanding sharpness-aware minimization
Maksym Andriushchenko and Nicolas Flammarion · 2022
Later among the works it cites.
Understanding gradient descent on edge of stability in deep learning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
On implicit regularization: Morse functions and applications to matrix factorization
Mohamed Ali Belabbas · 2020
Cited alongside, same era.
Sharpness-aware minimization for efficiently improving generalization
Pierre Foret, Ariel Kleiner, Hossein Mobahi, and Behnam Neyshabur · 2020
Cited alongside, same era.
Zhiyuan Li, Yuping Luo, and Kaifeng Lyu · 2020
Cited alongside, same era.
Implicit regularization in deep learning may not be explainable by norms
Noam Razin and Nadav Cohen · 2020
Cited alongside, same era.
Adversarial weight perturbation helps robust generalization
Dongxian Wu, Shu-Tao Xia, and Yisen Wang · 2020
Cited alongside, same era.
Gradient descent on neural networks typically occurs at the edge of stability, 2021
Jeremy M. Cohen, Simran Kaur, Yuanzhi Li, J. Zico Kolter, and Ameet Talwalkar · 2021
Cited alongside, same era.
Label noise sgd provably prefers flat global minimizers, 2021
Alex Damian, Tengyu Ma, and Jason Lee · 2021
Cited alongside, same era.
Sanjeev Arora, Zhiyuan Li, and Abhishek Panigrahi · 2022
Later among the works it cites.
Adaptive gradient methods at the edge of stability
Jeremy M Cohen, Behrooz Ghorbani, Shankar Krishnan, Naman Agarwal, Sourabh Medapati, Michal Badura, Daniel Suo, David Cardoze, Zachary Nado, George E Dahl, et al · 2022
Later among the works it cites.
Self-stabilization: The implicit bias of gradient descent at the edge of stability
Alex Damian, Eshaan Nichani, and Jason D Lee · 2022
Later among the works it cites.
Flat minima generalize for low-rank matrix recovery
Lijun Ding, Dmitriy Drusvyatskiy, and Maryam Fazel · 2022
Later among the works it cites.
Analyzing sharpness along gd trajectory: Progressive sharpening and edge of stability
Zhouzi Li, Zixuan Wang, and Jian Li · 2022
Later among the works it cites.
Understanding the generalization benefit of normalization layers: Sharpness reduction
Kaifeng Lyu, Zhiyuan Li, and Sanjeev Arora · 2022
Later among the works it cites.
The multiscale structure of neural network loss functions: The effect on optimization and origin
Chao Ma, Lei Wu, and Lexing Ying · 2022
Later among the works it cites.
Implicit bias of the step size in linear diagonal neural networks
Mor Shpigel Nacson, Kavya Ravichandran, Nathan Srebro, and Daniel Soudry · 2022
Later among the works it cites.
How does sharpness-aware minimization minimize sharpness?
Kaiyue Wen, Tengyu Ma, and Zhiyuan Li · 2022
Later among the works it cites.
Penalizing gradient norm for efficiently improving generalization in deep learning
Yang Zhao, Hao Zhang, and Xiuyuan Hu · 2022
Later among the works it cites.
Surrogate gap minimization improves sharpness-aware training
Juntang Zhuang, Boqing Gong, Liangzhe Yuan, Yin Cui, Hartwig Adam, Nicha Dvornek, Sekhar Tatikonda, James Duncan, and Ting Liu · 2022
Later among the works it cites.
A modern look at the relationship between sharpness and generalization
Maksym Andriushchenko, Francesco Croce, Maximilian Müller, Matthias Hein, and Nicolas Flammarion · 2023
Closest in time.