Fetching the paper…
Reading the bibliography…
Empirical evidence suggests that for a variety of overparameterized nonlinear models, most notably in neural network training, the growth of the loss around a minimizer strongly impacts its performance.
Flat minima
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
The sample complexity of pattern classification with neural networks: the size of the weights is more important than the size of the network
Peter L Bartlett · 1998
Earlier work this paper cites.
Rademacher and gaussian complexities: Risk bounds and structural results
Peter L Bartlett and Shahar Mendelson · 2002
Earlier work this paper cites.
Rank, trace-norm and max-norm
Nathan Srebro and Adi Shraibman · 2005
Earlier work this paper cites.
Exact matrix completion via convex optimization
Emmanuel J Candès and Benjamin Recht · 2009
Earlier work this paper cites.
Guaranteed minimum-rank solutions of linear matrix equations via nuclear norm minimization
Benjamin Recht, Maryam Fazel, and Pablo A Parrilo · 2010
Earlier work this paper cites.
Robust principal component analysis?
Emmanuel J Candès, Xiaodong Li, Yi Ma, and John Wright · 2011
Earlier work this paper cites.
Tight oracle inequalities for low-rank matrix recovery from a minimal number of noisy random measurements
Emmanuel J Candes and Yaniv Plan · 2011
Earlier work this paper cites.
Rank-sparsity incoherence for matrix decomposition
Venkat Chandrasekaran, Sujay Sanghavi, Pablo A Parrilo, and Alan S Willsky · 2011
Earlier work this paper cites.
A simpler approach to matrix completion
Benjamin Recht · 2011
Earlier work this paper cites.
Blind deconvolution using convex programming
Ali Ahmed, Benjamin Recht, and Justin Romberg · 2013
Earlier work this paper cites.
Phaselift: Exact and stable signal recovery from magnitude measurements via convex programming
Emmanuel J Candes, Thomas Strohmer, and Vladislav Voroninski · 2013
Earlier work this paper cites.
Low-rank matrix recovery from errors and erasures
Yudong Chen, Ali Jalali, Sujay Sanghavi, and Constantine Caramanis · 2013
Earlier work this paper cites.
In search of the real inductive bias: On the role of implicit regularization in deep learning
Behnam Neyshabur, Ryota Tomioka, and Nathan Srebro · 2014
Earlier work this paper cites.
Rop: Matrix recovery via rank-one projections
T Tony Cai and Anru Zhang · 2015
Earlier work this paper cites.
Incoherence-optimal matrix completion
Yudong Chen · 2015
Earlier work this paper cites.
Exact and stable covariance estimation from quadratic sampling via convex programming
Yuxin Chen, Yuejie Chi, and Andrea J Goldsmith · 2015
Earlier work this paper cites.
Self-calibration and biconvex compressive sensing
Shuyang Ling and Thomas Strohmer · 2015
Earlier work this paper cites.
On large-batch training for deep learning: Generalization gap and sharp minima
Keskar Nitish Shirish, Mudigere Dheevatsa, Nocedal Jorge, Smelyanskiy Mikhail, and Tang Ping Tak Peter · 2016
Earlier work this paper cites.
Sharp minima can generalize for deep nets
Laurent Dinh, Razvan Pascanu, Samy Bengio, and Yoshua Bengio · 2017
Cited alongside, same era.
Gintare Karolina Dziugaite and Daniel M Roy · 2017
Cited alongside, same era.
No spurious local minima in nonconvex low rank problems: A unified geometric analysis
Rong Ge, Chi Jin, and Yi Zheng · 2017
Cited alongside, same era.
Elad Hoffer, Itay Hubara, and Daniel Soudry · 2017
Cited alongside, same era.
Three factors influencing minima in sgd
Stanislaw Jastrzebski, Zachary Kenton, Devansh Arpit, Nicolas Ballas, Asja Fischer, Yoshua Bengio, and Amos Storkey · 2017
Reconciling modern machine-learning practice and the classical bias–variance trade-off
Mikhail Belkin, Daniel Hsu, Siyuan Ma, and Soumik Mandal · 2019
Later among the works it cites.
Entropy-sgd: Biasing gradient descent into wide valleys
Pratik Chaudhari, Anna Choromanska, Stefano Soatto, Yann LeCun, Carlo Baldassi, Christian Borgs, Jennifer Chayes, Levent Sagun, and Riccardo Zecchina · 2019
Later among the works it cites.
Gpipe: Efficient training of giant neural networks using pipeline parallelism
Yanping Huang, Youlong Cheng, Ankur Bapna, Orhan Firat, Dehao Chen, Mia Chen, HyoukJoong Lee, Jiquan Ngiam, Quoc V Le, Yonghui Wu, et al · 2019
Later among the works it cites.
A function space view of bounded norm infinite width relu nets: The multivariate case
Greg Ongie, Rebecca Willett, Daniel Soudry, and Nathan Srebro · 2019
Later among the works it cites.
How do infinite width bounded norm networks look in function space?
Pedro Savarese, Itay Evron, Daniel Soudry, and Nathan Srebro · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Visualizing the loss landscape of neural nets
Hao Li, Zheng Xu, Gavin Taylor, Christoph Studer, and Tom Goldstein · 2017
Cited alongside, same era.
Exploring generalization in deep learning
Behnam Neyshabur, Srinadh Bhojanapalli, David McAllester, and Nathan Srebro · 2017
Cited alongside, same era.
A bayesian perspective on generalization and stochastic gradient descent
Samuel L Smith and Quoc V Le · 2017
Cited alongside, same era.
Algorithmic regularization in learning deep homogeneous models: Layers are automatically balanced
Simon S Du, Wei Hu, and Jason D Lee · 2018
Cited alongside, same era.
Implicit bias of gradient descent on linear convolutional networks
Suriya Gunasekar, Jason Lee, Daniel Soudry, and Nathan Srebro · 2018
Cited alongside, same era.
Implicit regularization in matrix factorization
Suriya Gunasekar, Blake Woodworth, Srinadh Bhojanapalli, Behnam Neyshabur, and Nathan Srebro · 2018
Cited alongside, same era.
Averaging weights leads to wider optima and better generalization
Pavel Izmailov, Dmitrii Podoprikhin, Timur Garipov, Dmitry Vetrov, and Andrew Gordon Wilson · 2018
Cited alongside, same era.
Efficientnet: Rethinking model scaling for convolutional neural networks
Mingxing Tan and Quoc Le · 2019
Later among the works it cites.
High-dimensional statistics: A non-asymptotic viewpoint
Martin J Wainwright · 2019
Later among the works it cites.
Leave-one-out approach for matrix completion: Primal and dual analysis
Lijun Ding and Yudong Chen · 2020
Later among the works it cites.
Sharpness-aware minimization for efficiently improving generalization
Pierre Foret, Ariel Kleiner, Hossein Mobahi, and Behnam Neyshabur · 2020
Later among the works it cites.
An equivalence between critical points for rank constraints versus low-rank factorizations
Wooseok Ha, Haoyang Liu, and Rina Foygel Barber · 2020
Later among the works it cites.
Compressive sensing with un-trained neural networks: Gradient descent finds a smooth approximation
Reinhard Heckel and Mahdi Soltanolkotabi · 2020
Later among the works it cites.
Big transfer (bit): General visual representation learning
Alexander Kolesnikov, Lucas Beyer, Xiaohua Zhai, Joan Puigcerver, Jessica Yung, Sylvain Gelly, and Neil Houlsby · 2020
Later among the works it cites.
Unique properties of flat minima in deep networks
Rotem Mulayoff and Tomer Michaeli · 2020
Later among the works it cites.
Noisy gradient descent converges to flat minima for nonconvex matrix factorization
Tianyi Liu, Yan Li, Song Wei, Enlu Zhou, and Tuo Zhao · 2021
Later among the works it cites.
Beyond procrustes: Balancing-free gradient descent for asymmetric low-rank matrix sensing
Cong Ma, Yuanxin Li, and Yuejie Chi · 2021
Later among the works it cites.
Diametrical risk minimization: Theory and computations
Matthew D Norton and Johannes O Royset · 2021
Later among the works it cites.
Global convergence of gradient descent for asymmetric low-rank matrix factorization
Tian Ye and Simon S Du · 2021
Later among the works it cites.
Understanding deep learning (still) requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2021
Later among the works it cites.
The role of linear layers in nonlinear interpolating networks
Greg Ongie and Rebecca Willett · 2022
Closest in time.