Fetching the paper…
Reading the bibliography…
Models with nonlinear architectures/parameterizations such as deep neural networks (DNNs) are well known for their mysteriously good generalization performance at overparameterization.
Communication in the presence of noise
Claude E Shannon · 1984
Earlier work this paper cites.
Reflections after refereeing papers for nips
Leo Breiman · 1995
Earlier work this paper cites.
Adaptive and learning systems for signal processing communications, and control
Vladimir Naumovich Vapnik · 1998
Earlier work this paper cites.
Rademacher and gaussian complexities: Risk bounds and structural results
Peter L Bartlett and Shahar Mendelson · 2002
Earlier work this paper cites.
Exact solutions to the nonlinear dynamics of learning in deep linear neural networks
Andrew M Saxe, James L McClelland, and Surya Ganguli · 2013
Earlier work this paper cites.
Understanding deep learning requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2017
Earlier work this paper cites.
Implicit regularization in matrix factorization
Suriya Gunasekar, Blake E Woodworth, Srinadh Bhojanapalli, Behnam Neyshabur, and Nati Srebro · 2017
Earlier work this paper cites.
Algorithmic regularization in over-parameterized matrix sensing and neural networks with quadratic activations
Yuanzhi Li, Tengyu Ma, and Hongyang Zhang · 2018
Earlier work this paper cites.
Implicit bias of gradient descent on linear convolutional networks
Suriya Gunasekar, Jason D Lee, Daniel Soudry, and Nati Srebro · 2018
Earlier work this paper cites.
How sgd selects the global minima in over-parameterized learning: A dynamical stability perspective
Lei Wu, Chao Ma, and Weinan E · 2018
Earlier work this paper cites.
Gradient descent quantizes relu network features
Hartmut Maennel, Olivier Bousquet, and Sylvain Gelly · 2018
Cited alongside, same era.
Implicit regularization in deep matrix factorization
Sanjeev Arora, Nadav Cohen, Wei Hu, and Yuping Luo · 2019
Cited alongside, same era.
An analytic theory of generalization dynamics and transfer learning in deep linear networks
Andrew K. Lampinen and Surya Ganguli · 2019
Cited alongside, same era.
Training behavior of deep neural network in frequency domain
Zhi-Qin J Xu, Yaoyu Zhang, and Yanyang Xiao · 2019
Cited alongside, same era.
Semi-flat minima and saddle points by embedding neural networks to overparameterization
Kenji Fukumizu, Shoichiro Yamaguchi, Yoh-ichi Mototake, and Mirai Tanaka · 2019
Cited alongside, same era.
Kernel and rich regimes in overparametrized models
Embedding principle of loss landscape of deep neural networks
Yaoyu Zhang, Zhongwang Zhang, Tao Luo, and Zhi-Qin John Xu · 2021
Later among the works it cites.
Phase diagram for two-layer relu neural networks at infinite-width limit
Tao Luo, Zhi-Qin John Xu, Zheng Ma, and Yaoyu Zhang · 2021
Later among the works it cites.
Geometry of the loss landscape in overparameterized neural networks: Symmetries and invariances
Berfin Simsek, François Ged, Arthur Jacot, Francesco Spadaro, Clement Hongler, Wulfram Gerstner, and Johanni Brea · 2021
Later among the works it cites.
Embedding principle: a hierarchical structure of loss landscape of deep neural networks
Yaoyu Zhang, Yuqing Li, Zhongwang Zhang, Tao Luo, and Zhi-Qin John Xu · 2022
Closest in time.
Towards understanding the condensation of neural networks at initial training
Hanxu Zhou, Qixuan Zhou, Tao Luo, Yaoyu Zhang, and Zhi-Qin John Xu · 2022
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Blake Woodworth, Suriya Gunasekar, Jason D Lee, Edward Moroshko, Pedro Savarese, Itay Golan, Daniel Soudry, and Nathan Srebro · 2020
Cited alongside, same era.
Frequency principle: Fourier analysis sheds light on deep neural networks
Zhi-Qin John Xu, Yaoyu Zhang, Tao Luo, Yanyang Xiao, and Zheng Ma · 2020
Cited alongside, same era.
An analytic theory of shallow networks dynamics for hinge loss classification
Franco Pellegrini and Giulio Biroli · 2020
Cited alongside, same era.
A linear frequency principle model to understand the absence of overfitting in neural networks
Yaoyu Zhang, Tao Luo, Zheng Ma, and Zhi-Qin John Xu · 2021
Cited alongside, same era.
The implicit bias of minima stability: A view from function space
Rotem Mulayoff, Tomer Michaeli, and Daniel Soudry · 2021
Cited alongside, same era.
Empirical phase diagram for three-layer neural networks with infinite width
Hanxu Zhou, Qixuan Zhou, Zhenyuan Jin, Tao Luo, Yaoyu Zhang, and Zhi-Qin John Xu · 2022
Closest in time.
Sgd with large step sizes learns sparse features
Maksym Andriushchenko, Aditya Varre, Loucas Pillaud-Vivien, and Nicolas Flammarion · 2022
Closest in time.
Implicit regularization of dropout
Zhongwang Zhang and Zhi-Qin John Xu · 2022
Closest in time.
Embedding principle in depth for the loss landscape analysis of deep neural networks
Zhiwei Bai, Tao Luo, Zhi-Qin John Xu, and Yaoyu Zhang · 2022
Closest in time.