Fetching the paper…
Reading the bibliography…
In this work, we reveal a strong implicit bias of stochastic gradient descent (SGD) that drives overly expressive networks to much simpler subnetworks, thereby dramatically reducing the number of independent parameters, and improving generalization.
Brownian motion in a field of force and the diffusion model of chemical reactions
Hendrik Anthony Kramers · 1940
Earlier work this paper cites.
Stochastic stability and control
Harold J Kushner · 1967
Earlier work this paper cites.
Neural networks and principal component analysis: Learning from examples without local minima
Pierre Baldi and Kurt Hornik · 1989
Earlier work this paper cites.
Local minima and plateaus in hierarchical structures of multilayer perceptrons
Kenji Fukumizu and Shun-ichi Amari · 2000
Earlier work this paper cites.
Algebraic geometrical methods for hierarchical learning machines
Sumio Watanabe · 2001
Earlier work this paper cites.
Singularities affect dynamics of learning in neuromanifolds
Shun-ichi Amari, Hyeyoung Park, and Tomoko Ozeki · 2006
Earlier work this paper cites.
Dynamics of learning near singularities in layered networks
Haikun Wei, Jun Zhang, Florent Cousseau, Tomoko Ozeki, and Shun-ichi Amari · 2007
Earlier work this paper cites.
The effective rank: A measure of effective dimensionality
Olivier Roy and Martin Vetterli · 2007
Earlier work this paper cites.
Dynamics of learning in multilayer perceptrons near singularities
Florent Cousseau, Tomoko Ozeki, and Shun-ichi Amari · 2008
Earlier work this paper cites.
Algebraic geometry and statistical learning theory , volume 25
Sumio Watanabe · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky, Geoffrey Hinton, et al · 2009
Earlier work this paper cites.
Rectified linear units improve restricted boltzmann machines
Vinod Nair and Geoffrey E. Hinton · 2010
Earlier work this paper cites.
How to modify a neural network gradually without changing its input-output functionality
Christopher DiMattina and Kechen Zhang · 2010
Earlier work this paper cites.
The singular values and vectors of low rank perturbations of large rectangular random matrices
Florent Benaych-Georges and Raj Rao Nadakuditi · 2012
Earlier work this paper cites.
Rectifier nonlinearities improve neural network acoustic models
Andrew L Maas, Awni Y Hannun, Andrew Y Ng, et al · 2013
Earlier work this paper cites.
Kramers’ law: Validity, derivations and generalisations
Nils Berglund · 2013
Earlier work this paper cites.
Exact solutions to the nonlinear dynamics of learning in deep linear neural networks
Andrew M Saxe, James L McClelland, and Surya Ganguli · 2013
Earlier work this paper cites.
Stochastic differential equations: an introduction with applications
Bernt Oksendal · 2013
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman · 2014
Earlier work this paper cites.
Fast and accurate deep network learning by exponential linear units (elus)
Djork-Arné Clevert, Thomas Unterthiner, and Sepp Hochreiter · 2015
Earlier work this paper cites.
On large-batch training for deep learning: Generalization gap and sharp minima
Nitish Shirish Keskar, Dheevatsa Mudigere, Jorge Nocedal, Mikhail Smelyanskiy, and Ping Tak Peter Tang · 2016
Cited alongside, same era.
Gaussian error linear units (gelus)
Dan Hendrycks and Kevin Gimpel · 2016
Cited alongside, same era.
A variational analysis of stochastic gradient algorithms
Stephan Mandt, Matthew Hoffman, and David Blei · 2016
Cited alongside, same era.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Cited alongside, same era.
Stochastic modified equations and adaptive stochastic gradient algorithms
Qianxiao Li, Cheng Tai, and Weinan E · 2017
Cited alongside, same era.
High-dimensional dynamics of generalization error in neural networks
Madhu S. Advani, Andrew M. Saxe, and Haim Sompolinsky · 2020
Later among the works it cites.
Understanding deep learning (still) requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2021
Later among the works it cites.
Implicit gradient regularization
David Barrett and Benoit Dherin · 2021
Later among the works it cites.
Shape matters: Understanding the implicit bias of the noise covariance
Jeff Z. HaoChen, Colin Wei, Jason Lee, and Tengyu Ma · 2021
Later among the works it cites.
Label noise sgd provably prefers flat global minimizers
Alex Damian, Tengyu Ma, and Jason D Lee · 2021
Later among the works it cites.
Implicit bias of sgd for diagonal linear networks: a provable benefit of stochasticity
Scott Pesme, Loucas Pillaud-Vivien, and Nicolas Flammarion · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Stanisław Jastrzębski, Zachary Kenton, Devansh Arpit, Nicolas Ballas, Asja Fischer, Yoshua Bengio, and Amos Storkey · 2017
Cited alongside, same era.
An alternative view: When does SGD escape local minima?
Bobby Kleinberg, Yuanzhi Li, and Yang Yuan · 2018
Cited alongside, same era.
Gradient descent happens in a tiny subspace, 2018
Guy Gur-Ari, Daniel A. Roberts, and Ethan Dyer · 2018
Cited alongside, same era.
Searching for activation functions, 2018
Quoc V. Le Prajit Ramachandran, Barret Zoph · 2018
Cited alongside, same era.
Stochastic gradient descent performs variational inference, converges to limit cycles for deep networks
Pratik Chaudhari and Stefano Soatto · 2018
Cited alongside, same era.
Control batch size and learning rate to generalize well: Theoretical and empirical evidence
Fengxiang He, Tongliang Liu, and Dacheng Tao · 2019
Cited alongside, same era.
The anisotropic noise in stochastic gradient descent: Its behavior of escaping from sharp minima and regularization effects
Zhanxing Zhu, Jingfeng Wu, Bing Yu, Lei Wu, and Jinwen Ma · 2019
Cited alongside, same era.
Later among the works it cites.
The low-rank simplicity bias in deep networks
Minyoung Huh, Hossein Mobahi, Richard Zhang, Brian Cheung, Pulkit Agrawal, and Phillip Isola · 2021
Later among the works it cites.
Sgd with a constant large learning rate can converge to local maxima
Liu Ziyin, Botao Li, James B Simon, and Masahito Ueda · 2021
Later among the works it cites.
Geometry of the loss landscape in overparameterized neural networks: Symmetries and invariances
Berfin Simsek, François Ged, Arthur Jacot, Francesco Spadaro, Clement Hongler, Wulfram Gerstner, and Johanni Brea · 2021
Later among the works it cites.
Stochastic training is not necessary for generalization
Jonas Geiping, Micah Goldblum, Phil Pope, Michael Moeller, and Tom Goldstein · 2022
Later among the works it cites.
What happens after SGD reaches zero loss? –a mathematical framework
Zhiyuan Li, Tianhao Wang, and Sanjeev Arora · 2022
Later among the works it cites.
Implicit bias of the step size in linear diagonal neural networks
Mor Shpigel Nacson, Kavya Ravichandran, Nathan Srebro, and Daniel Soudry · 2022
Later among the works it cites.
Sgd with large step sizes learns sparse features
Maksym Andriushchenko, Aditya Varre, Loucas Pillaud-Vivien, and Nicolas Flammarion · 2022
Later among the works it cites.
Label noise (stochastic) gradient descent implicitly solves the lasso for quadratic parametrisation
Loucas Pillaud Vivien, Julien Reygner, and Nicolas Flammarion · 2022
Later among the works it cites.
Exact solutions of a deep linear network
Liu Ziyin, Botao Li, and Xiangming Meng · 2022
Later among the works it cites.
Sgd and weight decay provably induce a low-rank bias in neural networks
Tomer Galanti, Zachary S Siegel, Aparna Gupte, and Tomaso Poggio · 2022
Later among the works it cites.
Implicit bias of large depth networks: a notion of rank for nonlinear functions
Arthur Jacot · 2022
Later among the works it cites.
The asymmetric maximum margin bias of quasi-homogeneous neural networks
Daniel Kunin, Atsushi Yamamura, Chao Ma, and Surya Ganguli · 2023
Closest in time.
Implicit bias of sgd in l 2 l_{2} -regularized linear dnns: One-way jumps from high to low rank
Zihan Wang and Arthur Jacot · 2023
Closest in time.