Fetching the paper…
Reading the bibliography…
The phenomenon that stochastic gradient descent (SGD) favors flat minima has played a critical role in understanding the implicit regularization of SGD.
Theory of reproducing kernels
Nachman Aronszajn · 1950
Earlier work this paper cites.
Flat minima
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Yann LeCun, Léon Bottou, Yoshua Bengio, and Patrick Haffner · 1998
Earlier work this paper cites.
Kernel methods for deep learning
Youngmin Cho and Lawrence K Saul · 2009
Earlier work this paper cites.
Stochastic methods
Crispin Gardiner · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images, 2009
Alex Krizhevsky and Geoffrey Hinton · 2009
Earlier work this paper cites.
Geometrical methods in the theory of ordinary differential equations
Vladimir Igorevich Arnold · 2012
Earlier work this paper cites.
What is an RKHS?
Dino Sejdinovic and Arthur Gretton · 2012
Earlier work this paper cites.
Probability in Banach Spaces: Isoperimetry and processes
Michel Ledoux and Michel Talagrand · 2013
Earlier work this paper cites.
In search of the real inductive bias: On the role of implicit regularization in deep learning
Behnam Neyshabur, Ryota Tomioka, and Nathan Srebro · 2014
Earlier work this paper cites.
Understanding machine learning: From theory to algorithms
Shai Shalev-Shwartz and Shai Ben-David · 2014
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Breaking the curse of dimensionality with convex neural networks
Francis Bach · 2017
Earlier work this paper cites.
Three factors influencing minima in SGD
Stanisław Jastrzębski, Zachary Kenton, Devansh Arpit, Nicolas Ballas, Asja Fischer, Yoshua Bengio, and Amos Storkey · 2017
Earlier work this paper cites.
On large-batch training for deep learning: Generalization gap and sharp minima
N. S. Keskar, D. Mudigere, J. Nocedal, M. Smelyanskiy, and P. T. P. Tang · 2017
Earlier work this paper cites.
Stochastic modified equations and adaptive stochastic gradient algorithms
Qianxiao Li, Cheng Tai, and Weinan E · 2017
Earlier work this paper cites.
Towards understanding generalization of deep learning: Perspective of loss landscapes
Lei Wu, Zhanxing Zhu, and Weinan E · 2017
Earlier work this paper cites.
Understanding deep learning requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2017
Cited alongside, same era.
Averaging weights leads to wider optima and better generalization
P Izmailov, AG Wilson, D Podoprikhin, D Vetrov, and T Garipov · 2018
Cited alongside, same era.
Neural tangent kernel: Convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel, and Clément Hongler · 2018
Cited alongside, same era.
How SGD selects the global minima in over-parameterized learning: A dynamical stability perspective
Lei Wu, Chao Ma, and Weinan E · 2018
Cited alongside, same era.
Implicit regularization in deep matrix factorization
Sanjeev Arora, Nadav Cohen, Wei Hu, and Yuping Luo · 2019
Cited alongside, same era.
The break-even point on optimization trajectories of deep neural networks
A diffusion theory for deep learning dynamics: Stochastic gradient descent exponentially favors flat minima
Zeke Xie, Issei Sato, and Masashi Sugiyama · 2020
Later among the works it cites.
Towards theoretically understanding why SGD generalizes better than Adam in deep learning
Pan Zhou, Jiashi Feng, Chao Ma, Caiming Xiong, Steven Chu Hong Hoi, et al · 2020
Later among the works it cites.
On the implicit bias of initialization shape: Beyond infinitesimal mirror descent
Shahar Azulay, Edward Moroshko, Mor Shpigel Nacson, Blake E Woodworth, Nathan Srebro, Amir Globerson, and Daniel Soudry · 2021
Later among the works it cites.
Label noise SGD provably prefers flat global minimizers
Alex Damian, Tengyu Ma, and Jason D Lee · 2021
Later among the works it cites.
The inverse variance–flatness relation in stochastic gradient descent is critical for finding flat minima
Yu Feng and Yuhai Tu · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Stanislaw Jastrzebski, Maciej Szymczak, Stanislav Fort, Devansh Arpit, Jacek Tabor, Kyunghyun Cho, and Krzysztof Geras · 2019
Cited alongside, same era.
Fisher-Rao metric, geometry, and complexity of neural networks
Tengyuan Liang, Tomaso Poggio, Alexander Rakhlin, and James Stokes · 2019
Cited alongside, same era.
A tail-index analysis of stochastic gradient noise in deep neural networks
Umut Simsekli, Levent Sagun, and Mert Gurbuzbalaban · 2019
Cited alongside, same era.
Residual learning without normalization via better initialization
Hongyi Zhang, Yann N. Dauphin, and Tengyu Ma · 2019
Cited alongside, same era.
The anisotropic noise in stochastic gradient descent: Its behavior of escaping from sharp minima and regularization effects
Zhanxing Zhu, Jingfeng Wu, Bing Yu, Lei Wu, and Jinwen Ma · 2019
Cited alongside, same era.
Implicit regularization for deep neural networks driven by an Ornstein-Uhlenbeck like process
Guy Blanc, Neha Gupta, Gregory Valiant, and Paul Valiant · 2020
Cited alongside, same era.
Gradient descent on neural networks typically occurs at the edge of stability
Jeremy Cohen, Simran Kaur, Yuanzhi Li, J Zico Kolter, and Ameet Talwalkar · 2020
Cited alongside, same era.
Shape matters: Understanding the implicit bias of the noise covariance
Jeff Z HaoChen, Colin Wei, Jason Lee, and Tengyu Ma · 2021
Later among the works it cites.
On the validity of modeling SGD with stochastic differential equations (SDEs)
Zhiyuan Li, Sadhika Malladi, and Sanjeev Arora · 2021
Later among the works it cites.
On linear stability of SGD and input-smoothness of neural networks
Chao Ma and Lexing Ying · 2021
Later among the works it cites.
Logarithmic landscape and power-law escape rate of SGD
Takashi Mori, Liu Ziyin, Kangqiao Liu, and Masahito Ueda · 2021
Later among the works it cites.
The implicit bias of minima stability: A view from function space
Rotem Mulayoff, Tomer Michaeli, and Daniel Soudry · 2021
Later among the works it cites.
Implicit bias of SGD for diagonal linear networks: a provable benefit of stochasticity
Scott Pesme, Loucas Pillaud-Vivien, and Nicolas Flammarion · 2021
Later among the works it cites.
Neurashed: A phenomenological model for imitating deep learning training
Weijie Su · 2021
Later among the works it cites.
Stochastic gradient descent with noise of machine learning type. part II: Continuous time analysis
Stephan Wojtowytsch · 2021
Later among the works it cites.
What happens after SGD reaches zero loss? –a mathematical framework
Zhiyuan Li, Tianhao Wang, and Sanjeev Arora · 2022
Closest in time.
Eliminating sharp minima from SGD with truncated heavy-tailed noise
Xingyu Wang, Sewoong Oh, and Chang-Han Rhee · 2022
Closest in time.
A spectral-based analysis of the separation between two-layer neural networks and linear methods
Lei Wu and Jihao Long · 2022
Closest in time.
Strength of minibatch noise in SGD
Liu Ziyin, Kangqiao Liu, Takashi Mori, and Masahito Ueda · 2022
Closest in time.