Fetching the paper…
Reading the bibliography…
In this paper, we provide a theoretical study of noise geometry for minibatch stochastic gradient descent (SGD), a phenomenon where noise aligns favorably with the geometry of local landscape.
Flat minima
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
Learning multiple layers of features from tiny images, 2009
Alex Krizhevsky and Geoffrey Hinton · 2009
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman · 2015
Earlier work this paper cites.
Deep learning
Ian Goodfellow, Yoshua Bengio, and Aaron Courville · 2016
Earlier work this paper cites.
Introduction to online convex optimization
Elad Hazan et al · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
On large-batch training for deep learning: Generalization gap and sharp minima
N. S. Keskar, D. Mudigere, J. Nocedal, M. Smelyanskiy, and P. T. P. Tang · 2017
Earlier work this paper cites.
SGDR: Stochastic gradient descent with warm restarts
Ilya Loshchilov and Frank Hutter · 2017
Earlier work this paper cites.
Cyclical learning rates for training neural networks
Leslie N Smith · 2017
Earlier work this paper cites.
Towards understanding generalization of deep learning: Perspective of loss landscapes
Lei Wu, Zhanxing Zhu, and Weinan E · 2017
Earlier work this paper cites.
Understanding deep learning requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2017
Earlier work this paper cites.
On exponential convergence of sgd in non-convex over-parametrized learning
Raef Bassily, Mikhail Belkin, and Siyuan Ma · 2018
Earlier work this paper cites.
Optimization methods for large-scale machine learning
Léon Bottou, Frank E Curtis, and Jorge Nocedal · 2018
Earlier work this paper cites.
Escaping saddles with stochastic gradients
Hadi Daneshmand, Jonas Kohler, Aurelien Lucchi, and Thomas Hofmann · 2018
Earlier work this paper cites.
Snapshot ensembles: Train 1, get M {M} for free
Gao Huang, Yixuan Li, Geoff Pleiss, Zhuang Liu, John E Hopcroft, and Kilian Q Weinberger · 2018
Earlier work this paper cites.
An alternative view: When does SGD escape local minima?
Bobby Kleinberg, Yuanzhi Li, and Yang Yuan · 2018
Earlier work this paper cites.
High-dimensional probability: An introduction with applications in data science , volume 47
Roman Vershynin · 2018
Earlier work this paper cites.
How SGD selects the global minima in over-parameterized learning: A dynamical stability perspective
Lei Wu, Chao Ma, and Weinan E · 2018
Cited alongside, same era.
Implicit regularization in deep matrix factorization
Sanjeev Arora, Nadav Cohen, Wei Hu, and Yuping Luo · 2019
Cited alongside, same era.
Bootstrapping upper confidence bound
Botao Hao, Yasin Abbasi Yadkori, Zheng Wen, and Guang Cheng · 2019
Cited alongside, same era.
A tail-index analysis of stochastic gradient noise in deep neural networks
Umut Simsekli, Levent Sagun, and Mert Gurbuzbalaban · 2019
Cited alongside, same era.
Residual learning without normalization via better initialization
Hongyi Zhang, Yann N. Dauphin, and Tengyu Ma · 2019
Cited alongside, same era.
The anisotropic noise in stochastic gradient descent: Its behavior of escaping from sharp minima and regularization effects
On the validity of modeling SGD with stochastic differential equations (SDEs)
Zhiyuan Li, Sadhika Malladi, and Sanjeev Arora · 2021
Later among the works it cites.
Noise and fluctuation of finite learning rate stochastic gradient descent
Kangqiao Liu, Liu Ziyin, and Masahito Ueda · 2021
Later among the works it cites.
Implicit bias of SGD for diagonal linear networks: a provable benefit of stochasticity
Scott Pesme, Loucas Pillaud-Vivien, and Nicolas Flammarion · 2021
Later among the works it cites.
Stochastic gradient descent with noise of machine learning type. part II: Continuous time analysis
Stephan Wojtowytsch · 2021
Later among the works it cites.
Understanding the unstable convergence of gradient descent
Kwangjun Ahn, Jingzhao Zhang, and Suvrit Sra · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Zhanxing Zhu, Jingfeng Wu, Bing Yu, Lei Wu, and Jinwen Ma · 2019
Cited alongside, same era.
Gradient descent on neural networks typically occurs at the edge of stability
Jeremy Cohen, Simran Kaur, Yuanzhi Li, J Zico Kolter, and Ameet Talwalkar · 2020
Cited alongside, same era.
Fast convergence of stochastic subgradient method under interpolation
Huang Fang, Zhenan Fan, and Michael Friedlander · 2020
Cited alongside, same era.
On the interplay between noise and curvature and its effect on optimization and generalization
Valentin Thomas, Fabian Pedregosa, Bart Merriënboer, Pierre-Antoine Manzagol, Yoshua Bengio, and Nicolas Le Roux · 2020
Cited alongside, same era.
Kernel and rich regimes in overparametrized models
Blake Woodworth, Suriya Gunasekar, Jason D Lee, Edward Moroshko, Pedro Savarese, Itay Golan, Daniel Soudry, and Nathan Srebro · 2020
Cited alongside, same era.
On the noisy gradient descent that generalizes as SGD
Jingfeng Wu, Wenqing Hu, Haoyi Xiong, Jun Huan, Vladimir Braverman, and Zhanxing Zhu · 2020
Cited alongside, same era.
A diffusion theory for deep learning dynamics: Stochastic gradient descent exponentially favors flat minima
Zeke Xie, Issei Sato, and Masashi Sugiyama · 2020
Cited alongside, same era.
Jian-Feng Cai, Meng Huang, Dong Li, and Yang Wang · 2022
Later among the works it cites.
The power of adaptivity in SGD: Self-tuning step sizes with unbounded gradients and affine variance
Matthew Faw, Isidoros Tziotis, Constantine Caramanis, Aryan Mokhtari, Sanjay Shakkottai, and Rachel Ward · 2022
Later among the works it cites.
What happens after SGD reaches zero loss? –a mathematical framework
Zhiyuan Li, Tianhao Wang, and Sanjeev Arora · 2022
Later among the works it cites.
Beyond the quadratic approximation: The multiscale structure of neural network loss landscapes
Chao Ma, Daniel Kunin, Lei Wu, and Lexing Ying · 2022
Later among the works it cites.
Power-law escape rate of SGD
Takashi Mori, Liu Ziyin, Kangqiao Liu, and Masahito Ueda · 2022
Later among the works it cites.
The alignment property of SGD noise and how it helps select flat minima: A stability analysis
Lei Wu, Mingze Wang, and Weijie J Su · 2022
Later among the works it cites.
Adaptive inertia: Disentangling the effects of adaptive learning rate and momentum
Zeke Xie, Xinrui Wang, Huishuai Zhang, Issei Sato, and Masashi Sugiyama · 2022
Later among the works it cites.
Strength of minibatch noise in SGD
Liu Ziyin, Kangqiao Liu, Takashi Mori, and Masahito Ueda · 2022
Later among the works it cites.
Aiming towards the minimizers: fast convergence of SGD for overparametrized problems
Chaoyue Liu, Dmitriy Drusvyatskiy, Mikhail Belkin, Damek Davis, and Yi-An Ma · 2023
Closest in time.
Stochastic gradient descent with noise of machine learning type part i: Discrete time analysis
Stephan Wojtowytsch · 2023
Closest in time.
The implicit regularization of dynamical stability in stochastic gradient descent
Lei Wu and Weijie J Su · 2023
Closest in time.