On the ineffectiveness of variance reduced optimization for deep learning
Aaron Defazio and Leon Bottou · 2019
Later among the works it cites.
On the diffusion approximation of nonconvex stochastic gradient descent
Wenqing Hu, Chris Junchi Li, Lei Li, and Jian-Guo Liu · 2019
Later among the works it cites.
Traditional and heavy-tailed self regularization in neural network models
Charles H Martin and Michael W Mahoney · 2019
Later among the works it cites.
First exit time analysis of stochastic gradient descent under heavy-tailed gradient noise
Thanh Huy Nguyen, Umut Şimşekli, Mert Gürbüzbalaban, and Gaël Richard · 2019
Later among the works it cites.
Non-Gaussianity of stochastic gradient noise
Original
Abhishek Panigrahi, Raghav Somani, Navin Goyal, and Praneeth Netrapalli · 2019
Later among the works it cites.
Finite-time error bounds for linear stochastic approximation and TD learning
Rayadurgam Srikant and Lei Ying · 2019
Later among the works it cites.
The anisotropic noise in stochastic gradient descent: Its behavior of escaping from minima and regularization effects
Zhanxing Zhu, Jingfeng Wu, Bing Yu, Lei Wu, and Jinwen Ma · 2019
Later among the works it cites.
The implicit regularization of stochastic gradient flow for least squares
Alnur Ali, Edgar Dobriban, and Ryan J Tibshirani · 2020
Closest in time.
Stochastic gradient and Langevin processes
Xiang Cheng, Dong Yin, Peter L Bartlett, and Michael I Jordan · 2020
Closest in time.
Quantitative propagation of chaos for SGD in wide neural networks
Valentin De Bortoli, Alain Durmus, Xavier Fontaine, and Umut Şimşekli · 2020
Closest in time.
Bridging the gap between constant step size stochastic gradient descent and Markov chains
Aymeric Dieuleveut, Alain Durmus, and Francis Bach · 2020
Closest in time.
Concentration inequalities for random matrix products
Amelia Henriksen and Rachel Ward · 2020
Closest in time.
Multiplicative noise and heavy tails in stochastic optimization
Original
Liam Hodgkinson and Michael W Mahoney · 2020
Closest in time.
Matrix concentration for products
Original
De Huang, Jonathan Niles-Weed, Joel A. Tropp, and Rachel Ward · 2020
Closest in time.
The large learning rate phase of deep learning: the catapult mechanism
Original
Aitor Lewkowycz, Yasaman Bahri, Ethan Dyer, Jascha Sohl-Dickstein, and Guy Gur-Ari · 2020
Closest in time.
Hausdorff dimension, stochastic differential equations, and generalization in neural networks
Umut Şimşekli, Ozan Sener, George Deligiannidis, and Murat A Erdogdu · 2020
Closest in time.
Why are adaptive methods good for attention models?
Jingzhao Zhang, Sai Praneeth Karimireddy, Andreas Veit, Seungyeon Kim, Sashank Reddi, Sanjiv Kumar, and Suvrit Sra · 2020
Closest in time.
Towards theoretically understanding why SGD generalizes better than ADAM in deep learning
Pan Zhou, Jiashi Feng, Chao Ma, Caiming Xiong, Steven Hoi, and Weinan E · 2020
Closest in time.
Global convergence of stochastic gradient hamiltonian monte carlo for non-convex stochastic optimization: Non-asymptotic performance bounds and momentum-based acceleration
Xuefeng Gao, Mert Gürbüzbalaban, and Lingjiong Zhu · 2021
Closest in time.
A diffusion theory for deep learning dynamics: Stochastic gradient descent exponentially favors flat minima
Zeke Xie, Issei Sato, and Masashi Sugiyama · 2021
Closest in time.