Fetching the paper…
Reading the bibliography…
It is generally recognized that finite learning rate (LR), in contrast to infinitesimal LR, is important for good generalization in real-life deep nets.
The euler scheme for lévy driven stochastic differential equations
Philip Protter, Denis Talay, et al · 1997
Earlier work this paper cites.
Lévy processes and infinitely divisible distributions
Sato Ken-Iti · 1999
Earlier work this paper cites.
Numerical Solution of Stochastic Differential Equations
P.E. Kloeden and E. Platen · 2011
Earlier work this paper cites.
Reading digits in natural images with unsupervised feature learning
Yuval Netzer, Tao Wang, Adam Coates, Alessandro Bissacco, Bo Wu, and Andrew Y Ng · 2011
Earlier work this paper cites.
Practical recommendations for gradient-based training of deep architectures
Yoshua Bengio · 2012
Earlier work this paper cites.
Efficient BackProp , pages 9–48
Yann A. LeCun, Léon Bottou, Genevieve B. Orr, and Klaus-Robert Müller · 2012
Earlier work this paper cites.
One weird trick for parallelizing convolutional neural networks
Alex Krizhevsky · 2014
Earlier work this paper cites.
Batch normalization: accelerating deep network training by reducing internal covariate shift
Sergey Ioffe and Christian Szegedy · 2015
Earlier work this paper cites.
On estimating the tail index and the spectral measure of multivariate α \alpha -stable distributions
Mohammad Mohammadi, Adel Mohammadpour, and Hiroaki Ogata · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Sgdr: Stochastic gradient descent with warm restarts
Ilya Loshchilov and Frank Hutter · 2016
Earlier work this paper cites.
Sharp minima can generalize for deep nets
Laurent Dinh, Razvan Pascanu, Samy Bengio, and Yoshua Bengio · 2017
Earlier work this paper cites.
Accurate, large minibatch sgd: Training imagenet in 1 hour
Priya Goyal, Piotr Dollár, Ross Girshick, Pieter Noordhuis, Lukasz Wesolowski, Aapo Kyrola, Andrew Tulloch, Yangqing Jia, and Kaiming He · 2017
Earlier work this paper cites.
Train longer, generalize better: closing the generalization gap in large batch training of neural networks
Elad Hoffer, Itay Hubara, and Daniel Soudry · 2017
Earlier work this paper cites.
Three factors influencing minima in sgd
Stanisław Jastrzebski, Zachary Kenton, Devansh Arpit, Nicolas Ballas, Asja Fischer, Yoshua Bengio, and Amos Storkey · 2017
Earlier work this paper cites.
On large-batch training for deep learning: Generalization gap and sharp minima
Nitish Shirish Keskar, Dheevatsa Mudigere, Jorge Nocedal, Mikhail Smelyanskiy, and Ping Tak Peter Tang · 2017
Cited alongside, same era.
Stochastic modified equations and adaptive stochastic gradient algorithms
Qianxiao Li, Cheng Tai, and E Weinan · 2017
Cited alongside, same era.
Stochastic gradient descent as approximate bayesian inference
Stephan Mandt, Matthew D Hoffman, and David M Blei · 2017
Cited alongside, same era.
Cyclical learning rates for training neural networks
L. N. Smith · 2017
Cited alongside, same era.
Understanding batch normalization
Nils Bjorck, Carla P Gomes, Bart Selman, and Kilian Q Weinberger · 2018
Cited alongside, same era.
Stochastic gradient descent performs variational inference, converges to limit cycles for deep networks
Pratik Chaudhari and Stefano Soatto · 2018
Measuring the effects of data parallelism on neural network training
Christopher J. Shallue, Jaehoon Lee, Joseph Antognini, Jascha Sohl-Dickstein, Roy Frostig, and George E. Dahl · 2019
Later among the works it cites.
A tail-index analysis of stochastic gradient noise in deep neural networks
Umut Simsekli, Levent Sagun, and Mert Gurbuzbalaban · 2019
Later among the works it cites.
The anisotropic noise in stochastic gradient descent: Its behavior of escaping from sharp minima and regularization effects
Zhanxing Zhu, Jingfeng Wu, Bing Yu, Lei Wu, and Jinwen Ma · 2019
Later among the works it cites.
Experiment tracking with weights and biases, 2020
Lukas Biewald · 2020
Later among the works it cites.
Stochastic gradient and langevin processes
Xiang Cheng, Dong Yin, Peter Bartlett, and Michael Jordan · 2020
Later among the works it cites.
Neural mechanics: Symmetry and broken conservation laws in deep learning dynamics, 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Fix your classifier: the marginal value of training the last weight layer
Elad Hoffer, Itay Hubara, and Daniel Soudry · 2018
Cited alongside, same era.
Neural tangent kernel: Convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel, and Clement Hongler · 2018
Cited alongside, same era.
A bayesian perspective on generalization and stochastic gradient descent
Samuel L. Smith and Quoc V. Le · 2018
Cited alongside, same era.
Don’t decay the learning rate, increase the batch size
Samuel L. Smith, Pieter-Jan Kindermans, and Quoc V. Le · 2018
Cited alongside, same era.
Yuxin Wu and Kaiming He · 2018
Cited alongside, same era.
Learning and generalization in overparameterized neural networks, going beyond two layers
Zeyuan Allen-Zhu, Yuanzhi Li, and Yingyu Liang · 2019
Cited alongside, same era.
Daniel Kunin, Javier Sagastuy-Brena, Surya Ganguli, Daniel L. K. Yamins, and Hidenori Tanaka · 2020
Later among the works it cites.
Reconciling modern deep learning with traditional optimization analyses: The intrinsic learning rate
Zhiyuan Li, Kaifeng Lyu, and Sanjeev Arora · 2020
Later among the works it cites.
On learning rates and schrödinger operators
Bin Shi, Weijie J Su, and Michael I Jordan · 2020
Later among the works it cites.
On the generalization benefit of noise in stochastic gradient descent, 2020
Samuel L. Smith, Erich Elsen, and Soham De · 2020
Later among the works it cites.
On the noisy gradient descent that generalizes as SGD
Jingfeng Wu, Wenqing Hu, Haoyi Xiong, Jun Huan, Vladimir Braverman, and Zhanxing Zhu · 2020
Later among the works it cites.
Why {adam} beats {sgd} for attention models, 2020
Jingzhao Zhang, Sai Praneeth Karimireddy, Andreas Veit, Seungyeon Kim, Sashank J Reddi, Sanjiv Kumar, and Suvrit Sra · 2020
Later among the works it cites.
Towards theoretically understanding why sgd generalizes better than adam in deep learning, 2020
Pan Zhou, Jiashi Feng, Chao Ma, Caiming Xiong, Steven HOI, and Weinan E · 2020
Later among the works it cites.
On the origin of implicit regularization in stochastic gradient descent, 2021
Samuel L. Smith, Benoit Dherin, David G. T. Barrett, and Soham De · 2021
Closest in time.
A diffusion theory for deep learning dynamics: Stochastic gradient descent exponentially favors flat minima, 2021
Zeke Xie, Issei Sato, and Masashi Sugiyama · 2021
Closest in time.