Fetching the paper…
Reading the bibliography…
Stochastic gradient descent (SGD) and its variants are mainstream methods to train deep neural networks.
Stochastic processes in physics and chemistry
Nicolaas Godfried Van Kampen · 1992
Earlier work this paper cites.
Anomalous diffusion in the presence of external forces: exact time-dependent solutions and entropy
Constantino Tsallis and Dirk Jan Bukman · 1995
Earlier work this paper cites.
Anomalous diffusion in the presence of external forces: Exact time-dependent solutions and their thermostatistical basis
Constantino Tsallis and Dirk Jan Bukman · 1996
Earlier work this paper cites.
Pac-bayesian model averaging
David A McAllester · 1999
Earlier work this paper cites.
The tradeoffs of large scale learning
Léon Bottou and Olivier Bousquet · 2008
Earlier work this paper cites.
Making gradient descent optimal for strongly convex stochastic optimization
Alexander Rakhlin, Ohad Shamir, and Karthik Sridharan · 2012
Earlier work this paper cites.
Are power-law distributions an equilibrium distribution or a stationary nonequilibrium distribution?
Ran Guo and Jiulin Du · 2014
Earlier work this paper cites.
Kramers escape rate in overdamped systems with the power-law distribution
Yanjun Zhou and Jiulin Du · 2014
Earlier work this paper cites.
The loss surfaces of multilayer networks
Anna Choromanska, Mikael Henaff, Michael Mathieu, Gérard Ben Arous, and Yann LeCun · 2015
Earlier work this paper cites.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2015
Earlier work this paper cites.
Lenet-5, convolutional neural networks
Yann LeCun et al · 2015
Earlier work this paper cites.
Dual learning for machine translation
Di He, Yingce Xia, Tao Qin, Liwei Wang, Nenghai Yu, Tie-Yan Liu, and Wei-Ying Ma · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
On large-batch training for deep learning: Generalization gap and sharp minima
Nitish Shirish Keskar, Dheevatsa Mudigere, Jorge Nocedal, Mikhail Smelyanskiy, and Ping Tak Peter Tang · 2016
Earlier work this paper cites.
Stochastic modified equations and adaptive stochastic gradient algorithms
Qianxiao Li, Cheng Tai, et al · 2017
Cited alongside, same era.
Stochastic gradient descent as approximate bayesian inference
Stephan Mandt, Matthew D Hoffman, and David M Blei · 2017
Cited alongside, same era.
A bayesian perspective on generalization and stochastic gradient descent
Samuel L Smith and Quoc V Le · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Cited alongside, same era.
Fashion-mnist: a novel image dataset for benchmarking machine learning algorithms
Han Xiao, Kashif Rasul, and Roland Vollgraf · 2017
Cited alongside, same era.
Asymmetric valleys: Beyond sharp and flat local minima
Haowei He, Gao Huang, and Yang Yuan · 2019
Later among the works it cites.
On the diffusion approximation of nonconvex stochastic gradient descent
Wenqing Hu, Chris Junchi Li, Lei Li, and Jian-Guo Liu · 2019
Later among the works it cites.
Traditional and heavy tailed self regularization in neural network models
Michael Mahoney and Charles Martin · 2019
Later among the works it cites.
On the heavy-tailed theory of stochastic gradient descent for deep neural networks
Umut Şimşekli, Mert Gürbüzbalaban, Thanh Huy Nguyen, Gaël Richard, and Levent Sagun · 2019
Later among the works it cites.
A tail-index analysis of stochastic gradient noise in deep neural networks
Umut Simsekli, Levent Sagun, and Mert Gurbuzbalaban · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Stochastic gradient descent performs variational inference, converges to limit cycles for deep networks
Pratik Chaudhari and Stefano Soatto · 2018
Cited alongside, same era.
Essentially no barriers in neural network energy landscape
Felix Draxler, Kambis Veschgini, Manfred Salmhofer, and Fred Hamprecht · 2018
Cited alongside, same era.
Differential equations for modeling asynchronous algorithms
Li He, Qi Meng, Wei Chen, Zhi-Ming Ma, and Tie-Yan Liu · 2018
Cited alongside, same era.
Over-parameterized deep neural networks have no strict local minima for any continuous activations
Dawei Li, Tian Ding, and Ruoyu Sun · 2018
Cited alongside, same era.
Visualizing the loss landscape of neural nets
Hao Li, Zheng Xu, Gavin Taylor, Christoph Studer, and Tom Goldstein · 2018
Cited alongside, same era.
Tianyi Liu, Zhehui Chen, Enlu Zhou, and Tuo Zhao · 2018
Cited alongside, same era.
Energy–entropy competition and the effectiveness of stochastic gradient descent in machine learning
Yao Zhang, Andrew M Saxe, Madhu S Advani, and Alpha A Lee · 2018
Cited alongside, same era.
Jingfeng Wu, Wenqing Hu, Haoyi Xiong, Jun Huan, Vladimir Braverman, and Zhanxing Zhu · 2019
Later among the works it cites.
The multiplicative noise in stochastic gradient descent: Data-dependent regularization, continuous and discrete approximation
Jingfeng Wu, Wenqing Hu, Haoyi Xiong, Jun Huan, and Zhanxing Zhu · 2019
Later among the works it cites.
Toward understanding the importance of noise in training neural networks
Mo Zhou, Tianyi Liu, Yan Li, Dachao Lin, Enlu Zhou, and Tuo Zhao · 2019
Later among the works it cites.
Zhanxing Zhu, Jingfeng Wu, Bing Yu, Lei Wu, and Jinwen Ma · 2019
Later among the works it cites.
The heavy-tail phenomenon in sgd
Mert Gurbuzbalaban, Umut Simsekli, and Lingjiong Zhu · 2020
Closest in time.
Shape matters: Understanding the implicit bias of the noise covariance
Jeff Z HaoChen, Colin Wei, Jason D Lee, and Tengyu Ma · 2020
Closest in time.
Multiplicative noise and heavy tails in stochastic optimization
Liam Hodgkinson and Michael W Mahoney · 2020
Closest in time.
Zeke Xie, Issei Sato, and Masashi Sugiyama · 2020
Closest in time.