Fetching the paper…
Reading the bibliography…
Gradient clipping is commonly used in training deep neural networks partly due to its practicability in relieving the exploding gradient problem.
Note on the derivatives with respect to a parameter of the solutions of a system of differential equations
T. H. Gronwall · 1919
Earlier work this paper cites.
On complexity of finding stationary points of nonsmooth nonconvex functions
Jingzhao Zhang, Hongzhou Lin, Suvrit Sra, and Ali Jadbabaie · 2002
Earlier work this paper cites.
Recurrent neural network based language model
Tomáš Mikolov, Martin Karafiát, Lukáš Burget, Jan Černockỳ, and Sanjeev Khudanpur · 2010
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
John Duchi, Elad Hazan, and Yoram Singer · 2011
Earlier work this paper cites.
Statistical language models based on neural networks
Tomáš Mikolov · 2012
Earlier work this paper cites.
Understanding the exploding gradient problem
Razvan Pascanu, Tomas Mikolov, and Yoshua Bengio · 2012
Earlier work this paper cites.
Stochastic first-and zeroth-order methods for nonconvex stochastic programming
Saeed Ghadimi and Guanghui Lan · 2013
Earlier work this paper cites.
On the difficulty of training recurrent neural networks
Razvan Pascanu, Tomas Mikolov, and Yoshua Bengio · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Deep learning with differential privacy
Martin Abadi, Andy Chu, Ian Goodfellow, H Brendan McMahan, Ilya Mironov, Kunal Talwar, and Li Zhang · 2016
Cited alongside, same era.
Deep learning
Ian Goodfellow, Yoshua Bengio, and Aaron Courville · 2016
Cited alongside, same era.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Cited alongside, same era.
The power of normalization: Faster evasion of saddle points
Kfir Y Levy · 2016
Cited alongside, same era.
Accurate, large minibatch sgd: Training imagenet in 1 hour
Priya Goyal, Piotr Dollár, Ross Girshick, Pieter Noordhuis, Lukasz Wesolowski, Aapo Kyrola, Andrew Tulloch, Yangqing Jia, and Kaiming He · 2017
Cited alongside, same era.
An alternative view: When does sgd escape local minima?
Bobby Kleinberg, Yuanzhi Li, and Yang Yuan · 2018
Later among the works it cites.
Quasi-hyperbolic momentum and adam for deep learning
Jerry Ma and Denis Yarats · 2018
Later among the works it cites.
Lower bounds for non-convex stochastic optimization
Yossi Arjevani, Yair Carmon, John C Duchi, Dylan J Foster, Nathan Srebro, and Blake Woodworth · 2019
Later among the works it cites.
Lower bounds for finding stationary points i
Yair Carmon, John C Duchi, Oliver Hinder, and Aaron Sidford · 2019
Later among the works it cites.
The complexity of finding stationary points with stochastic gradient descent
Yoel Drori and Ohad Shamir · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Gao Huang, Yixuan Li, Geoff Pleiss, Zhuang Liu, John E Hopcroft, and Kilian Q Weinberger · 2017
Cited alongside, same era.
Regularizing and optimizing lstm language models
Stephen Merity, Nitish Shirish Keskar, and Richard Socher · 2017
Cited alongside, same era.
Scaling sgd batch size to 32k for imagenet training
Yang You, Igor Gitman, and Boris Ginsburg · 2017
Cited alongside, same era.
Spider: Near-optimal non-convex optimization via stochastic path-integrated differential estimator
Cong Fang, Chris Junchi Li, Zhouchen Lin, and Tong Zhang · 2018
Cited alongside, same era.
Why gradient clipping accelerates training: A theoretical justification for adaptivity
Jingzhao Zhang, Tianxing He, Suvrit Sra, and Ali Jadbabaie
Cited in the paper.
Why adam beats sgd for attention models
Jingzhao Zhang, Sai Praneeth Karimireddy, Andreas Veit, Seungyeon Kim, Sashank J Reddi, Sanjiv Kumar, and Suvrit Sra · 2019
Later among the works it cites.
Momentum improves normalized sgd
Ashok Cutkosky and Harsh Mehta · 2020
Closest in time.
Stochastic optimization with heavy-tailed noise via accelerated gradient clipping
Eduard Gorbunov, Marina Danilova, and Alexander Gasnikov · 2020
Closest in time.
Can gradient clipping mitigate label noise
Aditya Krishna Menon, Ankit Singh Rawat, Sashank J Reddi, and Sanjiv Kumar · 2020
Closest in time.