Fetching the paper…
Reading the bibliography…
Gradient clipping is a standard training technique used in deep learning applications such as large-scale language modeling to mitigate exploding gradients.
THE MNIST DATABASE of handwritten digits
Yann LeCun, Corinna Cortes, and Christopher J. C. Burges · 1998
Earlier work this paper cites.
Introductory lectures on convex optimization: A basic course , volume 87
Yurii Nesterov · 2003
Earlier work this paper cites.
A stochastic gradient method with an exponential convergence _rate for finite training sets
Nicolas Roux, Mark Schmidt, and Francis Bach · 2012
Earlier work this paper cites.
Stochastic dual coordinate ascent methods for regularized loss minimization
Shai Shalev-Shwartz and Tong Zhang · 2012
Earlier work this paper cites.
Saga: A fast incremental gradient method with support for non-strongly convex composite objectives
Aaron Defazio, Francis Bach, and Simon Lacoste-Julien · 2014
Earlier work this paper cites.
Variance reduction for faster non-convex optimization
Zeyuan Allen-Zhu and Elad Hazan · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Stochastic variance reduction for nonconvex optimization
Sashank J Reddi, Ahmed Hefny, Suvrit Sra, Barnabás Póczos, and Alex Smola · 2016
Earlier work this paper cites.
Convolutional sequence to sequence learning
Jonas Gehring, Michael Auli, David Grangier, Denis Yarats, and Yann N Dauphin · 2017
Earlier work this paper cites.
Non-convex finite-sum optimization via scsg methods
Lihua Lei, Cheng Ju, Jianbo Chen, and Michael I Jordan · 2017
Earlier work this paper cites.
Regularizing and optimizing lstm language models
Stephen Merity, Nitish Shirish Keskar, and Richard Socher · 2017
Earlier work this paper cites.
Stochastic recursive gradient algorithm for nonconvex optimization
Lam M Nguyen, Jie Liu, Katya Scheinberg, and Martin Takáč · 2017
Cited alongside, same era.
Spider: Near-optimal non-convex optimization via stochastic path-integrated differential estimator
Cong Fang, Chris Junchi Li, Zhouchen Lin, and Tong Zhang · 2018
Cited alongside, same era.
Deep contextualized word representations. arxiv 2018
ME Peters, M Neumann, M Iyyer, M Gardner, C Clark, K Lee, and L Zettlemoyer · 2018
Cited alongside, same era.
Momentum-based variance reduction in non-convex sgd
Ashok Cutkosky and Francesco Orabona · 2019
Cited alongside, same era.
On the ineffectiveness of variance reduced optimization for deep learning
Aaron Defazio and Léon Bottou · 2019
Cited alongside, same era.
Understanding gradient clipping in incremental gradient methods
Jiang Qian, Yuren Wu, Bojin Zhuang, Shaojun Wang, and Jing Xiao · 2021
Later among the works it cites.
On the convergence and improvement of stochastic normalized gradient descent
Shen-Yi Zhao, Yin-Peng Xie, and Wu-Jun Li · 2021
Later among the works it cites.
Lower bounds for non-convex stochastic optimization
Yossi Arjevani, Yair Carmon, John C Duchi, Dylan J Foster, Nathan Srebro, and Blake Woodworth · 2022
Later among the works it cites.
Robustness to unbounded smoothness of generalized signsgd
Michael Crawshaw, Mingrui Liu, Francesco Orabona, Wei Zhang, and Zhenxun Zhuang · 2022
Later among the works it cites.
Adaptive stochastic variance reduction for non-convex finite-sum minimization
Ali Kavis, Stratis Skoulakis, Kimon Antonakopoulos, Leello Tadesse Dadi, and Volkan Cevher · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Spiderboost and momentum: Faster variance reduction algorithms
Zhe Wang, Kaiyi Ji, Yi Zhou, Yingbin Liang, and Vahid Tarokh · 2019
Cited alongside, same era.
Why gradient clipping accelerates training: A theoretical justification for adaptivity
Jingzhao Zhang, Tianxing He, Suvrit Sra, and Ali Jadbabaie · 2019
Cited alongside, same era.
Adaptivity of stochastic gradient methods for nonconvex optimization
Samuel Horváth, Lihua Lei, Peter Richtárik, and Michael I Jordan · 2020
Cited alongside, same era.
Proxsarah: An efficient algorithmic framework for stochastic composite nonconvex optimization
Nhan H Pham, Lam M Nguyen, Dzung T Phan, and Quoc Tran-Dinh · 2020
Cited alongside, same era.
Improved analysis of clipping algorithms for non-convex optimization
Bohang Zhang, Jikai Jin, Cong Fang, and Liwei Wang · 2020
Cited alongside, same era.
Stochastic nested variance reduction for nonconvex optimization
Dongruo Zhou, Pan Xu, and Quanquan Gu · 2020
Cited alongside, same era.
Page: A simple and optimal probabilistic gradient estimator for nonconvex optimization
Zhize Li, Hongyan Bao, Xiangliang Zhang, and Peter Richtárik
Cited in the paper.
A communication-efficient distributed gradient clipping algorithm for training deep neural networks
Mingrui Liu, Zhenxun Zhuang, Yunwei Lei, and Chunyang Liao · 2022
Later among the works it cites.
Convergence of stein variational gradient descent under a weaker smoothness condition
Lukang Sun, Avetik Karagulyan, and Peter Richtarik · 2022
Later among the works it cites.
Bohan Wang, Yushun Zhang, Huishuai Zhang, Qi Meng, Zhi-Ming Ma, Tie-Yan Liu, and Wei Chen · 2022
Later among the works it cites.
Differentially private learning with per-sample adaptive clipping
Tianyu Xia, Shuheng Shen, Su Yao, Xinyi Fu, Ke Xu, Xiaolong Xu, Xing Fu, and Weiqiang Wang · 2022
Later among the works it cites.
Normalized/clipped sgd with perturbation for differentially private non-convex optimization
Xiaodong Yang, Huishuai Zhang, Wei Chen, and Tie-Yan Liu · 2022
Later among the works it cites.
Beyond uniform smoothness: A stopped analysis of adaptive sgd
Matthew Faw, Litu Rout, Constantine Caramanis, and Sanjay Shakkottai · 2023
Closest in time.