Fetching the paper…
Reading the bibliography…
Traditional analyses in non-convex optimization typically rely on the smoothness assumption, namely requiring the gradients to be Lipschitz.
A stochastic approximation method
Herbert Robbins and Sutton Monro · 1951
Earlier work this paper cites.
Stochastic quasigradient methods
Yuri Ermoliev · 1988
Earlier work this paper cites.
A direct adaptive method for faster backpropagation learning: The RPROP algorithm
Martin Riedmiller and Heinrich Braun · 1993
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
On the projected subgradient method for nonsmooth convex optimization in a Hilbert space
Ya I. Alber, Alfredo N. Iusem, and Mikhail V. Solodov · 1998
Earlier work this paper cites.
Convex optimization
Stephen Boyd and Lieven Vandenberghe · 2004
Earlier work this paper cites.
Adaptive bound optimization for online convex optimization
H Brendan McMahan and Matthew Streeter · 2010
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
John Duchi, Elad Hazan, and Yoram Singer · 2011
Earlier work this paper cites.
Deep neural networks for acoustic modeling in speech recognition
Geoffrey Hinton, Li Deng, Dong Yu, George Dahl, Abdel-rahman Mohamed, Navdeep Jaitly, Andrew Senior, Vincent Vanhoucke, Patrick Nguyen, and Brian Kingsbury · 2012
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton · 2012
Earlier work this paper cites.
Understanding the exploding gradient problem. corr abs/1211.5063 (2012)
Razvan Pascanu, Tomas Mikolov, and Yoshua Bengio · 2012
Earlier work this paper cites.
Minimization methods for non-differentiable functions , volume 3
Naum Zuselevich Shor · 2012
Earlier work this paper cites.
Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude
Tijmen Tieleman, Geoffrey Hinton, et al · 2012
Earlier work this paper cites.
Stochastic first-and zeroth-order methods for nonconvex stochastic programming
Saeed Ghadimi and Guanghui Lan · 2013
Earlier work this paper cites.
On the difficulty of training recurrent neural networks
Razvan Pascanu, Tomas Mikolov, and Yoshua Bengio · 2013
Earlier work this paper cites.
Scaling distributed machine learning with the parameter server
Mu Li, David G Andersen, Jun Woo Park, Alexander J Smola, Amr Ahmed, Vanja Josifovski, James Long, Eugene J Shekita, and Bor-Yiing Su · 2014
Earlier work this paper cites.
Iteration complexity of randomized block-coordinate descent methods for minimizing a composite function
Peter Richtárik and Martin Takác · 2014
Earlier work this paper cites.
Beyond convexity: Stochastic quasi-convex optimization
Elad Hazan, Kfir Y Levy, and Shai Shalev-Shwartz · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Deep learning
Yann LeCun, Yoshua Bengio, and Geoffrey Hinton · 2015
Cited alongside, same era.
Deeply-supervised nets
Chen-Yu Lee, Saining Xie, Patrick Gallagher, Zhengyou Zhang, and Zhuowen Tu · 2015
Cited alongside, same era.
Scale-free algorithms for online linear optimization
Francesco Orabona and Dávid Pál · 2015
Cited alongside, same era.
Scale-free online learning
Francesco Orabona and Dávid Pál · 2015
Cited alongside, same era.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Cited alongside, same era.
Non-convex optimization for machine learning
Prateek Jain and Purushottam Kar · 2017
Cited alongside, same era.
Building a large annotated corpus of English: The Penn Treebank
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter · 2019
Later among the works it cites.
Pytorch: An imperative style, high-performance deep learning library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al · 2019
Later among the works it cites.
Language models are unsupervised multitask learners
Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever · 2019
Later among the works it cites.
Adagrad stepsizes: Sharp convergence over nonconvex landscapes
Rachel Ward, Xiaoxia Wu, and Leon Bottou · 2019
Later among the works it cites.
Transfer Learning in Natural Language Processing , 2019
Thomas Wolf · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Mitchell P. Marcus, Mary Ann Marcinkiewicz, and Beatrice Santorini · 2017
Cited alongside, same era.
Pointer sentinel mixture models
Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Cited alongside, same era.
Dissecting Adam: The sign, magnitude and variance of stochastic gradients
Lukas Balles and Philipp Hennig · 2018
Cited alongside, same era.
signSGD: Compressed optimisation for non-convex problems
Jeremy Bernstein, Yu-Xiang Wang, Kamyar Azizzadenesheli, and Animashree Anandkumar · 2018
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Cited alongside, same era.
Fangyu Zou, Li Shen, Zequn Jie, Weizhong Zhang, and Wei Liu · 2019
Later among the works it cites.
Momentum improves normalized SGD
Ashok Cutkosky and Harsh Mehta · 2020
Later among the works it cites.
A simple convergence proof of Adam and Adagrad
Alexandre Défossez, Léon Bottou, Francis Bach, and Nicolas Usunier · 2020
Later among the works it cites.
Stochastic optimization with heavy-tailed noise via accelerated gradient clipping
Eduard Gorbunov, Marina Danilova, and Alexander Gasnikov · 2020
Later among the works it cites.
Better theory for SGD in the nonconvex world
Ahmed Khaled and Peter Richtárik · 2020
Later among the works it cites.
Improved analysis of clipping algorithms for non-convex optimization
Bohang Zhang, Jikai Jin, Cong Fang, and Liwei Wang · 2020
Later among the works it cites.
High-probability bounds for non-convex stochastic optimization with heavy tails
Ashok Cutkosky and Harsh Mehta · 2021
Later among the works it cites.
Non-convex distributionally robust optimization: Non-asymptotic analysis
Jikai Jin, Bohang Zhang, Haiyang Wang, and Liwei Wang · 2021
Later among the works it cites.
Stability and convergence of stochastic gradient clipping: Beyond lipschitz continuity and smoothness
Vien V Mai and Mikael Johansson · 2021
Later among the works it cites.
Adaptive federated optimization
Sashank J. Reddi, Zachary Charles, Manzil Zaheer, Zachary Garrett, Keith Rush, Jakub Konečný, Sanjiv Kumar, and Hugh Brendan McMahan · 2021
Later among the works it cites.
Rethinking adam: A twofold exponential moving average approach
Yizhou Wang, Yue Kang, Can Qin, Huan Wang, Yilun Xu, Yulun Zhang, and Yun Raymond Fu · 2021
Later among the works it cites.
Understanding the generalization of Adam in learning neural networks with proper regularization
Difan Zou, Yuan Cao, Yuanzhi Li, and Quanquan Gu · 2021
Later among the works it cites.
Understanding AdamW through proximal methods and scale-freeness
Zhenxun Zhuang, Mingrui Liu, Ashok Cutkosky, and Francesco Orabona · 2022
Closest in time.