Fetching the paper…
Reading the bibliography…
Accelerated gradient-based methods are being extensively used for solving non-convex machine learning problems, especially when the data points are abundant or the available data is distributed across several agents.
Systemes d’équations différentielles d’oscillations non linéaires
I Barbalat · 1959
Earlier work this paper cites.
Parallel and distributed computation: numerical methods
Dimitri P Bertsekas and John N Tsitsiklis · 1989
Earlier work this paper cites.
Iterative methods for optimization
Carl T Kelley · 1999
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
John Duchi, Elad Hazan, and Yoram Singer · 2011
Earlier work this paper cites.
Adadelta: An adaptive learning rate method
Matthew D Zeiler · 2012
Earlier work this paper cites.
Learning from data
Yaser S Abu-Mostafa, Malik Magdon-Ismail, and Hsuan-Tien Lin · 2012
Earlier work this paper cites.
Estimation, optimization, and parallelism when data is sparse
John Duchi, Michael I Jordan, and Brendan McMahan · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Unsupervised representation learning with deep convolutional generative adversarial networks
Alec Radford, Luke Metz, and Soumith Chintala · 2015
Earlier work this paper cites.
Incorporating nesterov momentum into adam
Timothy Dozat · 2016
Earlier work this paper cites.
Google’s neural machine translation system: Bridging the gap between human and machine translation
Yonghui Wu, Mike Schuster, Zhifeng Chen, Quoc V Le, Mohammad Norouzi, Wolfgang Macherey, Maxim Krikun, Yuan Cao, Qin Gao, Klaus Macherey, et al · 2016
Cited alongside, same era.
The marginal value of adaptive gradient methods in machine learning
Ashia C Wilson, Rebecca Roelofs, Mitchell Stern, Nathan Srebro, and Benjamin Recht · 2017
Cited alongside, same era.
Optimization methods for large-scale machine learning
Léon Bottou, Frank E Curtis, and Jorge Nocedal · 2018
Cited alongside, same era.
Deep contextualized word representations
Matthew E Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer · 2018
Cited alongside, same era.
Adashift: Decorrelation and convergence of adaptive learning rate methods
On the convergence of stochastic gradient descent with adaptive stepsizes
Xiaoyu Li and Francesco Orabona · 2019
Later among the works it cites.
Adagrad stepsizes: Sharp convergence over nonconvex landscapes
Rachel Ward, Xiaoxia Wu, and Leon Bottou · 2019
Later among the works it cites.
Calibrating the adaptive learning rate to improve convergence of ADAM
Qianqian Tong, Guannan Liang, and Jinbo Bi · 2019
Later among the works it cites.
Kushal Chakrabarti, Nirupam Gupta, and Nikhil Chopra · 2020
Later among the works it cites.
Adabelief optimizer: Adapting stepsizes by the belief in observed gradients
Juntang Zhuang, Tommy Tang, Sekhar Tatikonda, Nicha Dvornek, Yifan Ding, Xenophon Papademetris, and James S Duncan · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Zhiming Zhou, Qingru Zhang, Guansong Lu, Hongwei Wang, Weinan Zhang, and Yong Yu · 2018
Cited alongside, same era.
Adaptive methods for nonconvex optimization
S Reddi, Manzil Zaheer, Devendra Sachan, Satyen Kale, and Sanjiv Kumar · 2018
Cited alongside, same era.
WNGrad: Learn the learning rate in gradient descent
Xiaoxia Wu, Rachel Ward, and Léon Bottou · 2018
Cited alongside, same era.
Soham De, Anirbit Mukherjee, and Enayat Ullah · 2018
Cited alongside, same era.
On the convergence of a class of adam-type algorithms for non-convex optimization
Xiangyi Chen, Sijia Liu, Ruoyu Sun, and Mingyi Hong · 2018
Cited alongside, same era.
On the convergence of adam and beyond
Sashank J Reddi, Satyen Kale, and Sanjiv Kumar · 2019
Cited alongside, same era.
Later among the works it cites.
A simple convergence proof of Adam and Adagrad
Alexandre Défossez, Léon Bottou, Francis Bach, and Nicolas Usunier · 2020
Later among the works it cites.
Convergence rates of a momentum algorithm with bounded adaptive step size for nonconvex optimization
Anas Barakat and Pascal Bianchi · 2020
Later among the works it cites.
https://www.kaggle.com/oddrationale/mnist-in-csv
MNIST in CSV · 2020
Later among the works it cites.
Convergence and dynamical behavior of the ADAM algorithm for nonconvex stochastic optimization
Anas Barakat and Pascal Bianchi · 2021
Closest in time.