Fetching the paper…
Reading the bibliography…
Adaptive gradient methods have attracted much attention of machine learning communities due to the high efficiency.
Adaptive subgradient methods for online learning and stochastic optimization
John Duchi, Elad Hazan, and Yoram Singer · 2011
Earlier work this paper cites.
Nathan Halko, Per-Gunnar Martinsson, and Joel A Tropp · 2011
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Escaping from saddle points—online stochastic gradient for tensor decomposition
Rong Ge, Furong Huang, Chi Jin, and Yang Yuan · 2015
Earlier work this paper cites.
Non-convex finite-sum optimization via scsg methods
Lihua Lei, Cheng Ju, Jianbo Chen, and Michael I Jordan · 2017
Earlier work this paper cites.
On the convergence of adam and beyond
Sashank J. Reddi, Satyen Kale, and Sanjiv Kumar · 2018
Earlier work this paper cites.
On the convergence of adaptive gradient methods for nonconvex optimization
Dongruo Zhou, Yiqi Tang, Ziyan Yang, Yuan Cao, and Quanquan Gu · 2018
Earlier work this paper cites.
Adaptive methods for nonconvex optimization
Manzil Zaheer, Sashank Reddi, Devendra Sachan, Satyen Kale, and Sanjiv Kumar · 2018
Cited alongside, same era.
Adagrad stepsizes: Sharp convergence over nonconvex landscapes, from any initialization
Rachel Ward, Xiaoxia Wu, and Leon Bottou · 2018
Cited alongside, same era.
Sadagrad: Strongly adaptive stochastic gradient methods
Zaiyi Chen, Yi Xu, Enhong Chen, and Tianbao Yang · 2018
Cited alongside, same era.
Random shuffling beats sgd after finite epochs
Jeffery Z HaoChen and Suvrit Sra · 2018
Cited alongside, same era.
Natasha 2: Faster non-convex optimization than sgd
Zeyuan Allen-Zhu · 2018
Cited alongside, same era.
Sadam: A variant of adam for strongly convex functions
Guanghui Wang, Shiyin Lu, Weiwei Tu, and Lijun Zhang · 2019
Later among the works it cites.
Spiderboost and momentum: Faster variance reduction algorithms
Zhe Wang, Kaiyi Ji, Yi Zhou, Yingbin Liang, and Vahid Tarokh · 2019
Later among the works it cites.
Sgd without replacement: Sharper rates for general smooth convex functions
Dheeraj Nagaraj, Prateek Jain, and Praneeth Netrapalli · 2019
Later among the works it cites.
Efficient full-matrix adaptive regularization
Naman Agarwal, Brian Bullins, Xinyi Chen, Elad Hazan, Karan Singh, Cyril Zhang, and Yi Zhang · 2019
Later among the works it cites.
A unified convergence analysis for shuffling-type gradient methods
Lam M Nguyen, Quoc Tran-Dinh, Dzung T Phan, Phuong Ha Nguyen, and Marten van Dijk · 2020
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cong Fang, Chris Junchi Li, Zhouchen Lin, and Tong Zhang · 2018
Cited alongside, same era.
On the convergence of a class of adam-type algorithms for non-convex optimization
Xiangyi Chen, Sijia Liu, Ruoyu Sun, and Mingyi Hong · 2019
Cited alongside, same era.
Acutum: When generalization meets adaptability, 2020
Xunpeng Huang, Zhengyang Liu, Zhe Wang, Yue Yu, and Lei Li · 2020
Closest in time.