Fetching the paper…
Reading the bibliography…
Adaptive gradient methods have shown excellent performances for solving many machine learning problems.
An iterative row-action method for interval convex programming
Y. Censor and A. Lent · 1981
Earlier work this paper cites.
Two-point step size gradient methods
J. Barzilai and J. M. Borwein · 1988
Earlier work this paper cites.
Proximal minimization algorithm withd-functions
Y. Censor and S. A. Zenios · 1992
Earlier work this paper cites.
Long short-term memory
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
Mirror descent and nonlinear projected subgradient methods for convex optimization
A. Beck and M. Teboulle · 2003
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
J. Duchi, E. Hazan, and Y. Singer · 2011
Earlier work this paper cites.
Adadelta: an adaptive learning rate method
M. D. Zeiler · 2012
Earlier work this paper cites.
Stochastic first-and zeroth-order methods for nonconvex stochastic programming
S. Ghadimi and G. Lan · 2013
Earlier work this paper cites.
Generative adversarial nets
I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2014
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
K. Simonyan and A. Zisserman · 2014
Earlier work this paper cites.
Mini-batch stochastic approximation methods for nonconvex stochastic composite optimization
S. Ghadimi, G. Lan, and H. Zhang · 2016
Earlier work this paper cites.
Deep learning
I. Goodfellow, Y. Bengio, and A. Courville · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Earlier work this paper cites.
Variants of rmsprop and adagrad with logarithmic regret bounds
M. C. Mukkamala and M. Hein · 2017
Cited alongside, same era.
Stochastic quasi-newton methods for nonconvex stochastic optimization
X. Wang, S. Ma, D. Goldfarb, and W. Liu · 2017
Cited alongside, same era.
The marginal value of adaptive gradient methods in machine learning
A. C. Wilson, R. Roelofs, M. Stern, N. Srebro, and B. Recht · 2017
Cited alongside, same era.
Closing the generalization gap of adaptive gradient methods in training deep neural networks
J. Chen, D. Zhou, Y. Tang, Z. Yang, and Q. Gu · 2018
Cited alongside, same era.
Sadagrad: Strongly adaptive stochastic gradient methods
Z. Chen, Y. Xu, E. Chen, and T. Yang · 2018
Cited alongside, same era.
Spider: Near-optimal non-convex optimization via stochastic path-integrated differential estimator
On the convergence of stochastic gradient descent with adaptive stepsizes
X. Li and F. Orabona · 2019
Later among the works it cites.
On the variance of the adaptive learning rate and beyond
L. Liu, H. Jiang, P. He, W. Chen, X. Liu, J. Gao, and J. Han · 2019
Later among the works it cites.
Adaptive gradient methods with dynamic bound of learning rate
L. Luo, Y. Xiong, Y. Liu, and X. Sun · 2019
Later among the works it cites.
Escaping saddle points with adaptive gradient methods
M. Staib, S. Reddi, S. Kale, S. Kumar, and S. Sra · 2019
Later among the works it cites.
Hybrid stochastic gradient descent algorithms for stochastic nonconvex optimization
Q. Tran-Dinh, N. H. Pham, D. T. Phan, and L. M. Nguyen · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
C. Fang, C. J. Li, Z. Lin, and T. Zhang · 2018
Cited alongside, same era.
Decoupled weight decay regularization
I. Loshchilov and F. Hutter · 2018
Cited alongside, same era.
On the convergence of adam and beyond
S. J. Reddi, S. Kale, and S. Kumar · 2018
Cited alongside, same era.
Adaptive methods for nonconvex optimization
M. Zaheer, S. Reddi, D. Sachan, S. Kale, and S. Kumar · 2018
Cited alongside, same era.
On the convergence of adaptive gradient methods for nonconvex optimization
D. Zhou, J. Chen, Y. Cao, Y. Tang, Z. Yang, and Q. Gu · 2018
Cited alongside, same era.
Lower bounds for non-convex stochastic optimization
Y. Arjevani, Y. Carmon, J. C. Duchi, D. J. Foster, N. Srebro, and B. Woodworth · 2019
Cited alongside, same era.
On the convergence of a class of adam-type algorithms for non-convex optimization
X. Chen, S. Liu, R. Sun, and M. Hong · 2019
Cited alongside, same era.
Adagrad stepsizes: Sharp convergence over nonconvex landscapes
R. Ward, X. Wu, and L. Bottou · 2019
Later among the works it cites.
Momentum improves normalized sgd
A. Cutkosky and H. Mehta · 2020
Later among the works it cites.
Practical quasi-newton methods for training deep neural networks
D. Goldfarb, Y. Ren, and A. Bahamou · 2020
Later among the works it cites.
Adam+: A stochastic method with adaptive variance reduction
M. Liu, W. Zhang, F. Orabona, and T. Yang · 2020
Later among the works it cites.
Adabelief optimizer: Adapting stepsizes by the belief in observed gradients
J. Zhuang, T. Tang, Y. Ding, S. C. Tatikonda, N. Dvornek, X. Papademetris, and J. Duncan · 2020
Later among the works it cites.
Large-margin contrastive learning with distance polarization regularizer
S. Chen, G. Niu, C. Gong, J. Li, J. Yang, and M. Sugiyama · 2021
Closest in time.
On stochastic moving-average estimators for non-convex optimization
Z. Guo, Y. Xu, W. Yin, R. Jin, and T. Yang · 2021
Closest in time.
Faster stochastic quasi-newton methods
Q. Zhang, F. Huang, C. Deng, and H. Huang · 2021
Closest in time.