Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude
T. Tieleman and G. Hinton · 2012
Cited alongside, same era.
On stochastic gradient and subgradient methods with adaptive steplength sequences
F. Yousefian, A. Nedić, and U. V. Shanbhag · 2012
Cited alongside, same era.
ADADELTA: an adaptive learning rate method
Original
M. D. Zeiler · 2012
Cited alongside, same era.
Stochastic first-and zeroth-order methods for nonconvex stochastic programming
S. Ghadimi and G. Lan · 2013
Cited alongside, same era.
Stochastic majorization-minimization algorithms for large-scale optimization
J. Mairal · 2013
Cited alongside, same era.
Optimization, learning, and games with predictable sequences
A. Rakhlin and K. Sridharan · 2013
Cited alongside, same era.
TensorFlow: Large-scale machine learning on heterogeneous systems, 2015
M. Abadi, A. Agarwal, P. Barham, E. Brevdo, Z. Chen, C. Citro, G. S. Corrado, A. Davis, J. Dean, M. Devin, S. Ghemawat, I. Goodfellow, A. Harp, G. Irving, M. Isard, Y. Jia, R. Jozefowicz, L. Kaiser, M. Kudlur, J. Levenberg, D. Mané, R. Monga, S. Moore, D. Murray, C. Olah, M. Schuster, J. Shlens, B. Steiner, I. Sutskever, K. Talwar, P. Tucker, V. Vanhoucke, V. Vasudevan, F. Viégas, O. Vinyals, P. Warden, M. Wattenberg, M. Wicke, Y. Yu, and X. Zheng · 2015
Cited alongside, same era.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2015
Cited alongside, same era.
Scale-free online learning
F. Orabona and D. Pál · 2015
Cited alongside, same era.
Linear convergence of gradient and proximal-gradient methods under the Polyak-Łojasiewicz condition
H. Karimi, J. Nutini, and M. Schmidt · 2016
Cited alongside, same era.