Online geometric optimization in the bandit setting against an adaptive adversary
H Brendan McMahan and Avrim Blum · 2004
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky, Geoffrey Hinton, et al · 2009
Earlier work this paper cites.
Robust stochastic approximation approach to stochastic programming
Arkadi Nemirovski, Anatoli Juditsky, Guanghui Lan, and Alexander Shapiro · 2009
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
John Duchi, Elad Hazan, and Yoram Singer · 2011
Earlier work this paper cites.
Solving variational inequalities with stochastic mirror-prox algorithm
Anatoli Juditsky, Arkadi Nemirovski, Claire Tauvel, et al · 2011
Earlier work this paper cites.
Adam + {}^{\mbox{+}} : A stochastic method with adaptive variance reduction
Original
Mingrui Liu, Wei Zhang, Francesco Orabona, and Tianbao Yang · 2011
Earlier work this paper cites.
How to make the gradients small
Yurii Nesterov · 2012
Earlier work this paper cites.
Lecture 6.5-rmsprop, coursera: Neural networks for machine learning
Tijmen Tieleman and Geoffrey Hinton · 2012
Earlier work this paper cites.
Stochastic first-and zeroth-order methods for nonconvex stochastic programming
Saeed Ghadimi and Guanghui Lan · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Original
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
TensorFlow: Large-scale machine learning on heterogeneous systems, 2015
Martín Abadi, Ashish Agarwal, Paul Barham, et al · 2015
Earlier work this paper cites.
Adding gradient noise improves learning for very deep networks
Original
Arvind Neelakantan, Luke Vilnis, Quoc V Le, Ilya Sutskever, Lukasz Kaiser, Karol Kurach, and James Martens · 2015
Earlier work this paper cites.
Accelerated gradient methods for nonconvex nonlinear and stochastic programming
Saeed Ghadimi and Guanghui Lan · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Linear convergence of gradient and proximal-gradient methods under the polyak-łojasiewicz condition
Hamed Karimi, Julie Nutini, and Mark Schmidt · 2016
Earlier work this paper cites.
Unified convergence analysis of stochastic momentum methods for convex and non-convex optimization
Original
Tianbao Yang, Qihang Lin, and Zhe Li · 2016
Earlier work this paper cites.
Stochastic online auc maximization
Yiming Ying, Longyin Wen, and Siwei Lyu · 2016
Earlier work this paper cites.
A sufficient condition for convergences of Adam and RMSProp
Fangyu Zou, Li Shen, Zequn Jie, Weizhong Zhang, and Wei Liu · 2016
Earlier work this paper cites.
Qsgd: Communication-efficient sgd via gradient quantization and encoding
Dan Alistarh, Demjan Grubic, Jerry Li, Ryota Tomioka, and Milan Vojnovic · 2017
Earlier work this paper cites.
Neural message passing for quantum chemistry
Justin Gilmer, Samuel S Schoenholz, Patrick F Riley, Oriol Vinyals, and George E Dahl · 2017
Earlier work this paper cites.
Stochastic compositional gradient descent: algorithms for minimizing compositions of expected-value functions
Mengdi Wang, Ethan X Fang, and Han Liu · 2017
Earlier work this paper cites.
ZipML: Training linear models with end-to-end low precision, and a little bit of deep learning
Hantian Zhang, Jerry Li, Kaan Kara, Dan Alistarh, Ji Liu, and Ce Zhang · 2017
Earlier work this paper cites.
How to make the gradients small stochastically: Even faster convex and nonconvex sgd
Zeyuan Allen-Zhu · 2018
Earlier work this paper cites.
Spider: Near-optimal non-convex optimization via stochastic path-integrated differential estimator
Cong Fang, Chris Junchi Li, Zhouchen Lin, and Tong Zhang · 2018
Earlier work this paper cites.