Fetching the paper…
Reading the bibliography…
RMSProp and ADAM continue to be extremely popular algorithms for training neural nets but their theoretical convergence properties have remained unclear.
A method of solving a convex programming problem with convergence rate o (1/k2)
Yurii Nesterov · 1983
Earlier work this paper cites.
Introduction to optimization. translations series in mathematics and engineering
Boris T Polyak · 1987
Earlier work this paper cites.
Heavy-ball method in nonconvex optimization problems
SK Zavriev and FV Kostyuk · 1993
Earlier work this paper cites.
Momentum and optimal stochastic search
Genevieve B Orr and Todd K Leen · 1994
Earlier work this paper cites.
Stochastic dynamics of learning with momentum in neural networks
Wim Wiegerinck, Andrzej Komoda, and Tom Heskes · 1994
Earlier work this paper cites.
ARPACK users’ guide: solution of large-scale eigenvalue problems with implicitly restarted Arnoldi methods
Richard B Lehoucq, Danny C Sorensen, and Chao Yang · 1998
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
Xavier Glorot and Yoshua Bengio · 2010
Earlier work this paper cites.
Stacked denoising autoencoders: Learning useful representations in a deep network with a local denoising criterion
Pascal Vincent, Hugo Larochelle, Isabelle Lajoie, Yoshua Bengio, and Pierre-Antoine Manzagol · 2010
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
John Duchi, Elad Hazan, and Yoram Singer · 2011
Earlier work this paper cites.
Autoencoders, unsupervised learning, and deep architectures
Pierre Baldi · 2012
Earlier work this paper cites.
Lecture 6.5-rmsprop, coursera: Neural networks for machine learning
Tijmen Tieleman and Geoffrey Hinton · 2012
Earlier work this paper cites.
Accelerating stochastic gradient descent using predictive variance reduction
Rie Johnson and Tong Zhang · 2013
Earlier work this paper cites.
On the importance of initialization and momentum in deep learning
Ilya Sutskever, James Martens, George Dahl, and Geoffrey Hinton · 2013
Earlier work this paper cites.
Saga: A fast incremental gradient method with support for non-strongly convex composite objectives
Aaron Defazio, Francis Bach, and Simon Lacoste-Julien · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization. arxiv. org, 2014
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman · 2014
Earlier work this paper cites.
Why regularized auto-encoders learn sparse representation?
Devansh Arpit, Yingbo Zhou, Hung Ngo, and Venu Govindaraju · 2015
Earlier work this paper cites.
Stop wasting my gradients: Practical svrg
Reza Babanezhad, Mohamed Osama Ahmed, Alim Virani, Mark Schmidt, Jakub Konecny, and Scott Sallinen · 2015
Cited alongside, same era.
Draw: A recurrent neural network for image generation
Karol Gregor, Ivo Danihelka, Alex Graves, Danilo Jimenez Rezende, and Daan Wierstra · 2015
Cited alongside, same era.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Sergey Ioffe and Christian Szegedy · 2015
Cited alongside, same era.
Skip-thought vectors
Ryan Kiros, Yukun Zhu, Ruslan R Salakhutdinov, Richard Zemel, Raquel Urtasun, Antonio Torralba, and Sanja Fidler · 2015
Cited alongside, same era.
Optimizing neural networks with kronecker-factored approximate curvature
James Martens and Roger Grosse · 2015
Cited alongside, same era.
Training deep autoencoders for collaborative filtering
Oleksii Kuchaiev and Boris Ginsburg · 2017
Later among the works it cites.
Nicolas Loizou and Peter Richtárik · 2017
Later among the works it cites.
On the state of the art of evaluation in neural language models
Gábor Melis, Chris Dyer, and Phil Blunsom · 2017
Later among the works it cites.
Behavior of accelerated gradient methods near critical points of nonconvex problems
Michael O’Neill and Stephen J Wright · 2017
Later among the works it cites.
Critical points of an autoencoder can provably recover sparsely used overcomplete dictionaries
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Unsupervised representation learning with deep convolutional generative adversarial networks
Alec Radford, Luke Metz, and Soumith Chintala · 2015
Cited alongside, same era.
Show, attend and tell: Neural image caption generation with visual attention
Kelvin Xu, Jimmy Ba, Ryan Kiros, Kyunghyun Cho, Aaron Courville, Ruslan Salakhudinov, Rich Zemel, and Yoshua Bengio · 2015
Cited alongside, same era.
Tensorflow: A system for large-scale machine learning
Martín Abadi, Paul Barham, Jianmin Chen, Zhifeng Chen, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Geoffrey Irving, Michael Isard, et al · 2016
Cited alongside, same era.
Local convergence of the heavy-ball method and ipiano for non-convex optimization
Peter Ochs · 2016
Cited alongside, same era.
Unified convergence analysis of stochastic momentum methods for convex and non-convex optimization
Tianbao Yang, Qihang Lin, and Zhe Li · 2016
Cited alongside, same era.
On the influence of momentum acceleration on online learning
Kun Yuan, Bicheng Ying, and Ali H Sayed · 2016
Cited alongside, same era.
Empirical investigation of optimization algorithms in neural machine translation
Parnia Bahar, Tamer Alkhouli, Jan-Thorsten Peter, Christopher Jan-Steffen Brix, and Hermann Ney · 2017
Cited alongside, same era.
Akshay Rangamani, Anirbit Mukherjee, Ashish Arora, Tejaswini Ganapathy, Amitabh Basu, Sang Chin, and Trac D Tran · 2017
Later among the works it cites.
The marginal value of adaptive gradient methods in machine learning
Ashia C Wilson, Rebecca Roelofs, Mitchell Stern, Nati Srebro, and Benjamin Recht · 2017
Later among the works it cites.
signsgd: compressed optimisation for non-convex problems
Jeremy Bernstein, Yu-Xiang Wang, Kamyar Azizzadenesheli, and Anima Anandkumar · 2018
Closest in time.
On the convergence of a class of adam-type algorithms for non-convex optimization
Xiangyi Chen, Sijia Liu, Ruoyu Sun, and Mingyi Hong · 2018
Closest in time.
Stochastic heavy ball
Sébastien Gadat, Fabien Panloup, Sofiane Saadane, et al · 2018
Closest in time.
On the insufficiency of existing momentum schemes for stochastic optimization
Rahul Kidambi, Praneeth Netrapalli, Prateek Jain, and Sham M. Kakade · 2018
Closest in time.
On the convergence of stochastic gradient descent with adaptive stepsizes
Xiaoyu Li and Francesco Orabona · 2018
Closest in time.
Aggregated momentum: Stability through passive damping
James Lucas, Richard Zemel, and Roger Grosse · 2018
Closest in time.
On the convergence of adam and beyond
Sashank J Reddi, Satyen Kale, and Sanjiv Kumar · 2018
Closest in time.
A new trick for calculating Jacobian vector products
Jamie Townsend · 2018
Closest in time.
Adaptive methods for nonconvex optimization
Manzil Zaheer, Sashank Reddi, Devendra Sachan, Satyen Kale, and Sanjiv Kumar · 2018
Closest in time.
On the convergence of adaptive gradient methods for nonconvex optimization
Dongruo Zhou, Yiqi Tang, Ziyan Yang, Yuan Cao, and Quanquan Gu · 2018
Closest in time.