Fetching the paper…
Reading the bibliography…
Momentum-based acceleration of stochastic gradient descent (SGD) is widely used in deep learning.
Relative and absolute strength of response as a function of frequency of reinforcement 1, 2
Shin-ho Chung and Richard J Hernstein · 1961
Earlier work this paper cites.
Some methods of speeding up the convergence of iteration methods
BT Polyak · 1964
Earlier work this paper cites.
On second-best national saving and game-equilibrium growth
Edmund S Phelps and Robert A Pollak · 1968
Earlier work this paper cites.
A method for solving the convex programming problem with convergence rate o (1/kˆ 2)
Yurii E Nesterov · 1983
Earlier work this paper cites.
PID Controllers: Theory, Design, and Tuning
K. J. Aaström and T. Hägglund · 1995
Earlier work this paper cites.
Golden eggs and hyperbolic discounting
David Laibson · 1997
Earlier work this paper cites.
The mnist database of handwritten digits
Yann LeCun · 1998
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky · 2009
Earlier work this paper cites.
Neural networks for machine learning: Lecture 6a, 2012
Geoffrey Hinton, Nitish Srivastava, and Kevin Swersky · 2012
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton · 2012
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
Emanuel Todorov, Tom Erez, and Yuval Tassa · 2012
Earlier work this paper cites.
Accelerating stochastic gradient descent using predictive variance reduction
Rie Johnson and Tong Zhang · 2013
Earlier work this paper cites.
Minimizing finite sums with the stochastic average gradient
Mark W. Schmidt, Nicolas Le Roux, and Francis R. Bach · 2013
Earlier work this paper cites.
On the importance of initialization and momentum in deep learning
Ilya Sutskever, James Martens, George Dahl, and Geoffrey Hinton · 2013
Earlier work this paper cites.
SAGA: A fast incremental gradient method with support for non-strongly convex composite objectives
Aaron Defazio, Francis R. Bach, and Simon Lacoste-Julien · 2014
Earlier work this paper cites.
TensorFlow: Large-scale machine learning on heterogeneous systems, 2015
Martín Abadi, Ashish Agarwal, Paul Barham, Eugene Brevdo, Zhifeng Chen, Craig Citro, Greg S. Corrado, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Ian Goodfellow, Andrew Harp, Geoffrey Irving, Michael Isard, Yangqing Jia, Rafal Jozefowicz, Lukasz Kaiser, Manjunath Kudlur, Josh Levenberg, Dandelion Mané, Rajat Monga, Sherry Moore, Derek Murray, Chris Olah, Mike Schuster, Jonathon Shlens, Benoit Steiner, Ilya Sutskever, Kunal Talwar, Paul Tucker, Vincent Vanhoucke, Vijay Vasudevan, Fernanda Viégas, Oriol Vinyals, Pete Warden, Martin Wattenberg, Martin Wicke, Yuan Yu, and Xiaoqiang Zheng · 2015
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Sergey Ioffe and Christian Szegedy · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Imagenet large scale visual recognition challenge
Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael S. Bernstein, Alexander C. Berg, and Fei-Fei Li · 2015
Cited alongside, same era.
Show, attend and tell: Neural image caption generation with visual attention
Kelvin Xu, Jimmy Ba, Ryan Kiros, Kyunghyun Cho, Aaron C. Courville, Ruslan Salakhutdinov, Richard S. Zemel, and Yoshua Bengio · 2015
Cited alongside, same era.
Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wojciech Zaremba · 2016
Cited alongside, same era.
Language modeling with gated convolutional networks
Yann N Dauphin, Angela Fan, Michael Auli Auli, and David Grangier · 2016
Cited alongside, same era.
Incorporating nesterov momentum into adam
Timothy Dozat · 2016
Cited alongside, same era.
Nicolas Loizou and Peter Richtárik · 2017
Later among the works it cites.
Fixing weight decay regularization in adam
Ilya Loshchilov and Frank Hutter · 2017
Later among the works it cites.
Automatic differentiation in pytorch
Adam Paszke, Sam Gross, Soumith Chintala, Gregory Chanan, Edward Yang, Zachary DeVito, Zeming Lin, Alban Desmaison, Luca Antiga, and Adam Lerer · 2017
Later among the works it cites.
The fastest known globally convergent first-order method for minimizing strongly convex functions
Bryan Van Scoy, Randy A. Freeman, and Kevin M. Lynch · 2017
Later among the works it cites.
Fast stochastic variance reduced gradient method with momentum acceleration for machine learning
Fanhua Shang, Yuanyuan Liu, James Cheng, and Jiacheng Zhuo · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Identity mappings in deep residual networks
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Cited alongside, same era.
Analysis and design of optimization algorithms via integral quadratic constraints
Laurent Lessard, Benjamin Recht, and Andrew Packard · 2016
Cited alongside, same era.
Pointer sentinel mixture models
Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher · 2016
Cited alongside, same era.
Asynchrony begets momentum, with an application to deep learning
Ioannis Mitliagkas, Ce Zhang, Stefan Hadjis, and Christopher Ré · 2016
Cited alongside, same era.
Pytorch examples
Adam Paszke, Sam Gross, Soumith Chintala, Gregory Chanan, Edward Yang, Zachary DeVito, Zeming Lin, Alban Desmaison, Luca Antiga, and Adam Lerer · 2016
Cited alongside, same era.
An overview of gradient descent optimization algorithms
Sebastian Ruder · 2016
Cited alongside, same era.
Katyusha: The first direct acceleration of stochastic gradient methods
Zeyuan Allen-Zhu · 2017
Cited alongside, same era.
Later among the works it cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin · 2017
Later among the works it cites.
Yellowfin and the art of momentum tuning
Jian Zhang, Ioannis Mitliagkas, and Christopher Ré · 2017
Later among the works it cites.
A pid controller approach for stochastic optimization of deep networks
Wangpeng An, Haoqian Wang, Qingyun Sun, Jun Xu, Qionghai Dai, and Lei Zhang · 2018
Closest in time.
Online learning rate adaptation with hypergradient descent
Atilim Gunes Baydin, Robert Cornish, David Martinez Rubio, Mark Schmidt, and Frank Wood · 2018
Closest in time.
A robust accelerated optimization algorithm for strongly convex functions
Saman Cyrus, Bin Hu, Bryan Van Scoy, and Laurent Lessard · 2018
Closest in time.
Addressing function approximation error in actor-critic methods
Scott Fujimoto, Herke van Hoof, and Dave Meger · 2018
Closest in time.
On the insufficiency of existing momentum schemes for stochastic optimization
Rahul Kidambi, Praneeth Netrapalli, Prateek Jain, and Sham M. Kakade · 2018
Closest in time.
Aggregated momentum: Stability through passive damping
James Lucas, Richard S. Zemel, and Roger Grosse · 2018
Closest in time.
Scaling neural machine translation
Myle Ott, Sergey Edunov, David Grangier Grangier, and Michael Auli · 2018
Closest in time.
The best things in life are model free
Ben Recht · 2018
Closest in time.
On the convergence of adam and beyond
Sashank J. Reddi, Satyen Kale, and Sanjiv Kumar · 2018
Closest in time.
Online variance-reducing optimization, 2018
Nicolas Le Roux, Reza Babanezhad, and Pierre-Antoine Manzagol · 2018
Closest in time.
Qanet: Combining local convolution with global self-attention for reading comprehension
Adams Wei Yu, David Dohan, Minh-Thang Luong, Rui Zhao, Kai Chen, Mohammad Norouzi, and Quoc V Le · 2018
Closest in time.