Fetching the paper…
Reading the bibliography…
Most popular optimizers for deep learning can be broadly categorized as adaptive methods (e.g.
“A stochastic approximation method,”
Herbert Robbins and Sutton Monro, · 1951
Earlier work this paper cites.
“Methods of conjugate gradients for solving linear systems,”
Magnus R Hestenes, Eduard Stiefel, et al., · 1952
Earlier work this paper cites.
“On minimizing a convex function subject to linear inequalities,”
Evelyn ML Beale, · 1955
Earlier work this paper cites.
“An automatic method for finding the greatest or least value of a function,”
HoHo Rosenbrock, · 1960
Earlier work this paper cites.
“Some methods of speeding up the convergence of iteration methods,”
Boris T Polyak, · 1964
Earlier work this paper cites.
“Quasi-likelihood functions, generalized linear models, and the gauss—newton method,”
Robert WM Wedderburn, · 1974
Earlier work this paper cites.
“Updating quasi-newton matrices with limited storage,”
Jorge Nocedal, · 1980
Earlier work this paper cites.
“A method of solving a convex programming problem with convergence rate o(1/k2̂),”
Yu Nesterov, · 1983
Earlier work this paper cites.
“First-and second-order methods for learning: between steepest descent and newton’s method,”
Roberto Battiti, · 1992
Earlier work this paper cites.
“Building a large annotated corpus of english: The penn treebank,”
Mitchell Marcus, Beatrice Santorini, and Mary Ann Marcinkiewicz, · 1993
Earlier work this paper cites.
“Natural gradient works efficiently in learning,”
Shun-Ichi Amari, · 1998
Earlier work this paper cites.
“Fast curvature matrix-vector products for second-order gradient descent,”
Nicol N Schraudolph, · 2002
Earlier work this paper cites.
Convex optimization
Stephen Boyd, Stephen P Boyd, and Lieven Vandenberghe, · 2004
Earlier work this paper cites.
“Topmoumoute online natural gradient algorithm,”
Nicolas L Roux, Pierre-Antoine Manzagol, and Yoshua Bengio, · 2008
Earlier work this paper cites.
“Learning multiple layers of features from tiny images,”
Alex Krizhevsky, Geoffrey Hinton, et al., · 2009
Earlier work this paper cites.
“Imagenet: A large-scale hierarchical image database,”
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei, · 2009
Earlier work this paper cites.
“Deep learning via hessian-free optimization.,”
James Martens, · 2010
Earlier work this paper cites.
“Adaptive bound optimization for online convex optimization,”
H Brendan McMahan and Matthew Streeter, · 2010
Earlier work this paper cites.
“Adaptive subgradient methods for online learning and stochastic optimization,”
John Duchi, Elad Hazan, and Yoram Singer, · 2011
Earlier work this paper cites.
“Deep sparse rectifier neural networks,”
Xavier Glorot, Antoine Bordes, and Yoshua Bengio, · 2011
Earlier work this paper cites.
“Adadelta: an adaptive learning rate method,”
Matthew D Zeiler, · 2012
Earlier work this paper cites.
“Lecture notes: Some notes on gradient descent,”
Marc Toussaint, · 2012
Earlier work this paper cites.
“On the importance of initialization and momentum in deep learning,”
Ilya Sutskever, James Martens, George Dahl, and Geoffrey Hinton, · 2013
Cited alongside, same era.
“Generating sequences with recurrent neural networks,”
Alex Graves, · 2013
Cited alongside, same era.
“Accelerating stochastic gradient descent using predictive variance reduction,”
Rie Johnson and Tong Zhang, · 2013
Cited alongside, same era.
“Revisiting natural gradient for deep networks,”
Razvan Pascanu and Yoshua Bengio, · 2013
Cited alongside, same era.
“Adam: A method for stochastic optimization,”
Diederik P Kingma and Jimmy Ba, · 2014
Cited alongside, same era.
“Gans trained by a two time-scale update rule converge to a local nash equilibrium,”
Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter, · 2017
Later among the works it cites.
“Attention is all you need,”
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin, · 2017
Later among the works it cites.
“Scaling sgd batch size to 32k for imagenet training,”
Yang You, Igor Gitman, and Boris Ginsburg, · 2017
Later among the works it cites.
“Adaptive methods for nonconvex optimization,”
Manzil Zaheer, Sashank Reddi, Devendra Sachan, Satyen Kale, and Sanjiv Kumar, · 2018
Later among the works it cites.
“signsgd: Compressed optimisation for non-convex problems,”
Jeremy Bernstein, Yu-Xiang Wang, Kamyar Azizzadenesheli, and Anima Anandkumar, · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Generative adversarial nets,”
Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio, · 2014
Cited alongside, same era.
“Very deep convolutional networks for large-scale image recognition,”
Karen Simonyan and Andrew Zisserman, · 2014
Cited alongside, same era.
“Imagenet large scale visual recognition challenge,”
Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, et al., · 2015
Cited alongside, same era.
“Long short-term memory neural network for traffic speed prediction using remote microwave sensor data,”
Xiaolei Ma, Zhimin Tao, Yinhai Wang, Haiyang Yu, and Yunpeng Wang, · 2015
Cited alongside, same era.
“Unsupervised representation learning with deep convolutional generative adversarial networks,”
Alec Radford, Luke Metz, and Soumith Chintala, · 2015
Cited alongside, same era.
“Deep residual learning for image recognition,”
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun, · 2016
Cited alongside, same era.
“Improved techniques for training gans,”
Tim Salimans, Ian Goodfellow, Wojciech Zaremba, Vicki Cheung, Alec Radford, and Xi Chen, · 2016
Cited alongside, same era.
“On the convergence of a class of adam-type algorithms for non-convex optimization,”
Xiangyi Chen, Sijia Liu, Ruoyu Sun, and Mingyi Hong, · 2018
Later among the works it cites.
“Closing the generalization gap of adaptive gradient methods in training deep neural networks,”
Jinghui Chen and Quanquan Gu, · 2018
Later among the works it cites.
“On the convergence of adaptive gradient methods for nonconvex optimization,”
Dongruo Zhou, Yiqi Tang, Ziyan Yang, Yuan Cao, and Quanquan Gu, · 2018
Later among the works it cites.
“Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor,”
Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine, · 2018
Later among the works it cites.
“Quasi-hyperbolic momentum and adam for deep learning,”
Jerry Ma and Denis Yarats, · 2018
Later among the works it cites.
“Nostalgic adam: Weighting more of the past gradients when designing the adaptive learning rate,”
Haiwen Huang, Chang Wang, and Bin Dong, · 2018
Later among the works it cites.
“Gradient descent maximizes the margin of homogeneous neural networks,”
Kaifeng Lyu and Jian Li, · 2019
Later among the works it cites.
“Adaptive gradient methods with dynamic bound of learning rate,”
Liangchen Luo, Yuanhao Xiong, Yan Liu, and Xu Sun, · 2019
Later among the works it cites.
“On the convergence of adam and beyond,”
Sashank J Reddi, Satyen Kale, and Sanjiv Kumar, · 2019
Later among the works it cites.
“On the variance of the adaptive learning rate and beyond,”
Liyuan Liu, Haoming Jiang, Pengcheng He, Weizhu Chen, Xiaodong Liu, Jianfeng Gao, and Jiawei Han, · 2019
Later among the works it cites.
“Chainerrl: A deep reinforcement learning library,”
Yasuhiro Fujita, Toshiki Kataoka, Prabhat Nagarajan, and Takahiro Ishikawa, · 2019
Later among the works it cites.
“Lookahead optimizer: k steps forward, 1 step back,”
Michael Zhang, James Lucas, Jimmy Ba, and Geoffrey E Hinton, · 2019
Later among the works it cites.
“Sadam: A variant of adam for strongly convex functions,”
Guanghui Wang, Shiyin Lu, Weiwei Tu, and Lijun Zhang, · 2019
Later among the works it cites.
“On the distance between two neural networks and the stability of learning,”
Jeremy Bernstein, Arash Vahdat, Yisong Yue, and Ming-Yu Liu, · 2020
Closest in time.
“Adahessian: An adaptive second order optimizer for machine learning,”
Zhewei Yao, Amir Gholami, Sheng Shen, Kurt Keutzer, and Michael W Mahoney, · 2020
Closest in time.
“Eadam optimizer: How epsilon impact adam,”
Wei Yuan and Kai-Xin Gao, · 2020
Closest in time.
“Adax: Adaptive gradient descent with exponential long term memory,”
Wenjie Li, Zhaoyang Zhang, Xinjiang Wang, and Ping Luo, · 2020
Closest in time.