Fetching the paper…
Reading the bibliography…
Adaptive regularization methods pre-multiply a descent direction by a preconditioning matrix.
On the limited memory bfgs method for large scale optimization
Dong C Liu and Jorge Nocedal · 1989
Earlier work this paper cites.
Acceleration of stochastic approximation by averaging
Boris T Polyak and Anatoli B Juditsky · 1992
Earlier work this paper cites.
The penn treebank: annotating predicate argument structure
Mitchell Marcus, Grace Kim, Mary Ann Marcinkiewicz, Robert MacIntyre, Ann Bies, Mark Ferguson, Karen Katz, and Britta Schasberger · 1994
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Natural gradient works efficiently in learning
Shun-Ichi Amari · 1998
Earlier work this paper cites.
Logarithmic regret algorithms for online convex optimization
Elad Hazan, Amit Agarwal, and Satyen Kale · 2007
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
John Duchi, Elad Hazan, and Yoram Singer · 2011
Earlier work this paper cites.
Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude
Tijmen Tieleman and Geoffrey Hinton · 2012
Earlier work this paper cites.
Adadelta: an adaptive learning rate method
Matthew D Zeiler · 2012
Earlier work this paper cites.
Stochastic first-and zeroth-order methods for nonconvex stochastic programming
Saeed Ghadimi and Guanghui Lan · 2013
Earlier work this paper cites.
On the difficulty of training recurrent neural networks
Razvan Pascanu, Tomas Mikolov, and Yoshua Bengio · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman · 2014
Earlier work this paper cites.
Convex optimization: Algorithms and complexity
Sébastien Bubeck et al · 2015
Earlier work this paper cites.
MxNet: A flexible and efficient machine learning library for heterogeneous distributed systems
Tianqi Chen, Mu Li, Yutian Li, Min Lin, Naiyan Wang, Minjie Wang, Tianjun Xiao, Bing Xu, Chiyuan Zhang, and Zheng Zhang · 2015
Earlier work this paper cites.
Convergence rates of sub-sampled newton methods
Murat A Erdogdu and Andrea Montanari · 2015
Earlier work this paper cites.
Optimizing neural networks with kronecker-factored approximate curvature
James Martens and Roger Grosse · 2015
Earlier work this paper cites.
Tensorflow: Large-scale machine learning on heterogeneous distributed systems
Martín Abadi, Ashish Agarwal, Paul Barham, Eugene Brevdo, Zhifeng Chen, Craig Citro, Greg S Corrado, Andy Davis, Jeffrey Dean, Matthieu Devin, et al · 2016
Cited alongside, same era.
Variance reduction for faster non-convex optimization
Zeyuan Allen-Zhu and Elad Hazan · 2016
Cited alongside, same era.
Incorporating Nesterov momentum into Adam
Timothy Dozat · 2016
Cited alongside, same era.
Zoneout: Regularizing rnns by randomly preserving hidden activations
David Krueger, Tegan Maharaj, János Kramár, Mohammad Pezeshki, Nicolas Ballas, Nan Rosemary Ke, Anirudh Goyal, Yoshua Bengio, Aaron Courville, and Chris Pal · 2016
Cited alongside, same era.
Scalable adaptive stochastic optimization using random projections
Gabriel Krummenacher, Brian McWilliams, Yannic Kilcher, Joachim M Buhmann, and Nicolai Meinshausen · 2016
Cited alongside, same era.
Automatic differentiation in PyTorch
Adam Paszke, Sam Gross, Soumith Chintala, Gregory Chanan, Edward Yang, Zachary DeVito, Zeming Lin, Alban Desmaison, Luca Antiga, and Adam Lerer · 2017
Later among the works it cites.
Svcca: Singular vector canonical correlation analysis for deep understanding and improvement
Maithra Raghu, Justin Gilmer, Jason Yosinski, and Jascha Sohl-Dickstein · 2017
Later among the works it cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Later among the works it cites.
The marginal value of adaptive gradient methods in machine learning
Ashia C Wilson, Rebecca Roelofs, Mitchell Stern, Nati Srebro, and Benjamin Recht · 2017
Later among the works it cites.
Stronger generalization bounds for deep nets via a compression approach
Sanjeev Arora, Rong Ge, Behnam Neyshabur, and Yi Zhang · 2018
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Efficient second order online learning by sketching
Haipeng Luo, Alekh Agarwal, Nicolo Cesa-Bianchi, and John Langford · 2016
Cited alongside, same era.
Sgdr: stochastic gradient descent with restarts
Ilya Loshchilov and Frank Hutter · 2016
Cited alongside, same era.
Compadagrad: A compressed, complementary, computationally-efficient adaptive gradient method
Nishant A Mehta, Alistair Rendell, Anish Varghese, and Christfried Webers · 2016
Cited alongside, same era.
Finding approximate local minima faster than gradient descent
Naman Agarwal, Zeyuan Allen-Zhu, Brian Bullins, Elad Hazan, and Tengyu Ma · 2017
Cited alongside, same era.
Second-order stochastic optimization for machine learning in linear time
Naman Agarwal, Brian Bullins, and Elad Hazan · 2017
Cited alongside, same era.
Katyusha: The first direct acceleration of stochastic gradient methods
Zeyuan Allen-Zhu · 2017
Cited alongside, same era.
“convex until proven guilty”: Dimension-free acceleration of gradient descent on non-convex functions
Yair Carmon, John C Duchi, Oliver Hinder, and Aaron Sidford · 2017
Cited alongside, same era.
On the convergence of a class of adam-type algorithms for non-convex optimization
Xiangyi Chen, Sijia Liu, Ruoyu Sun, and Mingyi Hong · 2018
Closest in time.
Universal stagewise learning for non-convex problems with convergence on averaged solutions
Zaiyi Chen, Tianbao Yang, Jinfeng Yi, Bowen Zhou, and Enhong Chen · 2018
Closest in time.
Shampoo: Preconditioned stochastic tensor optimization
Vineet Gupta, Tomer Koren, and Yoram Singer · 2018
Closest in time.
On the convergence of stochastic gradient descent with adaptive stepsizes
Xiaoyu Li and Francesco Orabona · 2018
Closest in time.
Kronecker-factored curvature approximations for recurrent neural networks
James Martens, Jimmy Ba, and Matt Johnson · 2018
Closest in time.
An analysis of neural language modeling at multiple scales
Stephen Merity, Nitish Shirish Keskar, and Richard Socher · 2018
Closest in time.
On the convergence of Adam and beyond
Sashank J Reddi, Satyen Kale, and Sanjiv Kumar · 2018
Closest in time.
Adagrad stepsizes: Sharp convergence over nonconvex landscapes, from any initialization
Rachel Ward, Xiaoxia Wu, and Leon Bottou · 2018
Closest in time.
On the convergence of adagrad with momentum for training deep neural networks
Fangyu Zou and Li Shen · 2018
Closest in time.
On the convergence of adaptive gradient methods for nonconvex optimization
Dongruo Zhou, Yiqi Tang, Ziyan Yang, Yuan Cao, and Quanquan Gu · 2018
Closest in time.
Escaping saddle points with adaptive gradient methods
Matthew Staib, Sashank J Reddi, Satyen Kale, Sanjiv Kumar, and Suvrit Sra · 2019
Closest in time.