Fetching the paper…
Reading the bibliography…
In this paper, we introduce Apollo, a quasi-Newton method for nonconvex stochastic optimization, which dynamically incorporates the curvature of the loss function by approximating the Hessian via a diagonal matrix.
A stochastic approximation method
Herbert Robbins and Sutton Monro · 1951
Earlier work this paper cites.
The secant method for simultaneous nonlinear equations
Philip Wolfe · 1959
Earlier work this paper cites.
Quasi-newton methods and their application to function minimisation
Charles G Broyden · 1967
Earlier work this paper cites.
The convergence of a class of double-rank minimization algorithms
Charles George Broyden · 1970
Earlier work this paper cites.
A new approach to variable metric algorithms
Roger Fletcher · 1970
Earlier work this paper cites.
A family of variable metric updates derived by variational means
D Goldfarb · 1970
Earlier work this paper cites.
Conditioning of quasi-newton methods for function minimization
David F Shanno · 1970
Earlier work this paper cites.
Quasi-newton methods, motivation and theory
John E Dennis, Jr and Jorge J Moré · 1977
Earlier work this paper cites.
Learning internal representations by error propagation
David E Rumelhart, Geoffrey E Hinton, and Ronald J Williams · 1985
Earlier work this paper cites.
Practical methods of optimization
Roger Fletcher · 1987
Earlier work this paper cites.
Improving the convergence of back-propagation learning with second order methods
Sue Becker, Yann Le Cun, et al · 1988
Earlier work this paper cites.
Backpropagation applied to handwritten zip code recognition
Yann LeCun, Bernhard Boser, John S Denker, Donnie Henderson, Richard E Howard, Wayne Hubbard, and Lawrence D Jackel · 1989
Earlier work this paper cites.
On the limited memory bfgs method for large scale optimization
Dong C Liu and Jorge Nocedal · 1989
Earlier work this paper cites.
Variable metric method for minimization
William C Davidon · 1991
Earlier work this paper cites.
Sizing and least-change secant methods
John E Dennis, Jr and Henry Wolkowicz · 1993
Earlier work this paper cites.
A limited memory algorithm for bound constrained optimization
Richard H Byrd, Peihuang Lu, Jorge Nocedal, and Ciyou Zhu · 1995
Earlier work this paper cites.
If quasi-newton then why not quasi-cauchy
JL Nazareth · 1995
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Yann LeCun, Léon Bottou, Yoshua Bengio, and Patrick Haffner · 1998
Earlier work this paper cites.
On the momentum term in gradient descent learning algorithms
Ning Qian · 1999
Earlier work this paper cites.
The quasi-cauchy relation and diagonal updating
Mingfa Zhu, John Lawrence Nazareth, and Henry Wolkowicz · 1999
Earlier work this paper cites.
Online convex programming and generalized infinitesimal gradient ascent
Martin Zinkevich · 2003
Earlier work this paper cites.
Reducing the dimensionality of data with neural networks
Geoffrey E Hinton and Ruslan R Salakhutdinov · 2006
Earlier work this paper cites.
Numerical optimization
Jorge Nocedal and Stephen Wright · 2006
Earlier work this paper cites.
A stochastic quasi-newton method for online convex optimization
Nicol N Schraudolph, Jin Yu, and Simon Günter · 2007
Earlier work this paper cites.
The tradeoffs of large scale learning
Léon Bottou and Olivier Bousquet · 2008
Cited alongside, same era.
SGD-QN: Careful quasi-newton stochastic gradient descent
Antoine Bordes, Léon Bottou, and Patrick Gallinari · 2009
Cited alongside, same era.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Cited alongside, same era.
Learning multiple layers of features from tiny images
Alex Krizhevsky and Geoffrey Hinton · 2009
Cited alongside, same era.
Deep learning via hessian-free optimization
James Martens · 2010
Cited alongside, same era.
Rectified linear units improve restricted boltzmann machines
Vinod Nair and Geoffrey E Hinton · 2010
Cited alongside, same era.
End-to-end sequence labeling via bi-directional LSTM-CNNs-CRF
Xuezhe Ma and Eduard Hovy · 2016
Later among the works it cites.
An overview of gradient descent optimization algorithms
Sebastian Ruder · 2016
Later among the works it cites.
Rethinking the inception architecture for computer vision
Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jon Shlens, and Zbigniew Wojna · 2016
Later among the works it cites.
SGDR: Stochastic gradient descent with warm restarts
Ilya Loshchilov and Frank Hutter · 2017
Later among the works it cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Later among the works it cites.
Stochastic quasi-newton methods for nonconvex stochastic optimization
Xiao Wang, Shiqian Ma, Donald Goldfarb, and Wei Liu · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
On the use of stochastic hessian information in optimization methods for machine learning
Richard H Byrd, Gillian M Chin, Will Neveitt, and Jorge Nocedal · 2011
Cited alongside, same era.
Adaptive subgradient methods for online learning and stochastic optimization
John Duchi, Elad Hazan, and Yoram Singer · 2011
Cited alongside, same era.
Learning recurrent neural networks with hessian-free optimization
James Martens and Ilya Sutskever · 2011
Cited alongside, same era.
Deep neural networks for acoustic modeling in speech recognition: The shared views of four research groups
Geoffrey Hinton, Li Deng, Dong Yu, George E Dahl, Abdel-rahman Mohamed, Navdeep Jaitly, Andrew Senior, Vincent Vanhoucke, Patrick Nguyen, Tara N Sainath, et al · 2012
Cited alongside, same era.
Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude
Tijmen Tieleman and Geoffrey Hinton · 2012
Cited alongside, same era.
Adadelta: an adaptive learning rate method
Matthew D Zeiler · 2012
Cited alongside, same era.
Later among the works it cites.
The marginal value of adaptive gradient methods in machine learning
Ashia C Wilson, Rebecca Roelofs, Mitchell Stern, Nati Srebro, and Benjamin Recht · 2017
Later among the works it cites.
Aggregated residual transformations for deep neural networks
Saining Xie, Ross Girshick, Piotr Dollár, Zhuowen Tu, and Kaiming He · 2017
Later among the works it cites.
Shampoo: Preconditioned stochastic tensor optimization
Vineet Gupta, Tomer Koren, and Yoram Singer · 2018
Later among the works it cites.
Implementation of stochastic quasi-newton’s method in pytorch
Yingkai Li and Huidong Liu · 2018
Later among the works it cites.
On the convergence of adam and beyond
Sashank J Reddi, Satyen Kale, and Sanjiv Kumar · 2018
Later among the works it cites.
Efficient full-matrix adaptive regularization
Naman Agarwal, Brian Bullins, Xinyi Chen, Elad Hazan, Karan Singh, Cyril Zhang, and Yi Zhang · 2019
Later among the works it cites.
Memory efficient adaptive optimization
Rohan Anil, Vineet Gupta, Tomer Koren, and Yoram Singer · 2019
Later among the works it cites.
On the convergence of a class of adam-type algorithms for non-convex optimization
X Chen, M Hong, S Liu, and R Sun · 2019
Later among the works it cites.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter · 2019
Later among the works it cites.
Adaptive gradient methods with dynamic bound of learning rate
Liangchen Luo, Yuanhao Xiong, Yan Liu, and Xu Sun · 2019
Later among the works it cites.
Flowseq: Non-autoregressive conditional sequence generation with generative flow
Xuezhe Ma, Chunting Zhou, Xian Li, Graham Neubig, and Eduard Hovy · 2019
Later among the works it cites.
FairSeq: A fast, extensible toolkit for sequence modeling
Myle Ott, Sergey Edunov, Alexei Baevski, Angela Fan, Sam Gross, Nathan Ng, David Grangier, and Michael Auli · 2019
Later among the works it cites.
Extreme tensoring for low-memory preconditioning
Xinyi Chen, Naman Agarwal, Elad Hazan, Cyril Zhang, and Yi Zhang · 2020
Closest in time.
On the variance of the adaptive learning rate and beyond
Liyuan Liu, Haoming Jiang, Pengcheng He, Weizhu Chen, Xiaodong Liu, Jianfeng Gao, and Jiawei Han · 2020
Closest in time.
ADAHESSIAN: An adaptive second order optimizer for machine learning
Zhewei Yao, Amir Gholami, Sheng Shen, Kurt Keutzer, and Michael W Mahoney · 2020
Closest in time.
EAdam optimizer: How e p s i l o n epsilon impact adam
Wei Yuan and Kai-Xin Gao · 2020
Closest in time.
Why are adaptive methods good for attention models?
Jingzhao Zhang, Sai Praneeth Karimireddy, Andreas Veit, Seungyeon Kim, Sashank Reddi, Sanjiv Kumar, and Suvrit Sra · 2020
Closest in time.
Adabelief optimizer: Adapting stepsizes by the belief in observed gradients
Juntang Zhuang, Tommy Tang, Yifan Ding, Sekhar C Tatikonda, Nicha Dvornek, Xenophon Papademetris, and James Duncan · 2020
Closest in time.