Fetching the paper…
Reading the bibliography…
We present a novel adaptive optimization algorithm for large-scale machine learning problems.
A stochastic approximation method
Herbert Robbins and Sutton Monro · 1951
Earlier work this paper cites.
Quasi-newton methods, motivation and theory
John E Dennis, Jr and Jorge J Moré · 1977
Earlier work this paper cites.
Practical Methods of Optimization
Roger Fletcher · 1987
Earlier work this paper cites.
Numerical Optimization
Jorge Nocedal and Stephen J. Wright · 2006
Earlier work this paper cites.
An estimator for the diagonal of a matrix
C. Bekas, E. Kokiopoulou, and Y. Saad · 2007
Earlier work this paper cites.
Deep learning via hessian-free optimization
James Martens · 2010
Earlier work this paper cites.
On the use of stochastic hessian information in optimization methods for machine learning
Richard H Byrd, Gillian M Chin, Will Neveitt, and Jorge Nocedal · 2011
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
John Duchi, Elad Hazan, and Yoram Singer · 2011
Earlier work this paper cites.
Hogwild: A lock-free approach to parallelizing stochastic gradient descent
Benjamin Recht, Christopher Re, Stephen Wright, and Feng Niu · 2011
Earlier work this paper cites.
Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude
Tijmen Tieleman and Geoffrey Hinton · 2012
Earlier work this paper cites.
Adadelta: an adaptive learning rate method
Matthew D Zeiler · 2012
Earlier work this paper cites.
Accelerating stochastic gradient descent using predictive variance reduction
Rie Johnson and Tong Zhang · 2013
Earlier work this paper cites.
Introductory lectures on convex optimization: A basic course
Yurii Nesterov · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2015
Earlier work this paper cites.
Global convergence of online limited memory bfgs
Aryan Mokhtari and Alejandro Ribeiro · 2015
Earlier work this paper cites.
A multi-batch l-bfgs method for machine learning
Albert S Berahas, Jorge Nocedal, and Martin Takáč · 2016
Cited alongside, same era.
A self-correcting variable-metric algorithm for stochastic optimization
Frank E. Curtis · 2016
Cited alongside, same era.
Sgdr: Stochastic gradient descent with warm restarts
Ilya Loshchilov and Frank Hutter · 2016
Cited alongside, same era.
Online learning rate adaptation with hypergradient descent
Atilim Gunes Baydin, Robert Cornish, David Martinez Rubio, Mark Schmidt, and Frank Wood · 2017
Cited alongside, same era.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter · 2017
Cited alongside, same era.
Entropy-sgd: Biasing gradient descent into wide valleys
Pratik Chaudhari, Anna Choromanska, Stefano Soatto, Yann LeCun, Carlo Baldassi, Christian Borgs, Jennifer Chayes, Levent Sagun, and Riccardo Zecchina · 2019
Later among the works it cites.
Sgd: General analysis and improved rates
Robert Mansel Gower, Nicolas Loizou, Xun Qian, Alibek Sailanbayev, Egor Shulgin, and Peter Richtárik · 2019
Later among the works it cites.
New convergence aspects of stochastic gradient algorithms
Lam M Nguyen, Phuong Ha Nguyen, Peter Richtárik, Katya Scheinberg, Martin Takáč, and Marten van Dijk · 2019
Later among the works it cites.
A newton-based method for nonconvex optimization with fast evasion of saddle points
Santiago Paternain, Aryan Mokhtari, and Alejandro Ribeiro · 2019
Later among the works it cites.
On the convergence of adam and beyond
Sashank J Reddi, Satyen Kale, and Sanjiv Kumar · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
SARAH: a novel method for machine learning problems using stochastic recursive gradient
L. M. Nguyen, J. Liu, K Scheinberg, and M. Takáč · 2017
Cited alongside, same era.
Minimizing finite sums with the stochastic average gradient
Mark Schmidt, Nicolas Le Roux, and Francis Bach · 2017
Cited alongside, same era.
Newton-type methods for non-convex optimization under inexact Hessian information
P. Xu, F. Roosta-Khorasani, and M. W. Mahoney · 2017
Cited alongside, same era.
Sgd and hogwild! convergence without the bounded gradients assumption
Lam Nguyen, Phuong Ha Nguyen, Marten Dijk, Peter Richtárik, Katya Scheinberg, and Martin Takáč · 2018
Cited alongside, same era.
Newton-MR: Newton’s method without smoothness or convexity
F. Roosta, Y. Liu, P. Xu, and M. W. Mahoney · 2018
Cited alongside, same era.
Sub-sampled newton methods
Farbod Roosta-Khorasani and Michael W. Mahoney · 2018
Cited alongside, same era.
Adafactor: Adaptive learning rates with sublinear memory cost
Noam Shazeer and Mitchell Stern · 2018
Cited alongside, same era.
Fast and faster convergence of sgd for over-parameterized models and an accelerated perceptron
Sharan Vaswani, Francis Bach, and Mark Schmidt · 2019
Later among the works it cites.
PyHessian: Neural networks through the lens of the Hessian
Z. Yao, A. Gholami, K. Keutzer, and M. W. Mahoney · 2019
Later among the works it cites.
An investigation of Newton-Sketch and subsampled Newton methods
Albert S. Berahas, Raghu Bollapragada, and Jorge Nocedal · 2020
Later among the works it cites.
Efficient distributed hessian free algorithm for large-scale empirical risk minimization via accumulating sample strategy
Majid Jahani, Xi He, Chenxin Ma, Aryan Mokhtari, Dheevatsa Mudigere, Alejandro Ribeiro, and Martin Takáč · 2020
Later among the works it cites.
Scaling up quasi-newton algorithms: Communication efficient distributed sr1
Majid Jahani, Mohammadreza Nazari, Sergey Rusakov, Albert S Berahas, and Martin Takáč · 2020
Later among the works it cites.
Stochastic polyak step-size for sgd: An adaptive learning rate for fast convergence
Nicolas Loizou, Sharan Vaswani, Issam Laradji, and Simon Lacoste-Julien · 2020
Later among the works it cites.
Adaptive gradient descent without descent
Konstantin Mishchenko and Yura Malitsky · 2020
Later among the works it cites.
Second-order optimization for non-convex machine learning: An empirical study
Peng Xu, Fred Roosta, and Michael W Mahoney · 2020
Later among the works it cites.
Adahessian: An adaptive second order optimizer for machine learning
Zhewei Yao, Amir Gholami, Sheng Shen, Kurt Keutzer, and Michael W Mahoney · 2020
Later among the works it cites.
Fast and safe: accelerated gradient methods with optimality certificates and underestimate sequences
Majid Jahani, Naga Venkata C Gudapati, Chenxin Ma, Rachael Tappenden, and Martin Takáč · 2021
Closest in time.
Sonia: A symmetric blockwise truncated optimization algorithm
Majid Jahani, Mohammadreza Nazari, Rachael Tappenden, Albert Berahas, and Martin Takáč · 2021
Closest in time.