Fetching the paper…
Reading the bibliography…
We present two new remarkably simple stochastic second-order methods for minimizing the average of a very large number of sufficiently smooth and strongly convex functions.
A stochastic approximation method
H. Robbins and S. Monro · 1951
Earlier work this paper cites.
On the limited memory BFGS method for large scale optimization
Dong C. Liu and Jorge Nocedal · 1989
Earlier work this paper cites.
Optimization transfer using surrogate objective functions
Kenneth Lange, David R. Hunter, and Ilsoon Yang · 2000
Earlier work this paper cites.
Nonlinear programming: analysis and methods
Mordecai Avriel · 2003
Earlier work this paper cites.
Introductory lectures on convex optimization: a basic course (Applied Optimization)
Yurii Nesterov · 2004
Earlier work this paper cites.
Cubic regularization of Newton method and its global performance
Yurii Nesterov and Boris T. Polyak · 2006
Earlier work this paper cites.
A stochastic gradient method with an exponential convergence rate for finite training sets
Nicolas L. Roux, Mark Schmidt, and Francis R. Bach · 2012
Earlier work this paper cites.
Accelerating stochastic gradient descent using predictive variance reduction
Rie Johnson and Tong Zhang · 2013
Earlier work this paper cites.
Optimization with first-order surrogate functions
Julien Mairal · 2013
Earlier work this paper cites.
Stochastic dual coordinate ascent methods for regularized loss
Shai Shalev-Shwartz and Tong Zhang · 2013
Earlier work this paper cites.
SAGA: A fast incremental gradient method with support for non-strongly convex composite objectives
Aaron Defazio, Francis Bach, and Simon Lacoste-Julien · 2014
Earlier work this paper cites.
Finito: A faster, permutable incremental gradient method for big data problems
Aaron Defazio, Tiberio Caetano, and Justin Domke · 2014
Earlier work this paper cites.
A universal catalyst for first-order optimization
Hongzhou Lin, Julien Mairal, and Zaid Harchaoui · 2015
Earlier work this paper cites.
A multi-batch L-BFGS method for machine learning
Albert S. Berahas, Jorge Nocedal, and Martin Takáč · 2016
Earlier work this paper cites.
Stochastic block BFGS: Squeezing more curvature out of data
Robert Gower, Donald Goldfarb, and Peter Richtárik · 2016
Earlier work this paper cites.
A proximal stochastic quasi-Newton algorithm
Luo Luo, Zihao Chen, Zhihua Zhang, and Wu-Jun Li · 2016
Earlier work this paper cites.
A linearly-convergent stochastic L-BFGS algorithm
Philipp Moritz, Robert Nishihara, and Michael Jordan · 2016
Cited alongside, same era.
A stochastic successive minimization method for nonsmooth nonconvex optimization with applications to transceiver design in wireless communication networks
Meisam Razaviyayn, Maziar Sanjabi, and Zhi-Quan Luo · 2016
Cited alongside, same era.
A superlinearly-convergent proximal Newton-type method for the optimization of finite sums
Anton Rodomanov and Dmitry Kropotov · 2016
Cited alongside, same era.
Katyusha: The first direct acceleration of stochastic gradient methods
Zeyuan Allen-Zhu · 2017
Cited alongside, same era.
Stochastic optimization with variance reduction for infinite datasets with finite sum structure
Alberto Bietti and Julien Mairal · 2017
Cited alongside, same era.
Surpassing gradient descent provably: A cyclic incremental method with linear convergence rate
Aryan Mokhtari, Mert Gürbüzbalaban, and Alejandro Ribeiro · 2018
Later among the works it cites.
SGD and Hogwild! Convergence without the bounded gradients assumption
Lam Nguyen, Phuong Ha Nguyen, Marten van Dijk, Peter Richtárik, Katya Scheinberg, and Martin Takáč · 2018
Later among the works it cites.
Stochastic cubic regularization for fast nonconvex optimization
Nilesh Tripuraneni, Mitchell Stern, Chi Jin, Jeffrey Regier, and Michael I. Jordan · 2018
Later among the works it cites.
Adaptive stochastic variance reduction for subsampled Newton method with cubic regularization
Junyu Zhang, Lin Xiao, and Shuzhong Zhang · 2018
Later among the works it cites.
Stochastic variance-reduced cubic regularized Newton method
Dongruo Zhou, Pan Xu, and Quanquan Gu · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jonas Moritz Kohler and Aurelien Lucchi · 2017
Cited alongside, same era.
SARAH: A novel method for machine learning problems using stochastic recursive gradient
Lam M. Nguyen, Jie Liu, Katya Scheinberg, and Martin Takáč · 2017
Cited alongside, same era.
Stochastic recursive gradient algorithm for nonconvex optimization
Lam M. Nguyen, Jie Liu, Katya Scheinberg, and Martin Takáč · 2017
Cited alongside, same era.
Proximal-proximal-gradient method
Ernest K. Ryu and Wotao Yin · 2017
Cited alongside, same era.
Exact and inexact subsampled Newton methods for optimization
Raghu Bollapragada, Richard H. Byrd, and Jorge Nocedal · 2018
Cited alongside, same era.
Randomized block cubic Newton method
Nikita Doikov and Peter Richtárik · 2018
Cited alongside, same era.
MISSO: Minimization by incremental stochastic surrogate for large-scale nonconvex optimization
Belhal Karimi and Eric Moulines · 2018
Cited alongside, same era.
Nikita Doikov and Yurii Nesterov · 2019
Closest in time.
A unified theory of SGD: Variance reduction, sampling, quantization and coordinate descent
Eduard Gorbunov, Filip Hanzely, and Peter Richtárik · 2019
Closest in time.
RSN: Randomized subspace Newton
Robert M. Gower, Dmitry Kovalev, Felix Lieder, and Peter Richtárik · 2019
Closest in time.
SGD: General analysis and improved rates
Robert Mansel Gower, Nicolas Loizou, Xun Qian, Alibek Sailanbayev, Egor Shulgin, and Peter Richtárik · 2019
Closest in time.
One method to rule them all: variance reduction for data, parameters and many new methods
Filip Hanzely and Peter Richtárik · 2019
Closest in time.
Don’t jump through hoops and remove those loops: SVRG and Katyusha are better without the outer loop
Dmitry Kovalev, Samuel Horváth, and Peter Richtárik · 2019
Closest in time.
A stochastic decoupling method for minimizing the sum of smooth and non-smooth functions
Konstantin Mishchenko and Peter Richtárik · 2019
Closest in time.
MISO is making a comeback with better proofs and rates
Xun Qian, Alibek Sailanbayev, Konstantin Mishchenko, and Peter Richtárik · 2019
Closest in time.
How good is SGD with random shuffling?
Itay Safran and Ohad Shamir · 2019
Closest in time.
Stochastic recursive variance-reduced cubic regularization methods
Dongruo Zhou and Quanquan Gu · 2019
Closest in time.