Fetching the paper…
Reading the bibliography…
First-order stochastic methods are the state-of-the-art in large-scale machine learning optimization owing to efficient per-iteration complexity.
A stochastic approximation method
Herbert Robbins and Sutton Monro · 1951
Earlier work this paper cites.
The convergence of a class of double-rank minimization algorithms: 2. the new algorithm
Charles G. Broyden · 1970
Earlier work this paper cites.
A new approach to variable metric algorithms
Roger Fletcher · 1970
Earlier work this paper cites.
A family of variable-metric methods derived by variational means
Donald Goldfarb · 1970
Earlier work this paper cites.
Conditioning of quasi-Newton methods for function minimization
David F. Shanno · 1970
Earlier work this paper cites.
A method of solving a convex programming problem with convergence rate O(1/k2)
Yurii Nesterov · 1983
Earlier work this paper cites.
Real-sim, 1997
Andrew McCallum · 1997
Earlier work this paper cites.
The MNIST database of handwritten digits
Yann LeCun and Corinna Cortes · 1998
Earlier work this paper cites.
Comparative accuracies of artificial neural networks and discriminant analysis in predicting forest cover types from cartographic variables
Jock A. Blackard and Denis J. Dean · 1999
Earlier work this paper cites.
Interior point polynomial time methods in convex programming
Arkadi Nemirovski · 2004
Earlier work this paper cites.
Numerical Optimization
Jorge Nocedal and Stephen Wright · 2006
Earlier work this paper cites.
A stochastic quasi-Newton method for online convex optimization
Nicol N. Schraudolph, Jin Yu, and Simon Günter · 2007
Earlier work this paper cites.
Deep learning via Hessian-free optimization
James Martens · 2010
Earlier work this paper cites.
On the use of stochastic Hessian information in optimization methods for machine learning
Richard H. Byrd, Gillian M. Chin, Will Neveitt, and Jorge Nocedal · 2011
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
John Duchi, Elad Hazan, and Yoram Singer · 2011
Cited alongside, same era.
A stochastic gradient method with an exponential convergence rate for finite training sets
Nicolas L. Roux, Mark Schmidt, and Francis R. Bach · 2012
Cited alongside, same era.
User-friendly tail bounds for sums of random matrices
Joel A Tropp · 2012
Cited alongside, same era.
Accelerating stochastic gradient descent using predictive variance reduction
Rie Johnson and Tong Zhang · 2013
Cited alongside, same era.
UCI machine learning repository, 2013
Moshe Lichman · 2013
Cited alongside, same era.
Iterative row sampling
Mu Li, Gary L. Miller, and Richard Peng · 2013
Cited alongside, same era.
Convergence rates of sub-sampled Newton methods
Murat A. Erdogdu and Andrea Montanari · 2015
Later among the works it cites.
A universal catalyst for first-order optimization
Hongzhou Lin, Julien Mairal, and Zaïd Harchaoui · 2015
Later among the works it cites.
Newton sketch: A linear-time optimization algorithm with linear-quadratic convergence
Mert Pilanci and Martin J. Wainwright · 2015
Later among the works it cites.
Oracle complexity of second-order methods for finite-sum problems
Yossi Arjevani and Ohad Shamir · 2016
Closest in time.
Katyusha: The first direct acceleration of stochastic gradient methods
Zeyuan Allen-Zhu · 2016
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Introductory Lectures on Convex Optimization: A Basic Course
Yurii Nesterov · 2013
Cited alongside, same era.
Stochastic dual coordinate ascent methods for regularized loss minimization
Shai Shalev-Shwartz and Tong Zhang · 2013
Cited alongside, same era.
Linear convergence with condition number independent access of full gradients
Lijun Zhang, Mehrdad Mahdavi, and Rong Jin · 2013
Cited alongside, same era.
Saga: A fast incremental gradient method with support for non-strongly convex composite objectives
Aaron Defazio, Francis Bach, and Simon Lacoste-Julien · 2014
Cited alongside, same era.
An accelerated proximal coordinate gradient method
Qihang Lin, Zhaosong Lu, and Lin Xiao · 2014
Cited alongside, same era.
RES: regularized stochastic BFGS algorithm
Aryan Mokhtari and Alejandro Ribeiro · 2014
Cited alongside, same era.
Raghu Bollapragada, Richard Byrd, and Jorge Nocedal · 2016
Closest in time.
A stochastic quasi-Newton method for large-scale optimization
Richard H. Byrd, Samantha L. Hansen, Jorge Nocedal, and Yoram Singer · 2016
Closest in time.
Nearly tight oblivious subspace embeddings by trace inequalities
Michael B. Cohen · 2016
Closest in time.
Efficient second order online learning by sketching
Haipeng Luo, Alekh Agarwal, Nicolò Cesa-Bianchi, and John Langford · 2016
Closest in time.
A linearly-convergent stochastic L-BFGS algorithm
Philipp Moritz, Robert Nishihara, and Michael Jordan · 2016
Closest in time.
Accelerated proximal stochastic dual coordinate ascent for regularized loss minimization
Shai Shalev-Shwartz and Tong Zhang · 2016
Closest in time.
Sub-sampled Newton methods with non-uniform sampling
Peng Xu, Jiyan Yang, Farbod Roosta-Khorasani, Christopher Ré, and Michael W. Mahoney · 2016
Closest in time.
Finding approximate local minima faster than gradient descent
Naman Agarwal, Zeyuan Allen-Zhu, Brian Bullins, Elad Hazan, and Tengyu Ma · 2017
Closest in time.
A unifying framework for convergence analysis of approximate Newton methods
Haishan Ye, Luo Luo, and Zhihua Zhang · 2017
Closest in time.