Fetching the paper…
Reading the bibliography…
In this paper, we study large-scale convex optimization algorithms based on the Newton method applied to regularized generalized self-concordant losses, which include logistic regression and softmax regression.
Theory of reproducing kernels
Nachman Aronszajn · 1950
Earlier work this paper cites.
Interior-point polynomial algorithms in convex programming
Arkadii Nemirovskii and Yurii Nesterov · 1994
Earlier work this paper cites.
Using the Nyström method to speed up kernel machines
Christopher K. I. Williams and Matthias Seeger · 2001
Earlier work this paper cites.
Iterative Methods for Sparse Linear Systems
Y. Saad · 2003
Earlier work this paper cites.
Convex optimization
Stephen Boyd and Lieven Vandenberghe · 2004
Earlier work this paper cites.
Kernel Methods for Pattern Analysis
John Shawe-Taylor and Nello Cristianini · 2004
Earlier work this paper cites.
Optimal rates for the regularized least-squares algorithm
A. Caponnetto and E. De Vito · 2007
Earlier work this paper cites.
The tradeoffs of large scale learning
Léon Bottou and Olivier Bousquet · 2008
Earlier work this paper cites.
Random features for large-scale kernel machines
Ali Rahimi and Benjamin Recht · 2008
Earlier work this paper cites.
Self-concordant analysis for logistic regression
Francis Bach · 2010
Earlier work this paper cites.
Newton Methods for Nonlinear Problems: Affine Invariance and Adaptive Algorithms
Peter Deuflhard · 2011
Earlier work this paper cites.
Faster least squares approximation
Petros Drineas, Michael W Mahoney, Shan Muthukrishnan, and Tamás Sarlós · 2011
Earlier work this paper cites.
Fast approximation of matrix coherence and statistical leverage
Petros Drineas, Malik Magdon-Ismail, Michael W Mahoney, and David P Woodruff · 2012
Cited alongside, same era.
Matrix Computations , volume 3
Gene H. Golub and Charles F. Van Loan · 2012
Cited alongside, same era.
An introduction to conditional random fields
Charles Sutton and Andrew McCallum · 2012
Cited alongside, same era.
Improved matrix algorithms via the subsampled randomized hadamard transform
Christos Boutsidis and Alex Gittens · 2013
Cited alongside, same era.
Adaptivity of averaged stochastic gradient descent to local strong convexity for logistic regression
Francis Bach · 2014
Cited alongside, same era.
SAGA: A fast incremental gradient method with support for non-strongly convex composite objectives
Aaron Defazio, Francis Bach, and Simon Lacoste-Julien · 2014
Second-order stochastic optimization for machine learning in linear time
Naman Agarwal, Brian Bullins, and Elad Hazan · 2017
Later among the works it cites.
Katyusha: The first direct acceleration of stochastic gradient methods
Zeyuan Allen-Zhu · 2017
Later among the works it cites.
Newton sketch: A near linear-time optimization algorithm with linear-quadratic convergence
Mert Pilanci and Martin J Wainwright · 2017
Later among the works it cites.
FALKON: An optimal large scale kernel method
Alessandro Rudi, Luigi Carratino, and Lorenzo Rosasco · 2017
Later among the works it cites.
Exact and inexact subsampled newton methods for optimization
Raghu Bollapragada, Richard H. Byrd, and Jorge Nocedal · 2018
Later among the works it cites.
Optimization methods for large-scale machine learning
Léon Bottou, Frank E. Curtis, and Jorge Nocedal · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Convergence rates of sub-sampled Newton methods
Murat A. Erdogdu and Andrea Montanari · 2015
Cited alongside, same era.
Fast randomized kernel ridge regression with statistical guarantees
Ahmed Alaoui and Michael W Mahoney · 2015
Cited alongside, same era.
A universal catalyst for first-order optimization
Hongzhou Lin, Julien Mairal, and Zaid Harchaoui · 2015
Cited alongside, same era.
Less is more: Nyström computational regularization
Alessandro Rudi, Raffaello Camoriano, and Lorenzo Rosasco · 2015
Cited alongside, same era.
A simple practical accelerated method for finite sums
Aaron Defazio · 2016
Cited alongside, same era.
Later among the works it cites.
Accelerated stochastic matrix inversion: general theory and speeding up BFGS rules for faster second-order optimization
Robert Gower, Filip Hanzely, Peter Richtárik, and Sebastian U. Stich · 2018
Later among the works it cites.
Global linear convergence of newton’s method without strong-convexity or lipschitz gradients
Sai Praneeth Karimireddy, Sebastian U. Stich, and Martin Jaggi · 2018
Later among the works it cites.
On fast leverage score sampling and optimal learning
Alessandro Rudi, Daniele Calandriello, Luigi Carratino, and Lorenzo Rosasco · 2018
Later among the works it cites.
Beyond least-squares: Fast rates for regularized empirical risk minimization through self-concordance
Ulysse Marteau-Ferey, Dmitrii Ostrovskii, Francis Bach, and Alessandro Rudi · 2019
Closest in time.
Sub-sampled Newton methods
Farbod Roosta-Khorasani and Michael W. Mahoney · 2019
Closest in time.