Fetching the paper…
Reading the bibliography…
We consider the problem of finding the minimizer of a convex function $F: \mathbb R^d \rightarrow \mathbb R$ of the form $F(w) := \sum_{i=1}^n f_i(w) + R(w)$ where a low-rank factorization of $\nabla^2 f_i(w)$ is readily available.
“Inexact Newton methods”
Ron˜S. Dembo, Stanley˜C. Eisenstat and Trond Steihaug · 1982
Earlier work this paper cites.
“On the limited memory BFGS method for large scale optimization”, 1989, pp. 503–528
Dong˜C. Liu and Jorge Nocedal · 1989
Earlier work this paper cites.
“Inductive principles of the search for empirical dependences (methods based on weak convergence of probability measures)”
Vladimir˜N. Vapnik · 1989
Earlier work this paper cites.
“Regression shrinkage and selection via the Lasso”
Robert Tibshirani · 1996
Earlier work this paper cites.
“The Elements of Statistical Learning”
Jerome Friedman, Trevor Hastie and Robert Tibshirani · 2001
Earlier work this paper cites.
“Stability and generalization”
Olivier Bousquet and Andr“’e Elisseeff · 2002
Earlier work this paper cites.
“Convex optimization”
Stephen Boyd and Lieven Vandenberghe · 2004
Earlier work this paper cites.
“Numerical optimization”
Jorge Nocedal and Stephen Wright · 2006
Earlier work this paper cites.
“Relative-error CUR matrix decompositions”
Petros Drineas, Michael˜W. Mahoney and S. Muthukrishnan · 2008
Earlier work this paper cites.
“Deep learning via Hessian-free optimization”
James Martens · 2010
Earlier work this paper cites.
“On the use of stochastic Hessian information in optimization methods for machine learning”
Richard˜H. Byrd, Gillian˜M. Chin, Will Neveitt and Jorge Nocedal · 2011
Earlier work this paper cites.
“Sparse sums of positive semidefinite matrices”
Marcel˜Kenji Carli˜Silva, Nicholas J.˜A. Harvey and Cristiane˜M. Sato · 2011
Earlier work this paper cites.
Michael˜W. Mahoney · 2011
Cited alongside, same era.
“Krylov subspace descent for deep learning”
Oriol Vinyals and Daniel Povey · 2011
Cited alongside, same era.
“Fast approximation of matrix coherence and statistical leverage”
Petros Drineas, Malik Magdon-Ismail, Michael˜W. Mahoney and David˜P. Woodruff · 2012
Cited alongside, same era.
“Matrix Computations”
Gene˜H. Golub and Charles˜F. Van˜Loan · 2012
Cited alongside, same era.
“Metric learning: a survey”
Brian Kulis · 2012
Cited alongside, same era.
“Efficient accelerated coordinate descent methods and faster algorithms for solving linear systems”
“Ridge leverage scores for low-rank approximation”
Michael˜B. Cohen, Cameron Musco and Christopher Musco · 2015
Later among the works it cites.
“Uniform sampling for matrix approximation”
Michael˜B. Cohen et al · 2015
Later among the works it cites.
“Convergence rates of sub-sampled Newton methods”
Murat˜A. Erdogdu and Andrea Montanari · 2015
Later among the works it cites.
“Randomized approximation of the Gram matrix: Exact computation and probabilistic bounds”
John˜T. Holodnak and Ilse C.˜F. Ipsen · 2015
Later among the works it cites.
“Newton Sketch: a linear-time optimization algorithm with linear-quadratic convergence”
Mert Pilanci and Martin˜J. Wainwright · 2015
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yin˜Tat Lee and Aaron Sidford · 2013
Cited alongside, same era.
“Introductory Lectures on Convex Optimization: A Basic Course”
Yurii Nesterov · 2013
Cited alongside, same era.
“The nature of statistical learning theory”
Vladimir˜N. Vapnik · 2013
Cited alongside, same era.
“Theory of convex optimization for machine learning”
S“’ebastien Bubeck · 2014
Cited alongside, same era.
“Sketching as a tool for numerical linear algebra”
David˜P. Woodruff · 2014
Cited alongside, same era.
“Fast randomized kernel ridge gegression with statistical guarantees”
Ahmed˜El Alaoui and Michael˜W. Mahoney · 2015
Cited alongside, same era.
Joel˜A. Tropp · 2015
Later among the works it cites.
“Second order stochastic optimization in linear time”
Naman Agarwal, Brian Bullins and Elad Hazan · 2016
Closest in time.
“Sub-sampled Newton methods I: globally convergent algorithms”
Farbod Roosta-Khorasani and Michael˜W. Mahoney · 2016
Closest in time.
“Sub-sampled Newton methods II: local convergence rates”
Farbod Roosta-Khorasani and Michael˜W Mahoney · 2016
Closest in time.
“Implementing randomized matrix algorithms in parallel and distributed environments”
Jiyan Yang, Xiangrui Meng and Michael˜W. Mahoney · 2016
Closest in time.
“On the generalization ability of on-line learning algorithms”
Nicolo Cesa-Bianchi, Alex Conconi and Claudio Gentile · 2057
Closest in time.