Fetching the paper…
Reading the bibliography…
Large scale optimization problems are ubiquitous in machine learning and data analysis and there is a plethora of algorithms for solving such problems.
A method for the solution of certain problems in least squares
Kenneth Levenberg · 1944
Earlier work this paper cites.
A stochastic approximation method
Herbert Robbins and Sutton Monro · 1951
Earlier work this paper cites.
An algorithm for least-squares estimation of nonlinear parameters
Donald W Marquardt · 1963
Earlier work this paper cites.
Minimization of functions having Lipschitz continuous first partial derivatives
Larry Armijo et al · 1966
Earlier work this paper cites.
Updating quasi-Newton matrices with limited storage
Jorge Nocedal · 1980
Earlier work this paper cites.
Inexact Newton methods
Ron S Dembo, Stanley C Eisenstat, and Trond Steihaug · 1982
Earlier work this paper cites.
On the limited memory BFGS method for large scale optimization
Dong C Liu and Jorge Nocedal · 1989
Earlier work this paper cites.
Generalized linear models
Peter McCullagh and John A. Nelder · 1989
Earlier work this paper cites.
An introduction to the conjugate gradient method without the agonizing pain, 1994
Jonathan Richard Shewchuk · 1994
Earlier work this paper cites.
Neuro-dynamic Programming
Dimitri P. Bertsekas and John N. Tsitsiklis · 1996
Earlier work this paper cites.
Online learning and stochastic approximations
Léon Bottou · 1998
Earlier work this paper cites.
Nonlinear programming
Dimitri P. Bertsekas · 1999
Earlier work this paper cites.
A survey of truncated-Newton methods
Stephen G Nash · 2000
Earlier work this paper cites.
Large scale online learning
Léon Bottou and Yann LeCun · 2004
Earlier work this paper cites.
Convex optimization
Stephen Boyd and Lieven Vandenberghe · 2004
Earlier work this paper cites.
Introductory lectures on convex optimization
Yurii Nesterov · 2004
Earlier work this paper cites.
Fast Monte Carlo algorithms for matrices I: Approximating matrix multiplication
Petros Drineas, Ravi Kannan, and Michael W Mahoney · 2006
Earlier work this paper cites.
Sampling algorithms for ℓ 2 \ell_{2} regression and applications
Petros Drineas, Michael W Mahoney, and S Muthukrishnan · 2006
Cited alongside, same era.
Numerical optimization
Jorge Nocedal and Stephen Wright · 2006
Cited alongside, same era.
Trust region Newton method for logistic regression
Chih-Jen Lin, Ruby C. Weng, and S. Sathiya Keerthi · 2008
Cited alongside, same era.
Optimizing costly functions with simple constraints: A limited-memory projected quasi-Newton algorithm
Mark W. Schmidt, Ewout Berg, Michael P. Friedlander, and Kevin P. Murphy · 2009
Cited alongside, same era.
Blendenpik: Supercharging LAPACK’s least-squares solver
Haim Avron, Petar Maymounkov, and Sivan Toledo · 2010
Cited alongside, same era.
Large-scale machine learning with stochastic gradient descent
Léon Bottou · 2010
Cited alongside, same era.
Fast approximation of matrix coherence and statistical leverage
Petros Drineas, Malik Magdon-Ismail, Michael W Mahoney, and David P Woodruff · 2012
Later among the works it cites.
User-friendly tail bounds for sums of random matrices
Joel A. Tropp · 2012
Later among the works it cites.
An inexact successive quadratic approximation method for convex L-1 regularized optimization
Richard H. Byrd, Jorge Nocedal, and Figen Oztoprak · 2013
Later among the works it cites.
Accelerating stochastic gradient descent using predictive variance reduction
Rie Johnson and Tong Zhang · 2013
Later among the works it cites.
Practical inexact proximal quasi-Newton method with global complexity analysis
Katya Scheinberg and Xiaocheng Tang · 2013
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Deep learning via Hessian-free optimization
James Martens · 2010
Cited alongside, same era.
A quasi-Newton approach to nonsmooth convex optimization problems in machine learning
Jin Yu, SVN Vishwanathan, Simon Günter, and Nicol N. Schraudolph · 2010
Cited alongside, same era.
On the use of stochastic Hessian information in optimization methods for machine learning
Richard H. Byrd, Gillian M. Chin, Will Neveitt, and Jorge Nocedal · 2011
Cited alongside, same era.
Better mini-batch algorithms via accelerated gradient methods
Andrew Cotter, Ohad Shamir, Nati Srebro, and Karthik Sridharan · 2011
Cited alongside, same era.
Faster least squares approximation
Petros Drineas, Michael W Mahoney, S Muthukrishnan, and Tamás Sarlós · 2011
Cited alongside, same era.
Finding structure with randomness: Probabilistic algorithms for constructing approximate matrix decompositions
Nathan Halko, Per-Gunnar Martinsson, and Joel A. Tropp · 2011
Cited alongside, same era.
Mark Schmidt, Nicolas L. Roux, and Francis R. Bach · 2013
Later among the works it cites.
An empirical study of learning rates in deep neural networks for speech recognition
Alan Senior, Georg Heigold, Marc’Aurelio Ranzato, and Ke Yang · 2013
Later among the works it cites.
The lost honour of ℓ 2 \ell_{2} -based regularization
Kees Van Den Doel, Uri Ascher, and Eldad Haber · 2013
Later among the works it cites.
Simultaneous source for non-uniform data variance and missing data
Eldad Haber and Mathias Chung · 2014
Later among the works it cites.
Efficient mini-batch training for stochastic optimization
Mu Li, Tong Zhang, Yuqiang Chen, and Alexander J Smola · 2014
Later among the works it cites.
LSRN: A parallel iterative solver for strongly over-or underdetermined systems
Xiangrui Meng, Michael A Saunders, and Michael W Mahoney · 2014
Later among the works it cites.
Stochastic algorithms for inverse problems involving PDEs and many measurements
Farbod Roosta-Khorasani, Kees van den Doel, and Uri Ascher · 2014
Later among the works it cites.
Convergence rates of sub-sampled newton methods
Murat A. Erdogdu and Andrea Montanari · 2015
Later among the works it cites.
Newton sketch: A linear-time optimization algorithm with linear-quadratic convergence
Mert Pilanci and Martin J. Wainwright · 2015
Later among the works it cites.
Implementing randomized matrix algorithms in parallel and distributed environments
Jiyan Yang, Xiangrui Meng, and Michael W Mahoney · 2015
Later among the works it cites.
Sub-sampled Newton methods II: Local convergence rates
Farbod Roosta-Khorasani and Michael W. Mahoney · 2016
Closest in time.