Fetching the paper…
Reading the bibliography…
Many data-fitting applications require the solution of an optimization problem involving a sum of large number of functions of high dimensional parameter.
A method for the solution of certain problems in least squares
Kenneth Levenberg · 1944
Earlier work this paper cites.
A stochastic approximation method
Herbert Robbins and Sutton Monro · 1951
Earlier work this paper cites.
An algorithm for least-squares estimation of nonlinear parameters
Donald W Marquardt · 1963
Earlier work this paper cites.
Updating quasi-Newton matrices with limited storage
Jorge Nocedal · 1980
Earlier work this paper cites.
On the limited memory BFGS method for large scale optimization
Dong C Liu and Jorge Nocedal · 1989
Earlier work this paper cites.
Generalized linear models
Peter McCullagh and John A. Nelder · 1989
Earlier work this paper cites.
An introduction to the conjugate gradient method without the agonizing pain, 1994
Jonathan Richard Shewchuk · 1994
Earlier work this paper cites.
Neuro-dynamic Programming
Dimitri P. Bertsekas and John N. Tsitsiklis · 1996
Earlier work this paper cites.
Online learning and stochastic approximations
Léon Bottou · 1998
Earlier work this paper cites.
A tutorial on support vector machines for pattern recognition
Christopher JC Burges · 1998
Earlier work this paper cites.
Statistical learning theory
Vladimir N. Vapnik · 1998
Earlier work this paper cites.
Nonlinear programming
Dimitri P. Bertsekas · 1999
Earlier work this paper cites.
On optimization techniques for solving nonlinear inverse problems
Eldad Haber, Uri M. Ascher, and Doug Oldenburg · 2000
Earlier work this paper cites.
A finite Newton method for classification
Olvi L Mangasarian · 2002
Earlier work this paper cites.
Learning with kernels: Support vector machines, regularization, optimization, and beyond
Bernhard Schölkopf and Alexander J. Smola · 2002
Earlier work this paper cites.
Large scale online learning
Léon Bottou and Yann LeCun · 2004
Earlier work this paper cites.
Convex optimization
Stephen Boyd and Lieven Vandenberghe · 2004
Earlier work this paper cites.
Introductory lectures on convex optimization
Yurii Nesterov · 2004
Earlier work this paper cites.
A modified finite Newton method for fast solution of large scale linear SVMs
S. Sathiya Keerthi and Dennis DeCoste · 2005
Earlier work this paper cites.
Fast Monte Carlo algorithms for matrices I: Approximating matrix multiplication
Petros Drineas, Ravi Kannan, and Michael W Mahoney · 2006
Earlier work this paper cites.
Numerical optimization
Jorge Nocedal and Stephen Wright · 2006
Cited alongside, same era.
Training a support vector machine in the primal
Olivier Chapelle · 2007
Cited alongside, same era.
Trust region Newton method for logistic regression
Chih-Jen Lin, Ruby C. Weng, and S. Sathiya Keerthi · 2008
Cited alongside, same era.
Matrix mathematics: theory, facts, and formulas
Dennis S Bernstein · 2009
Cited alongside, same era.
Robust stochastic approximation approach to stochastic programming
Arkadi Nemirovski, Anatoli Juditsky, Guanghui Lan, and Alexander Shapiro · 2009
Cited alongside, same era.
Large-scale machine learning with stochastic gradient descent
Léon Bottou · 2010
Cited alongside, same era.
A stochastic gradient method with an exponential convergence rate for finite training sets
Nicolas L. Roux, Mark Schmidt, and Francis R. Bach · 2012
Later among the works it cites.
Matrix analysis
Rajendra Bhatia · 2013
Later among the works it cites.
Revisiting Frank-Wolfe: Projection-free sparse convex optimization
Martin Jaggi · 2013
Later among the works it cites.
Accelerating stochastic gradient descent using predictive variance reduction
Rie Johnson and Tong Zhang · 2013
Later among the works it cites.
Minimizing finite sums with the stochastic average gradient
Mark Schmidt, Nicolas L. Roux, and Francis R. Bach · 2013
Later among the works it cites.
An empirical study of learning rates in deep neural networks for speech recognition
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
David Gross and Vincent Nesme · 2010
Cited alongside, same era.
Deep learning via Hessian-free optimization
James Martens · 2010
Cited alongside, same era.
A quasi-Newton approach to nonsmooth convex optimization problems in machine learning
Jin Yu, SVN Vishwanathan, Simon Günter, and Nicol N. Schraudolph · 2010
Cited alongside, same era.
On the use of stochastic Hessian information in optimization methods for machine learning
Richard H. Byrd, Gillian M. Chin, Will Neveitt, and Jorge Nocedal · 2011
Cited alongside, same era.
Better mini-batch algorithms via accelerated gradient methods
Andrew Cotter, Ohad Shamir, Nati Srebro, and Karthik Sridharan · 2011
Cited alongside, same era.
Finding structure with randomness: Probabilistic algorithms for constructing approximate matrix decompositions
Nathan Halko, Per-Gunnar Martinsson, and Joel A. Tropp · 2011
Cited alongside, same era.
Alan Senior, Georg Heigold, Marc’Aurelio Ranzato, and Ke Yang · 2013
Later among the works it cites.
Stochastic dual coordinate ascent methods for regularized loss
Shai Shalev-Shwartz and Tong Zhang · 2013
Later among the works it cites.
The lost honour of ℓ 2 \ell_{2} -based regularization
Kees Van Den Doel, Uri Ascher, and Eldad Haber · 2013
Later among the works it cites.
Solving large-scale PDE-constrained Bayesian inverse problems with Riemann manifold Hamiltonian Monte Carlo
Tan Bui-Thanh and Mark Girolami · 2014
Later among the works it cites.
Simultaneous source for non-uniform data variance and missing data
Eldad Haber and Mathias Chung · 2014
Later among the works it cites.
Efficient mini-batch training for stochastic optimization
Mu Li, Tong Zhang, Yuqiang Chen, and Alexander J Smola · 2014
Later among the works it cites.
Data completion and stochastic algorithms for PDE inversion problems with many measurements
Farbod Roosta-Khorasani, Kees van den Doel, and Uri Ascher · 2014
Later among the works it cites.
Stochastic algorithms for inverse problems involving PDEs and many measurements
Farbod Roosta-Khorasani, Kees van den Doel, and Uri Ascher · 2014
Later among the works it cites.
Convergence rates of sub-sampled newton methods
Murat A. Erdogdu and Andrea Montanari · 2015
Later among the works it cites.
Newton sketch: A linear-time optimization algorithm with linear-quadratic convergence
Mert Pilanci and Martin J. Wainwright · 2015
Later among the works it cites.
Assessing stochastic algorithms for large scale nonlinear least squares problems using extremal probabilities of linear combinations of gamma random variables
Farbod Roosta-Khorasani, Gábor J. Székely, and Uri Ascher · 2015
Later among the works it cites.
Fast stochastic algorithms for SVD and PCA: Convergence properties and convexity
Ohad Shamir · 2015
Later among the works it cites.
An introduction to matrix concentration inequalities
Joel A Tropp · 2015
Later among the works it cites.
Sub-sampled Newton methods I: Globally convergent algorithms
Farbod Roosta-Khorasani and Michael W. Mahoney · 2016
Closest in time.