Fetching the paper…
Reading the bibliography…
In recent years, stochastic gradient descent (SGD) methods and randomized linear algebra (RLA) algorithms have been applied to many large-scale problems in machine learning and data analysis.
Condition numbers and equilibration of matrices
A. van der Sluis · 1969
Earlier work this paper cites.
Approximating center points with iterated radon points
K. L. Clarkson, D. Eppstein, G. L. Miller, C. Sturtivant, and S. Teng · 1993
Earlier work this paper cites.
Templates for the Solution of Linear Systems: Building Blocks for Iterative Methods, 2nd Edition
R. Barrett, M. Berry, T. F. Chan, J. Demmel, J. Donato, J. Dongarra, V. Eijkhout, R. Pozo, C. Romine, and H. Van der Vorst · 1994
Earlier work this paper cites.
Iterative Methods for Solving Linear and Nonlinear Equations
C. T. Kelley · 1995
Earlier work this paper cites.
Matrix Computations
G. H. Golub and C. F. Van Loan · 1996
Earlier work this paper cites.
On computation of regression quantiles: Making the Laplacian tortoise faster
S. Portnoy · 1997
Earlier work this paper cites.
The Gaussian hare and the Laplacian tortoise: Computability of squared-error versus absolute-error estimators, with discussion
S. Portnoy and R. Koenker · 1997
Earlier work this paper cites.
Atomic decomposition by basis pursuit
S. S. Chen, D. L. Donoho, and M. A. Saunders · 2001
Earlier work this paper cites.
Iterative Methods for Sparse Linear Systems
Y. Saad · 2003
Earlier work this paper cites.
Large scale online learning
L. Bottou and Y. Le Cun · 2004
Earlier work this paper cites.
Subgradient and sampling algorithms for ℓ 1 \ell_{1} regression
K. L. Clarkson · 2005
Earlier work this paper cites.
Pegasos: Primal estimated sub–gradient solver for SVM
S. Shalev-Shwartz, Y. Singer, and N. Srebro · 2007
Earlier work this paper cites.
The tradeoffs of large scale learning
L. Bottou and O. Bousquet · 2008
Earlier work this paper cites.
SVM optimization: inverse dependence on training set size
S. Shalev-Shwartz and N. Srebro · 2008
Earlier work this paper cites.
Sampling algorithms and coresets for ℓ p \ell_{p} regression
A. Dasgupta, P. Drineas, B. Harb, R. Kumar, and M. W. Mahoney · 2009
Earlier work this paper cites.
Accelerated gradient methods for stochastic optimization and online learning
C. Hu, J. T. Kwok, and W. Pan · 2009
Earlier work this paper cites.
Stochastic methods for ℓ 1 \ell_{1} regularized loss minimization
S. Shalev-Shwartz and A. Tewari · 2009
Cited alongside, same era.
A randomized Kaczmarz algorithm with exponential convergence
T. Strohmer and R. Vershynin · 2009
Cited alongside, same era.
Blendenpik: Supercharging LAPACK’s least-squares solver
H. Avron, P. Maymounkov, and S. Toledo · 2010
Cited alongside, same era.
Large-scale machine learning with stochastic gradient descent
L. Bottou · 2010
Cited alongside, same era.
Composite objective mirror descent
J. C. Duchi, S. Shalev-Shwartz, Y. Singer, and A. Tewari · 2010
Cited alongside, same era.
Faster least squares approximation
P. Drineas, M. W. Mahoney, S. Muthukrishnan, and T. Sarlós · 2011
Cited alongside, same era.
The Fast Cauchy Transform and faster robust linear regression
K. L. Clarkson, P. Drineas, M. Magdon-Ismail, M. W. Mahoney, X. Meng, and D. P. Woodruff · 2013
Later among the works it cites.
Accelerating stochastic gradient descent using predictive variance reduction
R. Johnson and T. Zhang · 2013
Later among the works it cites.
OSNAP: faster numerical linear algebra algorithms via sparser subspace embeddings
J. Nelson and H. L. Nguyen · 2013
Later among the works it cites.
Subspace embeddings and ℓ p \ell_{p} -regression using exponential random variables
D. P. Woodruff and Q. Zhang · 2013
Later among the works it cites.
A statistical perspective on algorithmic leveraging
P. Ma, M. W. Mahoney, and B. Yu · 2014
Later among the works it cites.
LSRN: A parallel iterative solver for strongly over- or under-determined systems
X. Meng, M. A. Saunders, and M. W. Mahoney · 2014
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Adaptive subgradient methods for online learning and stochastic optimization
J. C. Duchi, E. Hazan, and Y. Singer · 2011
Cited alongside, same era.
A unified framework for approximating and clustering data
D. Feldman and M. Langberg · 2011
Cited alongside, same era.
Randomized algorithms for matrices and data
M. W. Mahoney · 2011
Cited alongside, same era.
Hogwild: A lock-free approach to parallelizing stochastic gradient descent
F. Niu, B. Recht, C. Ré, and J. S Wright · 2011
Cited alongside, same era.
Subspace embedding for the ℓ 1 \ell_{1} -norm with applications
C. Sohler and D. P. Woodruff · 2011
Cited alongside, same era.
Improved analysis of the subsampled randomized Hadamard transform
J. A. Tropp · 2011
Cited alongside, same era.
Later among the works it cites.
Stochastic gradient descent, weighted sampling, and the randomized Kaczmarz algorithm
D. Needell, R. Ward, and N. Srebro · 2014
Later among the works it cites.
Iterative Hessian sketch: Fast and accurate solution approximation for constrained least-squares
M. Pilanci and M. J. Wainwright · 2014
Later among the works it cites.
Quantile regression for large-scale applications
J. Yang, X. Meng, and M. W. Mahoney · 2014
Later among the works it cites.
ℓ p \ell_{p} row sampling by lewis weights
M. B. Cohen and R. Peng · 2015
Closest in time.
Statistical Learning with Sparsity: The Lasso and Generalizations
T. Hastie, R. Tibshirani, and M. Wainwright · 2015
Closest in time.
Stochastic optimization with importance sampling
P. Zhao and T. Zhang · 2015
Closest in time.
A stochastic quasi-newton method for large-scale optimization
R. H. Byrd, S. L. Hansen, J. Nocedal, and Y. Singer · 2016
Closest in time.
Nearly tight oblivious subspace embeddings by trace inequalities
M. B. Cohen · 2016
Closest in time.
A self-correcting variable-metric algorithm for stochastic optimization
F. Curtis · 2016
Closest in time.
A linearly-convergent stochastic L-BFGS algorithm
P. Moritz, R. Nishihara, and M. I. Jordan · 2016
Closest in time.