Fetching the paper…
Reading the bibliography…
We analyze a batched variant of Stochastic Gradient Descent (SGD) with weighted sampling distribution for smooth and non-smooth objective functions.
“A stochastic approximation method,”
H. Robbins and S. Monroe, · 1951
Earlier work this paper cites.
“Efficient approximation algorithms for semidefinite programs arising from max cut and coloring,”
P. Klein and H.-I. Lu, · 1996
Earlier work this paper cites.
Introductory Lectures on Convex Optimization
Y. Nesterov, · 2004
Earlier work this paper cites.
“Decoding by linear programming,”
E. J. Candès and T. Tao, · 2005
Earlier work this paper cites.
“Regularization tools version 4.0 for matlab 7.3,”
P. C. Hansen, · 2007
Earlier work this paper cites.
“SVM optimization: inverse dependence on training set size,”
S. Shalev-Shwartz and N. Srebro, · 2008
Earlier work this paper cites.
“Robust stochastic approximation approach to stochastic programming,”
A. Nemirovski, A. Juditsky, G. Lan, and A. Shapiro, · 2009
Earlier work this paper cites.
“A randomized Kaczmarz algorithm with exponential convergence,”
T. Strohmer and R. Vershynin, · 2009
Earlier work this paper cites.
“Large-scale machine learning with stochastic gradient descent,”
L. Bottou, · 2010
Earlier work this paper cites.
“Randomized Kaczmarz solver for noisy linear systems,”
D. Needell, · 2010
Earlier work this paper cites.
“Better mini-batch algorithms via accelerated gradient methods,”
A. Cotter, O. Shamir, N. Srebro, and K. Sridharan, · 2011
Earlier work this paper cites.
“Distributed delayed stochastic optimization,”
A. Agarwal and J. C. Duchi, · 2011
Earlier work this paper cites.
“The tradeoffs of large-scale learning,”
L. Bottou and O. Bousquet, · 2011
Cited alongside, same era.
“Non-asymptotic analysis of stochastic approximation algorithms for machine learning,”
F. Bach and E. Moulines, · 2011
Cited alongside, same era.
“Pegasos: Primal estimated sub-gradient solver for svm,”
S. Shalev-Shwartz, Y. Singer, N. Srebro, and A. Cotter, · 2011
Cited alongside, same era.
“Optimal distributed online prediction using mini-batches,”
O. Dekel, R. Gilad-Bachrach, O. Shamir, and L. Xiao, · 2012
Cited alongside, same era.
“Efficiency of coordinate descent methods on huge-scale optimization problems,”
Y. Nesterov, · 2012
Cited alongside, same era.
“Sample size selection in optimization methods for machine learning,”
R. H. Byrd, G. M. Chin, J. Nocedal, and Y. Wu, · 2012
Cited alongside, same era.
“A proximal stochastic gradient method with progressive variance reduction,”
L. Xiao and T. Zhang, · 2014
Later among the works it cites.
“Efficient mini-batch training for stochastic optimization,”
M. Li, T. Zhang, Y. Chen, and A. J. Smola, · 2014
Later among the works it cites.
“Stochastic optimization with importance sampling for regularized loss minimization,”
P. Zhao and T. Zhang, · 2015
Later among the works it cites.
“On optimal probabilities in stochastic coordinate descent methods,”
P. Richtárik and M. Takáč, · 2015
Later among the works it cites.
“Quartz: Randomized dual coordinate ascent with arbitrary sampling,”
Z. Qu, P. Richtarik, and T. Zhang, · 2015
Later among the works it cites.
“Stochastic dual coordinate ascent with adaptive probabilities,”
D. Csiba, Z. Qu, and P. Richtarik, · 2015
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
O. Shamir and T. Zhang, · 2012
Cited alongside, same era.
“Making gradient descent optimal for strongly convex stochastic optimization,”
A. Rakhlin, O. Shamir, and K. Sridharan, · 2012
Cited alongside, same era.
“Mini-batch primal and dual methods for SVMs,”
M. Takac, A. Bijral, P. Richtarik, and N. Srebro, · 2013
Cited alongside, same era.
“Efficient accelerated coordinate descent methods and faster algorithms for solving linear systems,”
Y. T. Lee and A. Sidford, · 2013
Cited alongside, same era.
“Minimizing finite sums with the stochastic average gradient,”
M. Schmidt, N. Roux, and F. Bach, · 2013
Cited alongside, same era.
“Two-subspace projection method for coherent overdetermined linear systems,”
D. Needell and R. Ward, · 2013
Cited alongside, same era.
Later among the works it cites.
“Averaged least-mean-squares: Bias-variance trade-offs and optimal sampling distributions.,”
A. Défossez and F. R. Bach, · 2015
Later among the works it cites.
“Stochastic gradient descent and the randomized kaczmarz algorithm,”
D. Needell, N. Srebro, and R. Ward, · 2016
Closest in time.
“ms2gd: Mini-batch semi-stochastic gradient descent in the proximal setting,”
J. Konecnỳ, J. Liu, P. Richtarik, and M. Takac, · 2016
Closest in time.
“Importance sampling for minibatches,”
D. Csiba and P. Richtarik, · 2016
Closest in time.
“Randomized quasi-newton updates are linearly convergent matrix inversion algorithms,”
R. M. Gower and P. Richtárik, · 2016
Closest in time.
“Weighted sgd for ℓ p \ell_{p} regression with randomized preconditioning,”
J. Yang, Y.-L. Chow, C. Ré, and M. W. Mahoney, · 2016
Closest in time.