Fetching the paper…
Reading the bibliography…
We analyse an iterative algorithm to minimize quadratic functions whose Hessian matrix $H$ is the expectation of a random symmetric $d\times d$ matrix.
Variance-reduced methods for machine learning, Proceedings of the IEEE
[] Gower, R. M., Schmidt, M., Bach, F. and Richtarik, P. (2020) · 1983
Earlier work this paper cites.
Non-Uniform Random Variate Generation
[] Devroye, L. (1986) · 1986
Earlier work this paper cites.
Fast Monte-Carlo algorithms for finding low-rank approximations, Journal of the ACM (JACM)
[] Frieze, A., Kannan, R. and Vempala, S. (2004) · 2004
Earlier work this paper cites.
A fast randomized algorithm for overdetermined linear least-squares regression, Proceedings of the National Academy of Sciences
[] Rokhlin, V. and Tygert, M. (2008) · 2008
Earlier work this paper cites.
The elements of statistical learning: data mining, inference, and prediction
[] Hastie, T., Tibshirani, R. and Friedman, J. (2009) · 2009
Earlier work this paper cites.
Robust stochastic approximation approach to stochastic programming, SIAM Journal on optimization
[] Nemirovski, A., Juditsky, A., Lan, G. and Shapiro, A. (2009) · 2009
Earlier work this paper cites.
A randomized Kaczmarz algorithm with exponential convergence, Journal of Fourier Analysis and Applications
[] Strohmer, T. and Vershynin, R. (2009) · 2009
Earlier work this paper cites.
A stochastic gradient method with an exponential convergence rate for finite training sets, Proceedings of the 25th International Conference on Neural Information Processing Systems-Volume 2
[] Roux, N. L., Schmidt, M. and Bach, F. (2012) · 2012
Earlier work this paper cites.
Non-strongly-convex smooth stochastic approximation with convergence rate O(1/n), Advances in neural information processing systems
[] Bach, F. and Moulines, E. (2013) · 2013
Earlier work this paper cites.
Matrix computations
[] Golub, G. H. and Van Loan, C. F. (2013) · 2013
Earlier work this paper cites.
Accelerating stochastic gradient descent using predictive variance reduction, Advances in neural information processing systems
[] Johnson, R. and Zhang, T. (2013) · 2013
Earlier work this paper cites.
Stochastic dual coordinate ascent methods for regularized loss minimization, Journal of Machine Learning Research
[] Shalev-Shwartz, S. and Zhang, T. (2013) · 2013
Cited alongside, same era.
Stochastic proximal gradient descent with acceleration techniques, Advances in Neural Information Processing Systems
[] Nitanda, A. (2014) · 2014
Cited alongside, same era.
Accelerated proximal stochastic dual coordinate ascent for regularized loss minimization, International Conference on Machine Learning
[] Shalev-Shwartz, S. and Zhang, T. (2014) · 2014
Cited alongside, same era.
Averaged least-mean-squares: Bias-variance trade-offs and optimal sampling distributions, Artificial Intelligence and Statistics
[] Défossez, A. and Bach, F. (2015) · 2015
Cited alongside, same era.
Randomized iterative methods for linear systems, SIAM Journal on Matrix Analysis and Applications
[] Gower, R. M. and Richtárik, P. (2015) · 2015
Cited alongside, same era.
Katyusha: The first direct acceleration of stochastic gradient methods, The Journal of Machine Learning Research
[] Allen-Zhu, Z. (2018) · 2018
Later among the works it cites.
Optimization methods for large-scale machine learning, Siam Review
[] Bottou, L., Curtis, F. E. and Nocedal, J. (2018) · 2018
Later among the works it cites.
Parallelizing stochastic gradient descent for least squares regression: Mini-batching, averaging, and model misspecification, Journal of Machine Learning Research
[] Jain, P., Kakade, S. M., Kidambi, R., Netrapalli, P. and Sidford, A. (2018) · 2018
Later among the works it cites.
An optimal randomized incremental gradient method, Mathematical programming
[] Lan, G. and Zhou, Y. (2018) · 2018
Later among the works it cites.
Efficient simulation of high dimensional Gaussian vectors, Mathematics of Operations Research
[] Kahalé, N. (2019) · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Variance reduced stochastic gradient descent with neighbors, Advances in Neural Information Processing Systems
[] Hofmann, T., Lucchi, A., Lacoste-Julien, S. and McWilliams, B. (2015) · 2015
Cited alongside, same era.
Improved svrg for non-strongly-convex or sum-of-non-convex objectives, International conference on machine learning
[] Allen-Zhu, Z. and Yuan, Y. (2016) · 2016
Cited alongside, same era.
Iterative Hessian sketch: Fast and accurate solution approximation for constrained least-squares, The Journal of Machine Learning Research
[] Pilanci, M. and Wainwright, M. J. (2016) · 2016
Cited alongside, same era.
Harder, better, faster, stronger convergence rates for least-squares regression, The Journal of Machine Learning Research
[] Dieuleveut, A., Flammarion, N. and Bach, F. (2017) · 2017
Cited alongside, same era.
Less than a single pass: Stochastically controlled stochastic gradient, Artificial Intelligence and Statistics
[] Lei, L. and Jordan, M. (2017) · 2017
Cited alongside, same era.
Minimizing finite sums with the stochastic average gradient, Mathematical Programming
[] Schmidt, M., Le Roux, N. and Bach, F. (2017) · 2017
Cited alongside, same era.
Towards closing the gap between the theory and practice of SVRG, Advances in Neural Information Processing Systems
[] Sebbouh, O., Gazagnadou, N., Jelassi, S., Bach, F. and Gower, R. (2019) · 2019
Later among the works it cites.
Don’t jump through hoops and remove those loops: SVRG and Katyusha are better without the outer loop, Proceedings of the 31st International Conference on Algorithmic Learning Theory
[] Kovalev, D., Horváth, S. and Richtárik, P. (2020) · 2020
Closest in time.
Estimate sequences for stochastic composite optimization: Variance reduction, acceleration, and robustness to noise, Journal of Machine Learning Research
[] Kulunchakov, A. and Mairal, J. (2020) · 2020
Closest in time.
Momentum and stochastic momentum for stochastic gradient, newton, proximal point and subspace descent methods, Computational Optimization and Applications
[] Loizou, N. and Richtárik, P. (2020) · 2020
Closest in time.
A proximal stochastic gradient method with progressive variance reduction, SIAM Journal on Optimization
[] Xiao, L. and Zhang, T. (2014) · 2075
Closest in time.