Fetching the paper…
Reading the bibliography…
In this paper, we propose a unified view of gradient-based algorithms for stochastic convex composite optimization by extending the concept of estimate sequence introduced by Nesterov.
Fonctions convexes duales et points proximaux dans un espace hilbertien
Moreau, J.-J · 1962
Earlier work this paper cites.
Proximité et dualité dans un espace hilbertien
Moreau, J.-J · 1965
Earlier work this paper cites.
Convex analysis and minimization algorithms. II
Hiriart-Urruty, J.-B. and Lemaréchal, C · 1996
Earlier work this paper cites.
Introductory Lectures on Convex Optimization: A Basic Course
Nesterov, Y · 2004
Earlier work this paper cites.
A fast iterative shrinkage-thresholding algorithm for linear inverse problems
Beck, A. and Teboulle, M · 2009
Earlier work this paper cites.
Accelerated gradient methods for stochastic optimization and online learning
Hu, C., Pan, W., and Kwok, J. T · 2009
Earlier work this paper cites.
Robust stochastic approximation approach to stochastic programming
Nemirovski, A., Juditsky, A., Lan, G., and Shapiro, A · 2009
Earlier work this paper cites.
Stability selection
Meinshausen, N. and Bühlmann, P · 2010
Earlier work this paper cites.
Stochastic first order methods in smooth convex optimization
Devolder, O · 2011
Earlier work this paper cites.
Information-theoretic lower bounds on the oracle complexity of convex optimization
Agarwal, A., Wainwright, M. J., Bartlett, P. L., and Ravikumar, P. K · 2012
Earlier work this paper cites.
Optimal stochastic approximation algorithms for strongly convex stochastic composite optimization I: A generic algorithmic framework
Ghadimi, S. and Lan, G · 2012
Earlier work this paper cites.
An optimal method for stochastic composite optimization
Lan, G · 2012
Earlier work this paper cites.
Privacy aware learning
Wainwright, M. J., Jordan, M. I., and Duchi, J. C · 2012
Earlier work this paper cites.
Optimal stochastic approximation algorithms for strongly convex stochastic composite optimization II: Shrinking procedures and optimal algorithms
Ghadimi, S. and Lan, G · 2013
Earlier work this paper cites.
Accelerating stochastic gradient descent using predictive variance reduction
Johnson, R. and Zhang, T · 2013
Earlier work this paper cites.
Gradient methods for minimizing composite functions
Nesterov, Y · 2013
Cited alongside, same era.
Saga: A fast incremental gradient method with support for non-strongly convex composite objectives
Defazio, A., Bach, F., and Lacoste-Julien, S · 2014
Cited alongside, same era.
A sparsity preserving stochastic gradient methods for sparse regression
Lin, Q., Chen, X., and Peña, J · 2014
Cited alongside, same era.
Dropout: a simple way to prevent neural networks from overfitting
Srivastava, N., Hinton, G. E., Krizhevsky, A., Sutskever, I., and Salakhutdinov, R · 2014
Cited alongside, same era.
A proximal stochastic gradient method with progressive variance reduction
Xiao, L. and Zhang, T · 2014
Cited alongside, same era.
On the complexity analysis of randomized block-coordinate descent methods
Lu, Z. and Xiao, L · 2015
Non-convex finite-sum optimization via SCSG methods
Lei, L., Ju, C., Chen, J., and Jordan, M. I · 2017
Later among the works it cites.
Sarah: A novel method for machine learning problems using stochastic recursive gradient
Nguyen, L. M., Liu, J., Scheinberg, K., and Takáč, M · 2017
Later among the works it cites.
Minimizing finite sums with the stochastic average gradient
Schmidt, M., Le Roux, N., and Bach, F · 2017
Later among the works it cites.
Optimization methods for large-scale machine learning
Bottou, L., Curtis, F. E., and Nocedal, J · 2018
Later among the works it cites.
On acceleration with noise-corrupted gradients
Cohen, M. B., Diakonikolas, J., and Orecchia, L · 2018
Later among the works it cites.
Spider: Near-optimal non-convex optimization via stochastic path-integrated differential estimator
Fang, C., Li, C. J., Lin, Z., and Zhang, T · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Incremental majorization-minimization optimization with application to large-scale machine learning
Mairal, J · 2015
Cited alongside, same era.
Non-uniform stochastic average gradient method for training conditional random fields
Schmidt, M., Babanezhad, R., Ahmed, M., Defazio, A., Clifton, A., and Sarkar, A · 2015
Cited alongside, same era.
Dimension-free iteration complexity of finite sum optimization problems
Arjevani, Y. and Shamir, O · 2016
Cited alongside, same era.
End-to-end kernel learning with supervised convolutional kernel networks
Mairal, J · 2016
Cited alongside, same era.
Accelerated proximal stochastic dual coordinate ascent for regularized loss minimization
Shalev-Shwartz, S. and Zhang, T · 2016
Cited alongside, same era.
Improving the robustness of deep neural networks via stability training
Zheng, S., Song, Y., Leung, T., and Goodfellow, I · 2016
Cited alongside, same era.
Later among the works it cites.
SGD and Hogwild! convergence without the bounded gradients assumption
Nguyen, L. M., Nguyen, P. H., van Dijk, M., Richtárik, P., Scheinberg, K., and Takáč, M · 2018
Later among the works it cites.
Catalyst acceleration for gradient-based non-convex optimization
Paquette, C., Lin, H., Drusvyatskiy, D., Mairal, J., and Harchaoui, Z · 2018
Later among the works it cites.
Lightweight stochastic optimization for minimizing finite sums with infinite data
Zheng, S. and Kwok, J. T · 2018
Later among the works it cites.
A simple stochastic variance reduced algorithm with fast convergence rates
Zhou, K., Shang, F., and Cheng, J · 2018
Later among the works it cites.
A universally optimal multistage accelerated stochastic gradient method
Aybat, N. S., Fallah, A., Gurbuzbalaban, M., and Ozdaglar, A · 2019
Closest in time.
Don’t jump through hoops and remove those loops: SVRG and Katyusha are better without the outer loop
Kovalev, D., Horvath, S., and Richtarik, P · 2019
Closest in time.
Kulunchakov, A. and Mairal, J · 2019
Closest in time.
Direct acceleration of SAGA using sampled negative momentum
Zhou, K · 2019
Closest in time.