Fetching the paper…
Reading the bibliography…
We propose a stochastic variant of the classical Polyak step-size (Polyak, 1987) commonly used in the subgradient method.
Stochastic gradient descent with polyak’s learning rate
Oberman, A. M. and Prazeres, M. (2019) · 1903
Earlier work this paper cites.
Revisiting the polyak step size
Hazan, E. and Kakade, S. (2019) · 1905
Earlier work this paper cites.
Loizou, N. and Richtárik, P. (2019) · 1905
Earlier work this paper cites.
On the variance of the adaptive learning rate and beyond
Liu, L., Jiang, H., He, P., Chen, W., Liu, X., Gao, J., and Han, J. (2019) · 1908
Earlier work this paper cites.
Angenäherte auflösung von systemen linearer gleichungen
Kaczmarz, S. (1937) · 1937
Earlier work this paper cites.
A stochastic approximation method
Robbins, H. and Monro, S. (1951) · 1951
Earlier work this paper cites.
On Cezari’s convergence of the steepest descent method for approximating saddle point of convex-concave functions
Nemirovski, A. and Yudin, D. B. (1978) · 1978
Earlier work this paper cites.
Problem complexity and method efficiency in optimization
Nemirovski, A. and Yudin, D. B. (1983) · 1983
Earlier work this paper cites.
Introduction to optimization. translations series in mathematics and engineering
Polyak, B. (1987) · 1987
Earlier work this paper cites.
Neuro-Dynamic Programming
Bertsekas, D. P. and Tsitsiklis, J. N. (1996) · 1996
Earlier work this paper cites.
On the Lambert W function
Corless, R. M., Gonnet, G. H., Hare, D. E., Jeffrey, D. J., and Knuth, D. E. (1996) · 1996
Earlier work this paper cites.
Subgradient methods
Boyd, S., Xiao, L., and Mutapcic, A. (2003) · 2005
Earlier work this paper cites.
Convexity, classification, and risk bounds
Bartlett, P. L., Jordan, M. I., and McAuliffe, J. D. (2006) · 2006
Earlier work this paper cites.
Pegasos: primal estimated subgradient solver for SVM
Shalev-Shwartz, S., Singer, Y., and Srebro, N. (2007) · 2007
Earlier work this paper cites.
Robust stochastic approximation approach to stochastic programming
Nemirovski, A., Juditsky, A., Lan, G., and Shapiro, A. (2009) · 2009
Earlier work this paper cites.
A randomized Kaczmarz algorithm with exponential convergence
Strohmer, T. and Vershynin, R. (2009) · 2009
Earlier work this paper cites.
LIBSVM: A library for support vector machines
Chang, C.-C. and Lin, C.-J. (2011) · 2011
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
Duchi, J., Hazan, E., and Singer, Y. (2011) · 2011
Earlier work this paper cites.
Non-asymptotic analysis of stochastic approximation algorithms for machine learning
Moulines, E. and Bach, F. R. (2011) · 2011
Earlier work this paper cites.
Hogwild: A lock-free approach to parallelizing stochastic gradient descent
Recht, B., Re, C., Wright, S., and Niu, F. (2011) · 2011
Earlier work this paper cites.
11 projected Newton-type methods in machine learning
Schmidt, M., Kim, D., and Sra, S. (2011) · 2011
Cited alongside, same era.
Making gradient descent optimal for strongly convex stochastic optimization
Rakhlin, A., Shamir, O., and Sridharan, K. (2012) · 2012
Cited alongside, same era.
Stochastic first-and zeroth-order methods for nonconvex stochastic programming
Ghadimi, S. and Lan, G. (2013) · 2013
Cited alongside, same era.
Fast convergence of stochastic gradient descent under a strong growth condition
Schmidt, M. and Roux, N. (2013) · 2013
Cited alongside, same era.
Stochastic gradient descent for non-smooth optimization: Convergence results and optimal averaging schemes
Shamir, O. and Zhang, T. (2013) · 2013
Cited alongside, same era.
Nutini, J., Laradji, I., and Schmidt, M. (2017) · 2017
Later among the works it cites.
Reflections on random kitchen sinks - arg min blog
Rahimi, A. and Recht, B. (2017) · 2017
Later among the works it cites.
Understanding deep learning requires rethinking generalization
Zhang, C., Bengio, S., Hardt, M., Recht, B., and Vinyals, O. (2017) · 2017
Later among the works it cites.
Optimization methods for large-scale machine learning
Bottou, L., Curtis, F. E., and Nocedal, J. (2018) · 2018
Later among the works it cites.
The power of interpolation: Understanding the effectiveness of sgd in modern over-parametrized learning
Ma, S., Bassily, R., and Belkin, M. (2018) · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Hazan, E. and Kale, S. (2014) · 2014
Cited alongside, same era.
Rmsprop and equilibrated adaptive learning rates for nonconvex optimization
Bengio, Y. (2015) · 2015
Cited alongside, same era.
Randomized iterative methods for linear systems
Gower, R. and Richtárik, P. (2015) · 2015
Cited alongside, same era.
Stop wasting my gradients: Practical SVRG
Harikandeh, R., Ahmed, M. O., Virani, A., Schmidt, M., Konečnỳ, J., and Sallinen, S. (2015) · 2015
Cited alongside, same era.
Adam: A method for stochastic optimization
Kingma, D. and Ba, J. (2015) · 2015
Cited alongside, same era.
Train faster, generalize better: stability of stochastic gradient descent
Hardt, M., Recht, B., and Singer, Y. (2016) · 2016
Cited alongside, same era.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J. (2016) · 2016
Cited alongside, same era.
SGD and hogwild! Convergence without the bounded gradients assumption
Nguyen, L., Nguyen, P. H., van Dijk, M., Richtárik, P., Scheinberg, K., and Takáč, M. (2018) · 2018
Later among the works it cites.
L4: Practical loss-based stepsize adaptation for deep learning
Rolinek, M. and Martius, G. (2018) · 2018
Later among the works it cites.
The implicit bias of gradient descent on separable data
Soudry, D., Hoffer, E., Nacson, M. S., Gunasekar, S., and Srebro, N. (2018) · 2018
Later among the works it cites.
The importance of better models in stochastic optimization
Asi, H. and Duchi, J. C. (2019) · 2019
Later among the works it cites.
On the linear convergence of the stochastic gradient method with constant step-size
Cevher, V. and Vũ, B. C. (2019) · 2019
Later among the works it cites.
SGD: General analysis and improved rates
Gower, R. M., Loizou, N., Qian, X., Sailanbayev, A., Shulgin, E., and Richtárik, P. (2019) · 2019
Later among the works it cites.
Stochastic gradient descent for nonconvex learning without bounded gradient assumptions
Lei, Y., Hu, T., Li, G., and Tang, K. (2019) · 2019
Later among the works it cites.
On the convergence of stochastic gradient descent with adaptive stepsizes
Li, X. and Orabona, F. (2019) · 2019
Later among the works it cites.
Sampling: Design and Analysis: Design and Analysis
Lohr, S. L. (2019) · 2019
Later among the works it cites.
Randomized iterative methods for linear systems: momentum, inexactness and gossip
Loizou, N. (2019) · 2019
Later among the works it cites.
Adaptive gradient methods with dynamic bound of learning rate
Luo, L., Xiong, Y., Liu, Y., and Sun, X. (2019) · 2019
Later among the works it cites.
Adagrad stepsizes: Sharp convergence over nonconvex landscapes
Ward, R., Wu, X., and Bottou, L. (2019) · 2019
Later among the works it cites.
Lookahead optimizer: k steps forward, 1 step back
Zhang, M., Lucas, J., Ba, J., and Hinton, G. E. (2019) · 2019
Later among the works it cites.
Training neural networks for and by interpolation
Berrada, L., Zisserman, A., and Kumar, M. P. (2020) · 2020
Closest in time.
Stochastic reformulations of linear systems: algorithms and convergence theory
Richtárik, P. and Takác, M. (2020) · 2020
Closest in time.