Fetching the paper…
Reading the bibliography…
Virtually all state-of-the-art methods for training supervised machine learning models are variants of SGD enhanced with a number of additional tricks, such as minibatching, momentum, and adaptive stepsizes.
A stochastic approximation method
Robbins, H. and Monro, S · 1951
Earlier work this paper cites.
On the convergence of the LMS algorithm with adaptive learning rate for linear feedforward networks
Luo, Z.-Q · 1991
Earlier work this paper cites.
A class of unconstrained minimization methods for neural network training
Grippo, L · 1994
Earlier work this paper cites.
Serial and parallel backpropagation convergence via nonmonotone perturbed minimization
Mangasarian, O. and Solodov, M · 1994
Earlier work this paper cites.
Neuro-dynamic programming
Bertsekas, D. P. and Tsitsiklis, J. N · 1996
Earlier work this paper cites.
Gradient convergence in gradient methods with errors
Bertsekas, D. P. and Tsitsiklis, J. N · 2000
Earlier work this paper cites.
Incremental subgradient methods for nondifferentiable optimization
Nedić, A. and Bertsekas, D. P · 2001
Earlier work this paper cites.
Curiously fast convergence of some stochastic gradient descent algorithms
Bottou, L · 2009
Earlier work this paper cites.
Libsvm: a library for support vector machines
Chang, C.-C. and Lin, C.-J · 2011
Earlier work this paper cites.
Practical recommendations for gradient-based training of deep architectures
Bengio, Y · 2012
Earlier work this paper cites.
Making gradient descent optimal for strongly convex stochastic optimization
Rakhlin, A., Shamir, O., and Sridharan, K · 2012
Earlier work this paper cites.
A stochastic gradient method with an exponential convergence rate for finite training sets
Roux, N. L., Schmidt, M., and Bach, F · 2012
Earlier work this paper cites.
Accelerating stochastic gradient descent using predictive variance reduction
Johnson, R. and Zhang, T · 2013
Earlier work this paper cites.
Parallel stochastic gradient algorithms for large-scale matrix completion
Recht, B. and Ré, C · 2013
Cited alongside, same era.
Adaptive networks
Sayed, A. H · 2014
Cited alongside, same era.
Train faster, generalize better: Stability of stochastic gradient descent
Hardt, M., Recht, B., and Singer, Y · 2016
Cited alongside, same era.
Without-replacement sampling for stochastic gradient methods
Shamir, O · 2016
Cited alongside, same era.
On the convergence rate of incremental aggregated gradient algorithms
Gürbüzbalaban, M., Ozdaglar, A., and Parrilo, P. A · 2017
Cited alongside, same era.
Optimization methods for large-scale machine learning
Bottou, L., Curtis, F. E., and Nocedal, J · 2018
Cited alongside, same era.
SGD without replacement: sharper rates for general smooth convex functions
Nagaraj, D., Jain, P., and Netrapalli, P · 2019
Later among the works it cites.
Stochastic learning under random reshuffling with constant step-sizes
Ying, B., Yuan, K., Vlaski, S., and Sayed, A. H · 2019
Later among the works it cites.
A unified theory of SGD: Variance reduction, sampling, quantization and coordinate descent
Gorbunov, E., Hanzely, F., and Richtárik, P · 2020
Later among the works it cites.
Variance-reduced methods for machine learning
Gower, R. M., Schmidt, M., Bach, F., and Richtárik, P · 2020
Later among the works it cites.
Don’t jump through hoops and remove those loops: SVRG and katyusha are better without the outer loop
Kovalev, D., Horváth, S., and Richtárik, P · 2020
Later among the works it cites.
Random reshuffling: Simple analysis with vast improvements
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Surpassing gradient descent provably: A cyclic incremental method with linear convergence rate
Mokhtari, A., Gürbüzbalaban, M., and Ribeiro, A · 2018
Cited alongside, same era.
The complexity of finding stationary points with stochastic gradient descent
Drori, Y. and Shamir, O · 2019
Cited alongside, same era.
SGD: General analysis and improved rates
Gower, R. M., Loizou, N., Qian, X., Sailanbayev, A., Shulgin, E., and Richtárik, P · 2019
Cited alongside, same era.
Convergence rate of incremental gradient and incremental newton methods
Gürbüzbalaban, M., Ozdaglar, A., and Parrilo, P. A · 2019
Cited alongside, same era.
Why random reshuffling beats stochastic gradient descent
Gürbüzbalaban, M., Ozdaglar, A., and Parrilo, P. A · 2019
Cited alongside, same era.
Random shuffling beats SGD after finite epochs
Haochen, J. and Sra, S · 2019
Cited alongside, same era.
Mishchenko, K., Khaled, A., and Richtárik, P · 2020
Later among the works it cites.
A unified convergence analysis for shuffling-type gradient methods
Nguyen, L. M., Tran-Dinh, Q., Phan, D. T., Nguyen, P. H., and van Dijk, M · 2020
Later among the works it cites.
Linear convergence of cyclic SAGA
Park, Y. and Ryu, E. K · 2020
Later among the works it cites.
Closing the convergence gap of SGD without replacement
Rajput, S., Gupta, A., and Papailiopoulos, D · 2020
Later among the works it cites.
How good is SGD with random shuffling?
Safran, I. and Shamir, O · 2020
Later among the works it cites.
Optimization for deep learning: an overview
Sun, R.-Y · 2020
Later among the works it cites.
Variance-reduced stochastic learning under random reshuffling
Ying, B., Yuan, K., and Sayed, A. H · 2020
Later among the works it cites.