Fetching the paper…
Reading the bibliography…
We combine two advanced ideas widely used in optimization for machine learning: shuffling strategy and momentum technique to develop a novel shuffling gradient-based method with momentum, coined Shuffling Momentum Gradient (SMG), for non-convex finite-sum optimization problems.
A stochastic approximation method
Robbins, H. and Monro, S · 1951
Earlier work this paper cites.
Some methods of speeding up the convergence of iteration methods
Polyak, B. T · 1964
Earlier work this paper cites.
A method for unconstrained convex minimization problem with the rate of convergence 𝒪 ( 1 / k 2 ) \mathcal{O}(1/k^{2})
Nesterov, Y · 1983
Earlier work this paper cites.
Gradient-based learning applied to document recognition
LeCun, Y., Bottou, L., Bengio, Y., and Haffner, P · 1998
Earlier work this paper cites.
Convergence rate of incremental subgradient algorithms
Nedić, A. and Bertsekas, D · 2001
Earlier work this paper cites.
Incremental subgradient methods for nondifferentiable optimization
Nedic, A. and Bertsekas, D. P · 2001
Earlier work this paper cites.
Introductory lectures on convex optimization: A basic course , volume 87 of Applied Optimization
Nesterov, Y · 2004
Earlier work this paper cites.
Curiously fast convergence of some stochastic gradient descent algorithms
Bottou, L · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A. and Hinton, G · 2009
Earlier work this paper cites.
Incremental proximal methods for large scale convex optimization
Bertsekas, D · 2011
Earlier work this paper cites.
LIBSVM: A library for support vector machines
Chang, C.-C. and Lin, C.-J · 2011
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
Duchi, J., Hazan, E., and Singer, Y · 2011
Earlier work this paper cites.
Stochastic gradient descent tricks
Bottou, L · 2012
Earlier work this paper cites.
Stochastic first-and zeroth-order methods for nonconvex stochastic programming
Ghadimi, S. and Lan, G · 2013
Earlier work this paper cites.
Accelerating stochastic gradient descent using predictive variance reduction
Johnson, R. and Zhang, T · 2013
Earlier work this paper cites.
Saga: A fast incremental gradient method with support for non-strongly convex composite objectives
Defazio, A., Bach, F., and Lacoste-Julien, S · 2014
Earlier work this paper cites.
ADAM: A Method for Stochastic Optimization
Kingma, D. P. and Ba, J · 2014
Cited alongside, same era.
TensorFlow: Large-scale machine learning on heterogeneous systems, 2015
Abadi, M., Agarwal, A., Barham, P., Brevdo, E., Chen, Z., Citro, C., Corrado, G. S., Davis, A., Dean, J., Devin, M., Ghemawat, S., Goodfellow, I., Harp, A., Irving, G., Isard, M., Jia, Y., Jozefowicz, R., Kaiser, L., Kudlur, M., Levenberg, J., Mané, D., Monga, R., Moore, S., Murray, D., Olah, C., Schuster, M., Shlens, J., Steiner, B., Sutskever, I., Talwar, K., Tucker, P., Vanhoucke, V., Vasudevan, V., Viégas, F., Vinyals, O., Warden, P., Wattenberg, M., Wicke, M., Yu, Y., and Zheng, X · 2015
Cited alongside, same era.
Keras, 2015
Chollet, F. et al · 2015
Cited alongside, same era.
Incorporating nesterov momentum into ADAM
Dozat, T · 2016
Cited alongside, same era.
Without-replacement sampling for stochastic gradient methods
Shamir, O · 2016
Cited alongside, same era.
Random shuffling beats sgd after finite epochs
Haochen, J. and Sra, S · 2019
Later among the works it cites.
Incremental methods for weakly convex optimization
Li, X., Zhu, Z., So, A., and Lee, J. D · 2019
Later among the works it cites.
Convergence analysis of distributed stochastic gradient descent with shuffling
Meng, Q., Chen, W., Wang, Y., Ma, Z.-M., and Liu, T.-Y · 2019
Later among the works it cites.
Sgd without replacement: Sharper rates for general smooth convex functions
Nagaraj, D., Jain, P., and Netrapalli, P · 2019
Later among the works it cites.
New convergence aspects of stochastic gradient algorithms
Nguyen, L. M., Nguyen, P. H., Richtárik, P., Scheinberg, K., Takáč, M., and van Dijk, M · 2019
Later among the works it cites.
Pytorch: An imperative style, high-performance deep learning library
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Sgdr: Stochastic gradient descent with warm restarts, 2017
Loshchilov, I. and Hutter, F · 2017
Cited alongside, same era.
Sarah: A novel method for machine learning problems using stochastic recursive gradient
Nguyen, L. M., Liu, J., Scheinberg, K., and Takáč, M · 2017
Cited alongside, same era.
An overview of gradient descent optimization algorithms, 2017
Ruder, S · 2017
Cited alongside, same era.
Cyclical learning rates for training neural networks
Smith, L. N · 2017
Cited alongside, same era.
Fashion-mnist: a novel image dataset for benchmarking machine learning algorithms
Xiao, H., Rasul, K., and Vollgraf, R · 2017
Cited alongside, same era.
Convergence of variance-reduced stochastic learning under random reshuffling
Ying, B., Yuan, K., and Sayed, A. H · 2017
Cited alongside, same era.
Optimization Methods for Large-Scale Machine Learning
Bottou, L., Curtis, F. E., and Nocedal, J · 2018
Cited alongside, same era.
Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., Desmaison, A., Kopf, A., Yang, E., DeVito, Z., Raison, M., Tejani, A., Chilamkurthy, S., Steiner, B., Fang, L., Bai, J., and Chintala, S · 2019
Later among the works it cites.
Hybrid stochastic gradient descent algorithms for stochastic nonconvex optimization
Tran-Dinh, Q., Pham, N. H., Phan, D. T., and Nguyen, L. M · 2019
Later among the works it cites.
Spiderboost and momentum: Faster variance reduction algorithms
Wang, Z., Ji, K., Zhou, Y., Liang, Y., and Tarokh, V · 2019
Later among the works it cites.
Sgd with shuffling: optimal rates without component convexity and large epoch requirements
Ahn, K., Yun, C., and Sra, S · 2020
Closest in time.
Exponential step sizes for non-convex optimization
Li, L., Zhuang, Z., and Orabona, F · 2020
Closest in time.
Random reshuffling: Simple analysis with vast improvements
Mishchenko, K., Khaled Ragab Bayoumi, A., and Richtárik, P · 2020
Closest in time.
A unified convergence analysis for shuffling-type gradient methods
Nguyen, L. M., Tran-Dinh, Q., Phan, D. T., Nguyen, P. H., and van Dijk, M · 2020
Closest in time.
Closing the convergence gap of sgd without replacement
Rajput, S., Gupta, A., and Papailiopoulos, D · 2020
Closest in time.
How good is sgd with random shuffling?
Safran, I. and Shamir, O · 2020
Closest in time.
Scheduled restart momentum for accelerated stochastic gradient descent
Wang, B., Nguyen, T. M., Bertozzi, A. L., Baraniuk, R. G., and Osher, S. J · 2020
Closest in time.