Fetching the paper…
Reading the bibliography…
Stochastic gradient descent~(SGD) and its variants have been the dominating optimization methods in machine learning.
Robbins H, Monro S. A stochastic approximation method. The Annals of Mathematical Statistics, 1951, 22: 400-407
1951
Earlier work this paper cites.
Polyak B T. Some methods of speeding up the convergence of iteration methods. USSR Computational Mathematics and Mathematical Physics, 1964, 4:1-17
1964
Earlier work this paper cites.
Nesterov Y E. Introductory lectures on convex optimization: a basic course. Berlin: Springer Science & Business Media, 2004
2004
Earlier work this paper cites.
Ding F, Yang H Z, Liu F. Performance analysis of stochastic gradient algorithms under weak conditions. Sci China Ser F-Inf Sci, 2008, 51: 1269-1280
2008
Earlier work this paper cites.
Deng J, Dong W, Socher R, et al. ImageNet: A large-scale hierarchical image database. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Miami, 2009. 248-255
2009
Earlier work this paper cites.
LeCun Y A, Bottou L, Orr G B, et al. Neural networks: tricks of the trade. Berlin: Springer Science & Business Media, 2012. 9-48
2012
Earlier work this paper cites.
Li M, Andersen D G, Smola A J, et al. Communication efficient distributed machine learning with the parameter server. In: Proceedings of Advances in Neural Information Processing Systems, Montréal, 2014. 19-27
2014
Earlier work this paper cites.
Li M, Zhang T, Chen Y, et al. Efficient mini-batch training for stochastic optimization. In: Proceedings of the ACM SIGKDD international conference on Knowledge discovery and data mining, New York, 2014. 661-670
2014
Earlier work this paper cites.
Hazan E, Levy K, Shalev-Shwartz S. Beyond convexity: stochastic quasi-convex optimization. In: Proceedings of Advances in Neural Information Processing Systems, Montréal, 2015. 1594-1602
2015
Earlier work this paper cites.
Kingma D P, Ba J. Adam: A method for stochastic optimization. In: Proceedings of the International Conference on Learning Representations, San Diego, 2015
2015
Earlier work this paper cites.
He K, Zhang X, Ren S, et al. Deep residual learning for image recognition. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, 2016. 770-778
2016
Earlier work this paper cites.
2017
Earlier work this paper cites.
Guo H, Tang R, Ye Y, et al. DeepFM: a factorization-machine based neural network for CTR prediction. In: Proceedings of International Joint Conference on Artificial Intelligence, Melbourne, 2017. 1725-1731
2017
Earlier work this paper cites.
Hoffer E, Hubara I, Soudry D. Train longer, generalize better: closing the generalization gap in large batch training of neural networks. In: Proceedings of Advances in Neural Information Processing Systems, Long Beach, 2017. 1731-1741
2017
Cited alongside, same era.
Keskar N S, Mudigere D, Nocedal J, et al. On large-batch training for deep learning: generalization gap and sharp minima. In: Proceedings of International Conference on Learning Representations, Toulon, 2017
2017
Cited alongside, same era.
Lian X, Zhang C, Zhang H, et al. Can decentralized algorithms outperform centralized algorithms? A case study for decentralized parallel stochastic gradient descent. In: Proceedings of Advances in Neural Information Processing Systems, Long Beach, 2017. 5330-5340
2017
Cited alongside, same era.
Loshchilov I, Hutter F. SGDR: stochastic gradient descent with warm restarts. In: Proceedings of International Conference on Learning Representations, Toulon, 2017
2017
Yu H, Yang S, Zhu S. Parallel restarted SGD with faster convergence and less communication: Demystifying why model averaging works for deep learning. In: Proceedings of AAAI Conference on Artificial Intelligence, Hawaii, 2019. 5693-5700
2019
Later among the works it cites.
Lin T, Kong L, Stich S U, et al. Extrapolation for large-batch training in deep learning. In: Proceedings of International Conference on Machine Learning, Virtual, 2020. 6094-6104
2020
Closest in time.
Lin T, Stich S U, Patel K K, et al. Don’t use large mini-batches, use local SGD. In: Proceedings of International Conference on Learning Representations, Addis Ababa, 2020
2020
Closest in time.
You Y, Li J, Reddi S, et al. Large batch optimization for deep learning: training bert in 76 minutes. In: Proceedings of International Conference on Learning Representations, Addis Ababa, 2020
2020
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Merity S, Xiong C, Bradbury J, et al. Pointer sentinel mixture models. In: Proceedings of International Conference on Learning Representations, Toulon, 2017
2017
Cited alongside, same era.
2017
Cited alongside, same era.
Merity S, Keskar N S, Socher R. Regularizing and optimizing LSTM language models. In: Proceedings of International Conference on Learning Representations, Vancouver, 2018
2018
Cited alongside, same era.
Ott M, Edunov S, Grangier D, et al. Scaling neural machine translation. In: Proceedings of the Conference on Machine Translation, Brussels, 2018. 1-9
2018
Cited alongside, same era.
Yan Y, Yang T, Li Z, et al. A unified analysis of stochastic momentum methods for deep learning. In: Proceedings of International Joint Conference on Artificial Intelligence, Stockholm, 2018. 2955-2961
2018
Cited alongside, same era.
Chen C, Wang W, Zhang Y, et al. A convergence analysis for a class of practical variance-reduction stochastic gradient MCMC. Sci China Inf Sci, 2019, 62: 012101
2019
Cited alongside, same era.
2019
Cited alongside, same era.
You Y, Hseu J, Ying C, et al. Large-batch training for LSTM and beyond. In: Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis, Denver, 2019. 1-16
2019
Cited alongside, same era.
Zhang J, He T, Sra S, et al. Why gradient clipping accelerates training: a theoretical justification for adaptivity. In: Proceedings of International Conference on Learning Representations, Addis Ababa, 2020
2020
Closest in time.
Dosovitskiy A, Beyer L, Kolesnikov A, et al. An image is worth 16x16 words: Transformers for image recognition at scale. In: Proceedings of the International Conference on Learning Representations, Virtual, 2021
2021
Closest in time.
Gao H, Xu A, Huang H. On the convergence of communication-efficient local SGD for federated learning. In: Proceedings of AAAI Conference on Artificial Intelligence, Virtual, 2021. 7510-7518
2021
Closest in time.
Huo Z, Gu B, Huang H. Large batch optimization for deep learning using new complete layer-wise adaptive rate scaling. In: Proceedings of the AAAI Conference on Artificial Intelligence, Virtual, 2021. 7883-7890
2021
Closest in time.
Lin T, Karimireddy S P, Stich S U, et al. Quasi-global momentum: accelerating decentralized deep learning on heterogeneous data. In: Proceedings of International Conference on Machine Learning, Virtual, 2021. 6654-6665
2021
Closest in time.
Yuan K, Chen Y, Huang X, et al. DecentLaM: decentralized momentum SGD for large-batch deep training. In: Proceedings of the IEEE/CVF International Conference on Computer Vision, Montréal, 2021. 3009-3019
2021
Closest in time.
Zhao S-Y, Xie Y-P, Li W-J. On the convergence and improvement of stochastic normalized gradient descent. Sci China Inf Sci, 2021, 64: 132103
2021
Closest in time.
Liu R, Mozafari B. Communication-efficient distributed learning for large batch optimization. In: Proceedings of International Conference on Machine Learning, Baltimore, 2022. 13925-13946
2022
Closest in time.
Liu Y, Mai S, Chen X, et al. Towards efficient and scalable sharpness-aware minimization. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Louisiana, 2022. 12360-12370
2022
Closest in time.