Fetching the paper…
Reading the bibliography…
With the rapid growth of data, distributed momentum stochastic gradient descent~(DMSGD) has been widely used in distributed learning, especially for training large-scale deep models.
On the convergence and improvement of stochastic normalized gradient descent
Shen-Yi Zhao, Yin-Peng Xie, and Wu-Jun Li · 1905
Earlier work this paper cites.
A stochastic approximation method
Herbert Robbins and Sutton Monro · 1951
Earlier work this paper cites.
Some methods of speeding up the convergence of iteration methods
Boris Polyak · 1964
Earlier work this paper cites.
An incremental gradient(-projection) method with momentum term and adaptive stepsize rule
Paul Tseng · 1998
Earlier work this paper cites.
Large-scale machine learning with stochastic gradient descent
Léon Bottou · 2010
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E. Hinton · 2012
Earlier work this paper cites.
An optimal method for stochastic composite optimization
Guanghui Lan · 2012
Earlier work this paper cites.
Accelerating stochastic gradient descent using predictive variance reduction
Rie Johnson and Tong Zhang · 2013
Earlier work this paper cites.
On the importance of initialization and momentum in deep learning
Ilya Sutskever, James Martens, George E. Dahl, and Geoffrey E. Hinton · 2013
Earlier work this paper cites.
Scaling distributed machine learning with the parameter server
Mu Li, David G. Andersen, Jun Woo Park, Alexander J. Smola, Amr Ahmed, Vanja Josifovski, James Long, Eugene J. Shekita, and Bor-Yiing Su · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Sparse communication for distributed gradient descent
Alham Fikri Aji and Kenneth Heafield · 2017
Earlier work this paper cites.
QSGD: Communication-efficient SGD via gradient quantization and encoding
Dan Alistarh, Demjan Grubic, Jerry Li, Ryota Tomioka, and Milan Vojnovic · 2017
Cited alongside, same era.
SGDR: Stochastic gradient descent with warm restarts
Ilya Loshchilov and Frank Hutter · 2017
Cited alongside, same era.
Terngrad: Ternary gradients to reduce communication in distributed deep learning
Wei Wen, Cong Xu, Feng Yan, Chunpeng Wu, Yandan Wang, Yiran Chen, and Hai Li · 2017
Cited alongside, same era.
The convergence of sparsified gradient methods
Dan Alistarh, Torsten Hoefler, Mikael Johansson, Nikola Konstantinov, Sarit Khirirat, and Cédric Renggli · 2018
Cited alongside, same era.
A linear speedup analysis of distributed deep learning with sparse and quantized communication
Peng Jiang and Gagan Agrawal · 2018
Cited alongside, same era.
Deep gradient compression: Reducing the communication bandwidth for distributed training
Error feedback fixes signsgd and other gradient compression schemes
Sai Praneeth Karimireddy, Quentin Rebjock, Sebastian U. Stich, and Martin Jaggi · 2019
Closest in time.
Doublesqueeze: Parallel stochastic gradient descent with double-pass error-compensated compression
Hanlin Tang, Chen Yu, Xiangru Lian, Tong Zhang, and Ji Liu · 2019
Closest in time.
Powersgd: Practical low-rank gradient compression for distributed optimization
Thijs Vogels, Sai Praneeth Karimireddy, and Martin Jaggi · 2019
Closest in time.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Closest in time.
The non-iid data quagmire of decentralized machine learning
Kevin Hsieh, Amar Phanishayee, Onur Mutlu, and Phillip B. Gibbons · 2020
Closest in time.
Decentralized deep learning with arbitrary communication compression
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yujun Lin, Song Han, Huizi Mao, Yu Wang, and Bill Dally · 2018
Cited alongside, same era.
Sparsified SGD with memory
Sebastian U. Stich, Jean-Baptiste Cordonnier, and Martin Jaggi · 2018
Cited alongside, same era.
Group normalization
Yuxin Wu and Kaiming He · 2018
Cited alongside, same era.
Proximal scope for distributed sparse learning
Shen-Yi Zhao, Gong-Duo Zhang, Ming-Wei Li, and Wu-Jun Li · 2018
Cited alongside, same era.
Qsparse-local-sgd: Distributed SGD with quantization, sparsification and local computations
Debraj Basu, Deepesh Data, Can Karakus, and Suhas N. Diggavi · 2019
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Cited alongside, same era.
Measuring the effects of non-identical data distribution for federated visual classification
Tzu-Ming Harry Hsu, Hang Qi, and Matthew Brown · 2019
Cited alongside, same era.
Anastasia Koloskova, Tao Lin, Sebastian U Stich, and Martin Jaggi · 2020
Closest in time.
CSER: Communication-efficient SGD with error reset
Cong Xie, Shuai Zheng, Oluwasanmi Koyejo, Indranil Gupta, Mu Li, and Haibin Lin · 2020
Closest in time.
Stochastic normalized gradient descent with momentum for large batch training
Shen-Yi Zhao, Yin-Peng Xie, and Wu-Jun Li · 2020
Closest in time.
Vision transformer for small-size datasets
Seung Hoon Lee, Seunghyun Lee, and Byung Cheol Song · 2021
Closest in time.
Quasi-global momentum: Accelerating decentralized deep learning on heterogeneous data
Tao Lin, Sai Praneeth Karimireddy, Sebastian U. Stich, and Martin Jaggi · 2021
Closest in time.
Detached error feedback for distributed SGD with random sparsification
An Xu and Heng Huang · 2022
Closest in time.
Llama: Open and efficient foundation language models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurélien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample · 2023
Closest in time.