Fetching the paper…
Reading the bibliography…
To train large models (like BERT and GPT-3) on hundreds of GPUs, communication has become a major bottleneck, especially on commodity systems with limited-bandwidth TCP network.
1-bit stochastic gradient descent and application to data-parallel distributed training of speech dnns
Frank Seide, Hao Fu, Jasha Droppo, Gang Li, and Dong Yu · 2014
Earlier work this paper cites.
Asynchronous stochastic convex optimization: the noise is in the noise and sgd don t care
Sorathan Chaturapruek, John C Duchi, and Christopher Ré · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Asynchronous stochastic gradient descent with delay compensation for distributed deep learning
Shuxin Zheng, Qi Meng, Taifeng Wang, Wei Chen, Nenghai Yu, Zhiming Ma, and Tie-Yan Liu · 2016
Earlier work this paper cites.
QSGD: Communication-Efficient SGD via gradient quantization and encoding
Dan Alistarh, Demjan Grubic, Jerry Li, Ryota Tomioka, and Milan Vojnovic · 2017
Earlier work this paper cites.
Can decentralized algorithms outperform centralized algorithms? a case study for decentralized parallel stochastic gradient descent
Xiangru Lian, Ce Zhang, Huan Zhang, Cho-Jui Hsieh, Wei Zhang, and Ji Liu · 2017
Earlier work this paper cites.
Terngrad: Ternary gradients to reduce communication in distributed deep learning
Wei Wen, Cong Xu, Feng Yan, Chunpeng Wu, Yandan Wang, Yiran Chen, and Hai Li · 2017
Earlier work this paper cites.
ZipML: Training linear models with end-to-end low precision, and a little bit of deep learning
Hantian Zhang, Jerry Li, Kaan Kara, Dan Alistarh, Ji Liu, and Ce Zhang · 2017
Earlier work this paper cites.
cpSGD: Communication-efficient and differentially-private distributed SGD
Naman Agarwal, Ananda Theertha Suresh, Felix Xinnan X Yu, Sanjiv Kumar, and Brendan McMahan · 2018
Earlier work this paper cites.
A linear speedup analysis of distributed deep learning with sparse and quantized communication
Peng Jiang and Gagan Agrawal · 2018
Earlier work this paper cites.
Pipe-sgd: A decentralized pipelined sgd framework for distributed deep net training
Youjie Li, Mingchao Yu, Songze Li, Salman Avestimehr, Nam Sung Kim, and Alexander Schwing · 2018
Earlier work this paper cites.
Towards more efficient stochastic decentralized learning: Faster convergence and sparse communication
Zebang Shen, Aryan Mokhtari, Tengfei Zhou, Peilin Zhao, and Hui Qian · 2018
Earlier work this paper cites.
Sparsified sgd with memory
Sebastian U Stich, Jean-Baptiste Cordonnier, and Martin Jaggi · 2018
Cited alongside, same era.
GLUE: A multi-task benchmark and analysis platform for natural language understanding
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel Bowman · 2018
Cited alongside, same era.
Gradient sparsification for Communication-Efficient distributed optimization
Jianqiao Wangni, Jialei Wang, Ji Liu, and Tong Zhang · 2018
Cited alongside, same era.
Communication-Computation efficient gradient coding
Min Ye and Emmanuel Abbe · 2018
Cited alongside, same era.
Qsparse-local-sgd: Distributed sgd with quantization, sparsification and local computations
Debraj Basu, Deepesh Data, Can Karakus, and Suhas Diggavi · 2019
Cited alongside, same era.
signSGD with majority vote is communication efficient and fault tolerant
Jeremy Bernstein, Jiawei Zhao, Kamyar Azizzadenesheli, and Anima Anandkumar · 2019
Powersgd: Practical low-rank gradient compression for distributed optimization
Thijs Vogels, Sai Praneeth Karimireddy, and Martin Jaggi · 2019
Later among the works it cites.
Double quantization for communication-efficient distributed optimization
Yue Yu, Jiaxiang Wu, and Longbo Huang · 2019
Later among the works it cites.
Communication-efficient distributed blockwise momentum sgd with error-feedback
Shuai Zheng, Ziyue Huang, and James Kwok · 2019
Later among the works it cites.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei · 2020
Later among the works it cites.
Decentralized deep learning with arbitrary communication compression
Anastasia Koloskova*, Tao Lin*, Sebastian U Stich, and Martin Jaggi · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Cited alongside, same era.
Communication-efficient distributed sgd with sketching
Nikita Ivkin, Daniel Rothchild, Enayat Ullah, Vladimir braverman, Ion Stoica, and Raman Arora · 2019
Cited alongside, same era.
A distributed synchronous sgd algorithm with global top-k sparsification for low bandwidth networks
S. Shi, Q. Wang, K. Zhao, Z. Tang, Y. Wang, X. Huang, and X. Chu · 2019
Cited alongside, same era.
Compressing gradient optimizers via Count-Sketches
Ryan Spring, Anastasios Kyrillidis, Vijai Mohan, and Anshumali Shrivastava · 2019
Cited alongside, same era.
Communication-efficient distributed learning via lazily aggregated quantized gradients
Jun Sun, Tianyi Chen, Georgios Giannakis, and Zaiyue Yang · 2019
Cited alongside, same era.
Distributed sgd with flexible gradient compression
T. T. Phuong and L. T. Phong · 2020
Later among the works it cites.
On layer normalization in the transformer architecture
Ruibin Xiong, Yunchang Yang, Di He, Kai Zheng, Shuxin Zheng, Chen Xing, Huishuai Zhang, Yanyan Lan, Liwei Wang, and Tieyan Liu · 2020
Later among the works it cites.
Large batch optimization for deep learning: Training bert in 76 minutes
Yang You, Jing Li, Sashank Reddi, Jonathan Hseu, Sanjiv Kumar, Srinadh Bhojanapalli, Xiaodan Song, James Demmel, Kurt Keutzer, and Cho-Jui Hsieh · 2020
Later among the works it cites.
Accelerating training of transformer-based language models with progressive layer dropping
Minjia Zhang and Yuxiong He · 2020
Later among the works it cites.
Towards scalable distributed training of deep learning on public cloud clusters
Shaohuai Shi, Xianhao Zhou, Shutao Song, Xingyao Wang, Zilin Zhu, Xue Huang, Xinan Jiang, Feihu Zhou, Zhenyu Guo, Liqiang Xie, Rui Lan, Xianbin Ouyang, Yan Zhang, Jieqian Wei, Jing Gong, Weiliang Lin, Ping Gao, Peng Meng, Xiaomin Xu, Chenyang Guo, Bo Yang, Zhibo Chen, Yongjian Wu, and Xiaowen Chu · 2021
Closest in time.
1-bit Adam: Communication Efficient Large-Scale Training with Adam’s Convergence Speed
Hanlin Tang, Shaoduo Gan, Ammar Ahmad Awan, Samyam Rajbhandari, Conglong Li, Xiangru Lian, Ji Liu, Ce Zhang, and Yuxiong He · 2021
Closest in time.