Fetching the paper…
Reading the bibliography…
Communication overhead is one of the major obstacles to train large deep learning models at scale.
Two algorithms for barrier synchronization
Debra Hensgen, Raphael Finkel, and Udi Manber. 1988 · 1988
Earlier work this paper cites.
Environmental robustness in automatic speech recognition. In International Conference on Acoustics, Speech, and Signal Processing . IEEE, 849–852
Alejandro Acero and Richard M Stern. 1990 · 1990
Earlier work this paper cites.
Analysis of quickselect: An algorithm for order statistics
Hosam M Mahmoud, Reza Modarres, and Robert T Smythe. 1995 · 1995
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber. 1997 · 1997
Earlier work this paper cites.
InfiniBand network architecture
Tom Shanley. 2003 · 2003
Earlier work this paper cites.
Optimization of collective communication operations in MPICH
Rajeev Thakur, Rolf Rabenseifner, and William Gropp. 2005 · 2005
Earlier work this paper cites.
Collective communication: theory, practice, and experience
Ernie Chan, Marcel Heimlich, Avi Purkayastha, and Robert Van De Geijn. 2007 · 2007
Earlier work this paper cites.
Cray XC series network
Bob Alverson, Edwin Froese, Larry Kaplan, and Duncan Roweth. 2012 · 2012
Earlier work this paper cites.
NUMA-aware shared-memory collective communication for MPI. In Proceedings of the 22nd international symposium on High-performance parallel and distributed computing . 85–96
Shigang Li, Torsten Hoefler, and Marc Snir. 2013 · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba. 2014 · 2014
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman. 2014 · 2014
Earlier work this paper cites.
Communication quantization for data-parallel training of deep neural networks. In 2016 2nd Workshop on Machine Learning in HPC Environments (MLHPC) . IEEE, 1–8
Nikoli Dryden, Tim Moon, Sam Ade Jacobs, and Brian Van Essen. 2016 · 2016
Earlier work this paper cites.
Sparse communication for distributed gradient descent
Alham Fikri Aji and Kenneth Heafield. 2017 · 2017
Earlier work this paper cites.
QSGD: Communication-efficient SGD via gradient quantization and encoding
Dan Alistarh, Demjan Grubic, Jerry Li, Ryota Tomioka, and Milan Vojnovic. 2017 · 2017
Earlier work this paper cites.
Ultra-performance Pascal GPU and NVLink Interconnect
Denis Foley and John Danskin. 2017 · 2017
Earlier work this paper cites.
Accurate, large minibatch sgd: Training imagenet in 1 hour
Priya Goyal, Piotr Dollár, Ross Girshick, Pieter Noordhuis, Lukasz Wesolowski, Aapo Kyrola, Andrew Tulloch, Yangqing Jia, and Kaiming He. 2017 · 2017
Earlier work this paper cites.
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
TernGrad: Ternary Gradients to Reduce Communication in Distributed Deep Learning. In Proceedings of the 31st International Conference on Neural Information Processing Systems (Long Beach, California, USA) (NIPS’17) . Curran Associates Inc., Red Hook, NY, USA, 1508–1518
Wei Wen, Cong Xu, Feng Yan, Chunpeng Wu, Yandan Wang, Yiran Chen, and Hai Li. 2017 · 2017
Cited alongside, same era.
The convergence of sparsified gradient methods
Dan Alistarh, Torsten Hoefler, Mikael Johansson, Sarit Khirirat, Nikola Konstantinov, and Cédric Renggli. 2018 · 2018
Cited alongside, same era.
Building efficient convnets using redundant feature pruning
Babajide O Ayinde and Jacek M Zurada. 2018 · 2018
Cited alongside, same era.
signSGD: Compressed optimisation for non-convex problems. In International Conference on Machine Learning . PMLR, 560–569
Jeremy Bernstein, Yu-Xiang Wang, Kamyar Azizzadenesheli, and Animashree Anandkumar. 2018 · 2018
Cited alongside, same era.
Understanding top-k sparsification in distributed deep learning
Shaohuai Shi, Xiaowen Chu, Ka Chun Cheung, and Simon See. 2019a · 2019
Later among the works it cites.
A distributed synchronous SGD algorithm with global top-k sparsification for low bandwidth networks. In 2019 IEEE 39th International Conference on Distributed Computing Systems (ICDCS) . IEEE, 2238–2247
Shaohuai Shi, Qiang Wang, Kaiyong Zhao, Zhenheng Tang, Yuxin Wang, Xiang Huang, and Xiaowen Chu. 2019b · 2019
Later among the works it cites.
Network congestion avoidance through packet-chaining reservation. In Proceedings of the 48th International Conference on Parallel Processing . 1–10
Ke Wu, Dezun Dong, Cunlu Li, Shan Huang, and Yi Dai. 2019 · 2019
Later among the works it cites.
Large batch optimization for deep learning: Training bert in 76 minutes
Yang You, Jing Li, Sashank Reddi, Jonathan Hseu, Sanjiv Kumar, Srinadh Bhojanapalli, Xiaodan Song, James Demmel, Kurt Keutzer, and Cho-Jui Hsieh. 2019b · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Optimization methods for large-scale machine learning
Léon Bottou, Frank E Curtis, and Jorge Nocedal. 2018 · 2018
Cited alongside, same era.
Long live time: improving lifetime for training-in-memory engines by structured gradient sparsification. In Proceedings of the 55th Annual Design Automation Conference . 1–6
Yi Cai, Yujun Lin, Lixue Xia, Xiaoming Chen, Song Han, Yu Wang, and Huazhong Yang. 2018 · 2018
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018 · 2018
Cited alongside, same era.
Horovod: fast and easy distributed deep learning in TensorFlow
Alexander Sergeev and Mike Del Balso. 2018 · 2018
Cited alongside, same era.
Efficient top-k query processing on massively parallel hardware. In Proceedings of the 2018 International Conference on Management of Data . 1557–1570
Anil Shanbhag, Holger Pirk, and Samuel Madden. 2018 · 2018
Cited alongside, same era.
Yang You, Zhao Zhang, Cho-Jui Hsieh, James Demmel, and Kurt Keutzer. 2018 · 2018
Cited alongside, same era.
Demystifying parallel and distributed deep learning: An in-depth concurrency analysis
Tal Ben-Nun and Torsten Hoefler. 2019 · 2019
Cited alongside, same era.
Stochastic distributed learning with gradient quantization and variance reduction
Samuel Horváth, Dmitry Kovalev, Konstantin Mishchenko, Sebastian Stich, and Peter Richtárik. 2019 · 2019
Cited alongside, same era.
Language models are few-shot learners
Tom B Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Later among the works it cites.
Adaptive gradient sparsification for efficient federated learning: An online learning approach. In 2020 IEEE 40th International Conference on Distributed Computing Systems (ICDCS) . IEEE, 300–310
Pengchao Han, Shiqiang Wang, and Kin K Leung. 2020 · 2020
Later among the works it cites.
Breaking (global) barriers in parallel stochastic optimization with wait-avoiding group averaging
Shigang Li, Tal Ben-Nun, Giorgi Nadiradze, Salvatore Di Girolamo, Nikoli Dryden, Dan Alistarh, and Torsten Hoefler. 2020b · 2020
Later among the works it cites.
An In-Depth Analysis of the Slingshot Interconnect. In Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis (SC20)
Daniele De Sensi, Salvatore Di Girolamo, Kim H. McMahon, Duncan Roweth, and Torsten Hoefler. 2020 · 2020
Later among the works it cites.
FFT-based Gradient Sparsification for the Distributed Training of Deep Neural Networks. In Proceedings of the 29th International Symposium on High-Performance Parallel and Distributed Computing . 113–124
Linnan Wang, Wei Wu, Junyu Zhang, Hang Liu, George Bosilca, Maurice Herlihy, and Rodrigo Fonseca. 2020 · 2020
Later among the works it cites.
DAPPLE: a pipelined data parallel approach for training large models. In Proceedings of the 26th ACM SIGPLAN Symposium on Principles and Practice of Parallel Programming . 431–445
Shiqing Fan, Yi Rong, Chen Meng, Zongyan Cao, Siyu Wang, Zhen Zheng, Chuan Wu, Guoping Long, Jun Yang, Lixue Xia, et al · 2021
Later among the works it cites.
Efficient sparse collective communication and its application to accelerate distributed deep learning. In Proceedings of the 2021 ACM SIGCOMM 2021 Conference . 676–691
Jiawei Fei, Chen-Yu Ho, Atal N Sahu, Marco Canini, and Amedeo Sapio. 2021 · 2021
Later among the works it cites.
Torsten Hoefler, Dan Alistarh, Tal Ben-Nun, Nikoli Dryden, and Alexandra Peste. 2021 · 2021
Later among the works it cites.
Chimera: Efficiently Training Large-Scale Neural Networks with Bidirectional Pipelines
Shigang Li and Torsten Hoefler. 2021 · 2021
Later among the works it cites.
Asynchronous decentralized SGD with quantized and local updates
Giorgi Nadiradze, Amirmojtaba Sabour, Peter Davies, Shigang Li, and Dan Alistarh. 2021 · 2021
Later among the works it cites.
Memory-efficient pipeline-parallel dnn training. In International Conference on Machine Learning . PMLR, 7937–7947
Deepak Narayanan, Amar Phanishayee, Kaiyu Shi, Xie Chen, and Matei Zaharia. 2021 · 2021
Later among the works it cites.
DeepReduce: A Sparse-tensor Communication Framework for Federated Deep Learning
Hang Xu, Kelly Kostopoulou, Aritra Dutta, Xin Li, Alexandros Ntoulas, and Panos Kalnis. 2021 · 2021
Later among the works it cites.