Fetching the paper…
Reading the bibliography…
The ability to scale out training workloads has been one of the key performance enablers of deep learning.
Transformer-xl: Attentive language models beyond a fixed-length context
Zihang Dai, Zhilin Yang, Yiming Yang, Jaime Carbonell, Quoc V Le, and Ruslan Salakhutdinov. 2019 · 1901
Earlier work this paper cites.
Large batch optimization for deep learning: Training bert in 76 minutes
Yang You, Jing Li, Sashank Reddi, Jonathan Hseu, Sanjiv Kumar, Srinadh Bhojanapalli, Xiaodan Song, James Demmel, Kurt Keutzer, and Cho-Jui Hsieh. 2019 · 1904
Earlier work this paper cites.
Priority-based parameter propagation for distributed DNN training
Anand Jayarajan, Jinliang Wei, Garth Gibson, Alexandra Fedorova, and Gennady Pekhimenko. 2019 · 1905
Earlier work this paper cites.
Some Methods for Classification and Analysis of MultiVariate Observations. In Proc. of the fifth Berkeley Symposium on Mathematical Statistics and Probability (Berkeley, CA, USA, June 21-July 18 1965), Vol. 1. University of California Press, 281–297
J. B. MacQueen. 1967 · 1965
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database. In 2009 IEEE conference on computer vision and pattern recognition (Miami, FL, June 20 - 25, 2009). IEEE, 248–255
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. 2009 · 2009
Earlier work this paper cites.
Project Adam: Building an Efficient and Scalable Deep Learning Training System. In Proceedings of the 11th USENIX Conference on Operating Systems Design and Implementation (Broomfield, CO, USA, October 6-8, 2014), Vol. 14. 571–582
Trishul M Chilimbi, Yutaka Suzue, Johnson Apacible, and Karthik Kalyanaraman. 2014 · 2014
Earlier work this paper cites.
Scaling distributed machine learning with the parameter server. In Proceedings 11th { \{ USENIX } \} Symposium on Operating Systems Design and Implementation ( { \{ OSDI } \} 14) (Broomfield, CO, USA, October 6-8, 2014). 583–598
Mu Li, David G Andersen, Jun Woo Park, Alexander J Smola, Amr Ahmed, Vanja Josifovski, James Long, Eugene J Shekita, and Bor-Yiing Su. 2014 · 2014
Earlier work this paper cites.
1-bit stochastic gradient descent and its application to data-parallel distributed training of speech dnns. In Fifteenth Annual Conference of the International Speech Communication Association (Singapore, September 14-18, 2014)
Frank Seide, Hao Fu, Jasha Droppo, Gang Li, and Dong Yu. 2014 · 2014
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman. 2014 · 2014
Earlier work this paper cites.
Scalable distributed DNN training using commodity GPU cloud computing. In Sixteenth Annual Conference of the International Speech Communication Association (Dresden, Germany, September 6-10, 2015)
Nikko Strom. 2015 · 2015
Earlier work this paper cites.
Tensorflow: A system for large-scale machine learning. In Proceedings of the 12th USENIX Conference on Operating Systems Design and Implementation (Savannah, GA, USA, November 2 - 4, 2016) (OSDI’16) . USENIX Association, USA, 265–283
Martín Abadi, Paul Barham, Jianmin Chen, Zhifeng Chen, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Geoffrey Irving, Michael Isard, et al · 2016
Earlier work this paper cites.
Communication quantization for data-parallel training of deep neural networks. In Proceedings of the Workshop on Machine Learning in High Performance Computing Environments (Salt Lake City, UT, USA, November 14 2016). IEEE Press, 1–8
Nikoli Dryden, Sam Ade Jacobs, Tim Moon, and Brian Van Essen. 2016 · 2016
Earlier work this paper cites.
QSGD: Communication-efficient SGD via gradient quantization and encoding. In Advances in Neural Information Processing Systems (Long Beach, CA, USA, December 4 - 7, 2017), Vol. 30. 1709–1720
Dan Alistarh, Demjan Grubic, Jerry Li, Ryota Tomioka, and Milan Vojnovic. 2017 · 2017
Earlier work this paper cites.
Scaling Deep Learning Workloads: NVIDIA DGX-1/Pascal and Intel Knights Landing
Nitin A. Gawande, Joshua B. Landwehr, Jeff A. Daily, Nathan R. Tallent, Abhinav Vishnu, and Darren J. Kerbyson. 2017 · 2017
Earlier work this paper cites.
Accurate, large minibatch sgd: Training imagenet in 1 hour
Priya Goyal, Piotr Dollár, Ross Girshick, Pieter Noordhuis, Lukasz Wesolowski, Aapo Kyrola, Andrew Tulloch, Yangqing Jia, and Kaiming He. 2017 · 2017
Earlier work this paper cites.
Deep gradient compression: Reducing the communication bandwidth for distributed training
Yujun Lin, Song Han, Huizi Mao, Yu Wang, and William J Dally. 2017 · 2017
Earlier work this paper cites.
Attention is all you need. In Advances in neural information processing systems . 5998–6008
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
Terngrad: Ternary gradients to reduce communication in distributed deep learning
Wei Wen, Cong Xu, Feng Yan, Chunpeng Wu, Yandan Wang, Yiran Chen, and Hai Li. 2017 · 2017
Earlier work this paper cites.
ADaComP: Adaptive residual gradient compression for data-parallel distributed training. In 32nd AAAI Conference on Artificial Intelligence, AAAI 2018 (New Orleans, USA, February 2–7, 2018). 2827–2835
Chia Yu Chen, Jungwook Choi, Daniel Brand, Ankur Agrawal, Wei Zhang, and Kailash Gopalakrishnan. 2018 · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018 · 2018
Cited alongside, same era.
Synchronous multi-gpu deep learning with low-precision communication: An experimental study. In Proceedings of the 21st International Conference on Extending Database Technology (Vienna, Austria, March 26-29, 2018). OpenProceedings, 145–156
Demjan Grubic, Leo K Tam, Dan Alistarh, and Ce Zhang. 2018 · 2018
Cited alongside, same era.
Tartan: Evaluating Modern GPU Interconnect via a Multi-GPU Benchmark Suite. In 2018 IEEE International Symposium on Workload Characterization (IISWC) (Raleigh, NC, USA, 30 September - 02 October 2018). 191–202
Ang Li, Shuaiwen Leon Song, Jieyang Chen, Xu Liu, Nathan Tallent, and Kevin Barker. 2018 · 2018
Cited alongside, same era.
3lc: Lightweight and effective traffic compression for distributed machine learning
Hyeontaek Lim, David G Andersen, and Michael Kaminsky. 2018 · 2018
A Unified Architecture for Accelerating Distributed DNN Training in Heterogeneous GPU/CPU Clusters. In 14th USENIX Symposium on Operating Systems Design and Implementation (OSDI 20) (Virtual Event, November 4–6, 2020). USENIX Association, 463–479
Yimin Jiang, Yibo Zhu, Chang Lan, Bairen Yi, Yong Cui, and Chuanxiong Guo. 2020 · 2020
Later among the works it cites.
MLPerf: An industry standard benchmark suite for machine learning performance
Peter Mattson, Vijay Janapa Reddi, Christine Cheng, Cody Coleman, Greg Diamos, David Kanter, Paulius Micikevicius, David Patterson, Guenther Schmuelling, Hanlin Tang, et al · 2020
Later among the works it cites.
GRACE: A Compressed Communication Framework for Distributed Machine Learning. In Proceedings of ICDCS’21 (Virtual event, July 7- 10, 2020)
Hang Xu, Chen-Yu Ho, Ahmed M. Abdelmoniem, Aritra Dutta, El Houcine Bergou, Konstantinos Karatsenidis, Marco Canini, and Panos Kalnis. 2021 · 2020
Later among the works it cites.
Fast Training of Deep Learning Models over Multiple GPUs. In Proceedings of the 21st International Middleware Conference (Delft, Netherlands, December 7 - 11, 2020). 105–118
Xiaodong Yi, Ziyue Luo, Chen Meng, Mengdi Wang, Guoping Long, Chuan Wu, Jun Yang, and Wei Lin. 2020 · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Image transformer. In International Conference on Machine Learning (Stockholm, Sweden, July 10-15, 2018). PMLR, 4055–4064
Niki Parmar, Ashish Vaswani, Jakob Uszkoreit, Lukasz Kaiser, Noam Shazeer, Alexander Ku, and Dustin Tran. 2018 · 2018
Cited alongside, same era.
Horovod: fast and easy distributed deep learning in TensorFlow
Alexander Sergeev and Mike Del Balso. 2018 · 2018
Cited alongside, same era.
ATOMO: Communication-efficient learning via atomic sparsification
Hongyi Wang, Scott Sievert, Zachary Charles, Shengchao Liu, Stephen Wright, and Dimitris Papailiopoulos. 2018 · 2018
Cited alongside, same era.
Error feedback fixes signsgd and other gradient compression schemes. In International Conference on Machine Learning (Long Beach, CA, USA, Jun 10 – 15, 2019). PMLR, 3252–3261
Sai Praneeth Karimireddy, Quentin Rebjock, Sebastian Stich, and Martin Jaggi. 2019 · 2019
Cited alongside, same era.
Evaluating Modern GPU Interconnect: PCIe, NVLink, NV-SLI, NVSwitch and GPUDirect
Ang Li, Shuaiwen Leon Song, Jieyang Chen, Jiajia Li, Xu Liu, Nathan R. Tallent, and Kevin J. Barker. 2020 · 2019
Cited alongside, same era.
Pytorch: An imperative style, high-performance deep learning library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al · 2019
Cited alongside, same era.
A generic communication scheduler for distributed dnn training acceleration. In Proceedings of the 27th ACM Symposium on Operating Systems Principles (Ontario, Canada, October 27 - 30, 2019). 16–29
Yanghua Peng, Yibo Zhu, Yangrui Chen, Yixin Bao, Bairen Yi, Chang Lan, Chuan Wu, and Chuanxiong Guo. 2019 · 2019
Cited alongside, same era.
SparCML: High-performance sparse communication for machine learning. In Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis (Denver, CO, USA, November 17–22, 2019)
Cèdric Renggli, Saleh Ashkboos, Mehdi Aghagolzadeh, Dan Alistarh, and Torsten Hoefler. 2019 · 2019
Cited alongside, same era.
Later among the works it cites.
Petrel: Heterogeneity-aware distributed deep learning via hybrid synchronization
Qihua Zhou, Song Guo, Zhihao Qu, Peng Li, Li Li, Minyi Guo, and Kun Wang. 2020 · 2020
Later among the works it cites.
Adaptive Gradient Communication via Critical Learning Regime Identification. In Proceedings of Machine Learning and Systems (Virtual event, USA, April 5 - 9, 2021), Vol. 3. 55–80
Saurabh Agarwal, Hongyi Wang, Kangwook Lee, Shivaram Venkataraman, and Dimitris Papailiopoulos. 2021 · 2021
Closest in time.
Gradient Compression Supercharged High-Performance Data Parallel DNN Training. In Proceedings of the ACM SIGOPS 28th Symposium on Operating Systems Principles (Virtual Event, Germany, October 26-29, 2021) (SOSP ’21) . Association for Computing Machinery, New York, NY, USA, 359–375
Youhui Bai, Cheng Li, Quan Zhou, Jun Yi, Ping Gong, Feng Yan, Ruichuan Chen, and Yinlong Xu. 2021 · 2021
Closest in time.
Efficient Sparse Collective Communication and Its Application to Accelerate Distributed Deep Learning. In Proceedings of the 35th ACM SIGCOMM 2021 Conference (Virtual Event,USA, August 23 - 27, 2021) (SIGCOMM ’21) . 676–691
Jiawei Fei, Chen-Yu Ho, Atal N. Sahu, Marco Canini, and Amedeo Sapio. 2021 · 2021
Closest in time.
Bagua: Scaling up Distributed Learning with System Relaxations
Shaoduo Gan, Jiawei Jiang, Binhang Yuan, Ce Zhang, Xiangru Lian, Rui Wang, Jianbin Chang, Chengjun Liu, Hongmei Shi, Shengzhuo Zhang, Xianghong Li, Tengxu Sun, Sen Yang, and Ji Liu. 2021 · 2021
Closest in time.
Sync-Switch: Hybrid Parameter Synchronization for Distributed Deep Learning
Shijian Li, Oren Mangoubi, Lijie Xu, and Tian Guo. 2021 · 2021
Closest in time.
An efficient statistical-based gradient compression technique for distributed training systems. In Proceedings of Machine Learning and Systems (Virtual event, USA, April 5 - 9, 2021), Vol. 3. 297–322
Ahmed M Abdelmoniem, Ahmed Elzanaty, Mohamed-Slim Alouini, and Marco Canini. 2021 · 2021
Closest in time.
NUQSGD: Provably Communication-efficient Data-parallel SGD via Nonuniform Quantization
Ali Ramezani-Kebrya, Fartash Faghri, Ilya Markov, Vitalii Aksenov, Dan Alistarh, and Daniel M Roy. 2021 · 2021
Closest in time.
NVIDIA AMPERE GA102 GPU ARCHITECTURE
2021 · 2022
Closest in time.
Genesis GPU Cloud Offering
Genesis. 2021 · 2022
Closest in time.
Dual NVIDIA GeForce RTX 3090 NVLink Performance Review
William Harmon. 2021 · 2022
Closest in time.
Huggingface Transformers Repository
Inc Huggingface. 2022 · 2022
Closest in time.
LambdaLabs GPU Cloud Offering
LambdaLabs. 2021 · 2022
Closest in time.
LeaderGPU Cloud Offering
LeaderGPU. 2021 · 2022
Closest in time.
NVIDIA Deep Learning Examples for Tensor Cores
Nvidia. 2020 · 2022
Closest in time.