Fetching the paper…
Reading the bibliography…
Training large deep learning models at scale is very challenging.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber. 1997 · 1997
Earlier work this paper cites.
InfiniBand network architecture
Tom Shanley. 2003 · 2003
Earlier work this paper cites.
Communication lower bounds for distributed-memory matrix multiplication
Dror Irony, Sivan Toledo, and Alexander Tiskin. 2004 · 2004
Earlier work this paper cites.
Optimization of collective reduction operations. In International Conference on Computational Science . Springer, 1–9
Rolf Rabenseifner. 2004 · 2004
Earlier work this paper cites.
Optimization of collective communication operations in MPICH
Rajeev Thakur, Rolf Rabenseifner, and William Gropp. 2005 · 2005
Earlier work this paper cites.
Implementation and Performance Analysis of Non-Blocking Collective Operations for MPI. In Proceedings of the 2007 International Conference on High Performance Computing, Networking, Storage and Analysis, SC07 (Reno, USA). IEEE Computer Society/ACM
Torsten Hoefler, Andrew Lumsdaine, and Wolfgang Rehm. 2007b · 2007
Earlier work this paper cites.
Message Progression in Parallel Computing - To Thread or not to Thread?. In Proceedings of the 2008 IEEE International Conference on Cluster Computing (Tsukuba, Japan). IEEE Computer Society
Torsten Hoefler and Andrew Lumsdaine. 2008 · 2008
Earlier work this paper cites.
Cray XC series network
Bob Alverson, Edwin Froese, Larry Kaplan, and Duncan Roweth. 2012 · 2012
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton. 2012 · 2012
Earlier work this paper cites.
NUMA-aware shared-memory collective communication for MPI. In Proceedings of the 22nd international symposium on High-performance parallel and distributed computing . 85–96
Shigang Li, Torsten Hoefler, and Marc Snir. 2013 · 2013
Earlier work this paper cites.
One weird trick for parallelizing convolutional neural networks
Alex Krizhevsky. 2014 · 2014
Earlier work this paper cites.
Training deep nets with sublinear memory cost
Tianqi Chen, Bing Xu, Chiyuan Zhang, and Carlos Guestrin. 2016 · 2016
Earlier work this paper cites.
QSGD: Communication-efficient SGD via gradient quantization and encoding. In Advances in Neural Information Processing Systems (NeurIPS)
Dan Alistarh, Demjan Grubic, Jerry Li, Ryota Tomioka, and Milan Vojnovic. 2017 · 2017
Earlier work this paper cites.
Ultra-performance Pascal GPU and NVLink Interconnect
Denis Foley and John Danskin. 2017 · 2017
Earlier work this paper cites.
AMPNet: Asynchronous model-parallel training for dynamic neural networks
Alexander L Gaunt, Matthew A Johnson, Maik Riechert, Daniel Tarlow, Ryota Tomioka, Dimitrios Vytiniotis, and Sam Webster. 2017 · 2017
Earlier work this paper cites.
Accurate, large minibatch sgd: Training imagenet in 1 hour
Priya Goyal, Piotr Dollár, Ross Girshick, Pieter Noordhuis, Lukasz Wesolowski, Aapo Kyrola, Andrew Tulloch, Yangqing Jia, and Kaiming He. 2017 · 2017
Earlier work this paper cites.
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
AI and Compute
Dario Amodei and Danny Hernandez. 2018 · 2018
Earlier work this paper cites.
Optimization methods for large-scale machine learning
Léon Bottou, Frank E Curtis, and Jorge Nocedal. 2018 · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018 · 2018
Cited alongside, same era.
Asynchronous Decentralized Parallel Stochastic Gradient Descent. In Proceedings of the 35th International Conference on Machine Learning (Proceedings of Machine Learning Research, Vol. 80) , Jennifer Dy and Andreas Krause (Eds.). PMLR, 3043–3052
Xiangru Lian, Wei Zhang, Ce Zhang, and Ji Liu. 2018 · 2018
Cited alongside, same era.
Horovod: fast and easy distributed deep learning in TensorFlow
Alexander Sergeev and Mike Del Balso. 2018 · 2018
Cited alongside, same era.
Yang You, Zhao Zhang, Cho-Jui Hsieh, James Demmel, and Kurt Keutzer. 2018 · 2018
Cited alongside, same era.
Large batch optimization for deep learning: Training bert in 76 minutes
Yang You, Jing Li, Sashank Reddi, Jonathan Hseu, Sanjiv Kumar, Srinadh Bhojanapalli, Xiaodan Song, James Demmel, Kurt Keutzer, and Cho-Jui Hsieh. 2019b · 2019
Later among the works it cites.
Language models are few-shot learners
Tom B Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Later among the works it cites.
End-to-end object detection with transformers. In European Conference on Computer Vision . Springer, 213–229
Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko. 2020 · 2020
Later among the works it cites.
Pre-trained image processing transformer
Hanting Chen, Yunhe Wang, Tianyu Guo, Chang Xu, Yiping Deng, Zhenhua Liu, Siwei Ma, Chunjing Xu, Chao Xu, and Wen Gao. 2020b · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Stochastic gradient push for distributed deep learning. In Proceedings of the Thirty-sixth International Conference on Machine Learning (ICML) . 344–353
Mahmoud Assran, Nicolas Loizou, Nicolas Ballas, and Michael Rabbat. 2019 · 2019
Cited alongside, same era.
Demystifying parallel and distributed deep learning: An in-depth concurrency analysis
Tal Ben-Nun and Torsten Hoefler. 2019 · 2019
Cited alongside, same era.
Gpipe: Efficient training of giant neural networks using pipeline parallelism. In Advances in neural information processing systems . 103–112
Yanping Huang, Youlong Cheng, Ankur Bapna, Orhan Firat, Dehao Chen, Mia Chen, HyoukJoong Lee, Jiquan Ngiam, Quoc V Le, Yonghui Wu, et al · 2019
Cited alongside, same era.
Beyond data and model parallelism for deep neural networks
Zhihao Jia, Matei Zaharia, and Alex Aiken. 2019 · 2019
Cited alongside, same era.
Red-Blue Pebbling Revisited: Near Optimal Parallel Matrix-Matrix Multiplication. In Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis (SC19)
Grzegorz Kwasniewski, Marko Kabić, Maciej Besta, Joost VandeVondele, Raffaele Solcà, and Torsten Hoefler. 2019 · 2019
Cited alongside, same era.
SwarmSGD: Scalable decentralized SGD with local updates
Giorgi Nadiradze, Amirmojtaba Sabour, Dan Alistarh, Aditya Sharma, Ilia Markov, and Vitaly Aksenov. 2019 · 2019
Cited alongside, same era.
PipeDream: generalized pipeline parallelism for DNN training. In Proceedings of the 27th ACM Symposium on Operating Systems Principles . 1–15
Deepak Narayanan, Aaron Harlap, Amar Phanishayee, Vivek Seshadri, Nikhil R Devanur, Gregory R Ganger, Phillip B Gibbons, and Matei Zaharia. 2019 · 2019
Cited alongside, same era.
Pytorch: An imperative style, high-performance deep learning library. In Advances in neural information processing systems . 8026–8037
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al · 2019
Cited alongside, same era.
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al · 2020
Later among the works it cites.
GEMS: GPU-Enabled Memory-Aware Model-Parallelism System for Distributed DNN Training. In 2020 SC20: International Conference for High Performance Computing, Networking, Storage and Analysis (SC) . IEEE Computer Society, 621–635
Arpan Jain, Ammar Awan, Asmaa Aljuhani, Jahanzeb Hashmi, Quentin Anthony, Hari Subramoni, Dhabaleswar Panda, Raghu Machiraju, and Anil Parwani. 2020 · 2020
Later among the works it cites.
Breaking (Global) Barriers in Parallel Stochastic Optimization with Wait-Avoiding Group Averaging
Shigang Li, Tal Ben-Nun, Giorgi Nadiradze, Salvatore Digirolamo, Nikoli Dryden, Dan Alistarh, and Torsten Hoefler. 2020b · 2020
Later among the works it cites.
Memory-efficient pipeline-parallel DNN training
Deepak Narayanan, Amar Phanishayee, Kaiyu Shi, Xie Chen, and Matei Zaharia. 2020 · 2020
Later among the works it cites.
NVIDIA Collective Communications Library
NVIDIA. 2020 · 2020
Later among the works it cites.
Zero: Memory optimizations toward training trillion parameter models. In SC20: International Conference for High Performance Computing, Networking, Storage and Analysis . IEEE, 1–16
Samyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, and Yuxiong He. 2020 · 2020
Later among the works it cites.
An In-Depth Analysis of the Slingshot Interconnect. In Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis (SC20)
Daniele De Sensi, Salvatore Di Girolamo, Kim H. McMahon, Duncan Roweth, and Torsten Hoefler. 2020 · 2020
Later among the works it cites.
Communication-efficient distributed deep learning: A comprehensive survey
Zhenheng Tang, Shaohuai Shi, Xiaowen Chu, Wei Wang, and Bo Li. 2020 · 2020
Later among the works it cites.
FPDeep: Scalable acceleration of CNN training on deeply-pipelined FPGA clusters
Tianqi Wang, Tong Geng, Ang Li, Xi Jin, and Martin Herbordt. 2020 · 2020
Later among the works it cites.
Why Dataset Properties Bound the Scalability of Parallel Machine Learning Training Algorithms
Daning Cheng, Shigang Li, Hanping Zhang, Fen Xia, and Yunquan Zhang. 2021 · 2021
Closest in time.
DAPPLE: a pipelined data parallel approach for training large models. In Proceedings of the 26th ACM SIGPLAN Symposium on Principles and Practice of Parallel Programming . 431–445
Shiqing Fan, Yi Rong, Chen Meng, Zongyan Cao, Siyu Wang, Zhen Zheng, Chuan Wu, Guoping Long, Jun Yang, Lixue Xia, et al · 2021
Closest in time.
Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity
William Fedus, Barret Zoph, and Noam Shazeer. 2021 · 2021
Closest in time.
Torsten Hoefler, Dan Alistarh, Tal Ben-Nun, Nikoli Dryden, and Alexandra Peste. 2021 · 2021
Closest in time.
ZeRO-Offload: Democratizing Billion-Scale Model Training
Jie Ren, Samyam Rajbhandari, Reza Yazdani Aminabadi, Olatunji Ruwase, Shuangyan Yang, Minjia Zhang, Dong Li, and Yuxiong He. 2021 · 2021
Closest in time.