Fetching the paper…
Reading the bibliography…
Deep neural networks (DNNs) have grown exponentially in size over the past decade, leaving only those who have massive datacenter-based resources with the ability to develop and train such models.
ImageNet: A large-scale hierarchical image database. In 2009 IEEE Conference on Computer Vision and Pattern Recognition (CVPR’09) . 248–255
Deng, Jia and Dong, Wei and Socher, Richard and Li, Li-Jia and Kai Li and Li Fei-Fei. 2009 · 2009
Earlier work this paper cites.
Imagenet Classification with Deep Convolutional Neural Networks. In Proceedings of the 25th International Conference on Neural Information Processing Systems (NeurIPS’12) . Lake Tahoe, NV, 1097–1105
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton. 2012 · 2012
Earlier work this paper cites.
Playing Atari with Deep Reinforcement Learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, and Martin A. Riedmiller. 2013 · 2013
Earlier work this paper cites.
One weird trick for parallelizing convolutional neural networks
Alex Krizhevsky. 2014 · 2014
Earlier work this paper cites.
Scaling distributed machine learning with the parameter server. In Proceedings of the 11th USENIX Symposium on Operating Systems Design and Implementation (OSDI’14) . Broomfield, CO, 583–598
Mu Li, David G Andersen, Jun Woo Park, Alexander J Smola, Amr Ahmed, Vanja Josifovski, James Long, Eugene J Shekita, and Bor-Yiing Su. 2014 · 2014
Earlier work this paper cites.
Deep learning for computational biology
Christof Angermueller, Tanel Pärnamaa, Leopold Parts, and Oliver Stegle. 2016 · 2016
Earlier work this paper cites.
Training deep nets with sublinear memory cost
Tianqi Chen, Bing Xu, Chiyuan Zhang, and Carlos Guestrin. 2016 · 2016
Earlier work this paper cites.
Deep residual learning for image recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR’16) . Las Vegas, NV, 770–778
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016 · 2016
Earlier work this paper cites.
Identity Mappings in Deep Residual Networks
Kaiming He and Xiangyu Zhang and Shaoqing Ren and Jian Sun. 2016 · 2016
Earlier work this paper cites.
Pointer sentinel mixture models
Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher. 2016 · 2016
Earlier work this paper cites.
NVLINK AND NVSWITCH, https://www.nvidia.com/en-us/data-center/nvlink/
NVIDIA. 2016 · 2016
Earlier work this paper cites.
vDNN: Virtualized deep neural networks for scalable, memory-efficient neural network design. In Proceedings of the 49th IEEE/ACM International Symposium on Microarchitecture (MICRO’16) . Taipei, Taiwan, 1–13
Minsoo Rhu, Natalia Gimelshein, Jason Clemons, Arslan Zulfiqar, and Stephen W Keckler. 2016 · 2016
Earlier work this paper cites.
Google’s neural machine translation system: Bridging the gap between human and machine translation
Yonghui Wu, Mike Schuster, Zhifeng Chen, Quoc V Le, Mohammad Norouzi, Wolfgang Macherey, Maxim Krikun, Yuan Cao, Qin Gao, Klaus Macherey, et al · 2016
Earlier work this paper cites.
Accurate, large minibatch sgd: Training imagenet in 1 hour
Priya Goyal, Piotr Dollár, Ross Girshick, Pieter Noordhuis, Lukasz Wesolowski, Aapo Kyrola, Andrew Tulloch, Yangqing Jia, and Kaiming He. 2017 · 2017
Earlier work this paper cites.
Train Longer, Generalize Better: Closing the Generalization Gap in Large Batch Training of Neural Networks. In Proceedings of the 31st International Conference on Neural Information Processing Systems (NeurIPS’17) (Long Beach, California, USA). 1729–1739
Elad Hoffer, Itay Hubara, and Daniel Soudry. 2017 · 2017
Earlier work this paper cites.
Intel Xeon Processors, https://www.intel.com/content/www/us/en/products/details/processors/xeon.html
Intel. 2017 · 2017
Earlier work this paper cites.
Large Batch Training of Convolutional Networks
Yang You, Igor Gitman, and Boris Ginsburg. 2017 · 2017
Earlier work this paper cites.
Large model support for deep learning in caffe and chainer
Minsik Cho, Tung D Le, U Finkler, Haruiki Imai, Yasushi Negishi, Taro Sekiyama, Saritha Vinod, Vladimir Zolotov, Kiyokuni Kawachiya, David S Kung, et al · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018 · 2018
Earlier work this paper cites.
TensorFlow code and pre-trained models for BERT, https://github.com/google-research/bert
Google. 2018 · 2018
Earlier work this paper cites.
PyTorch Pretrained Bert, https://github.com/maknotavailable/pytorch-pretrained-BERT
Huggingface. 2018 · 2018
Earlier work this paper cites.
TensorFlow Large-Model-Support, https://github.com/IBM/tensorflow-large-model-support
IBM. 2018 · 2018
Cited alongside, same era.
Gist: Efficient data encoding for deep neural network training. In The ACM/IEEE 45th Annual International Symposium on Computer Architecture (ISCA’18) . Los Angeles, CA, 7132–7141
Animesh Jain, Amar Phanishayee, Jason Mars, Lingjia Tang, and Gennady Pekhimenko. 2018 · 2018
Cited alongside, same era.
Layer-centric memory reuse and data migration for extreme-scale deep learning on many-core architectures
Hai Jin, Bo Liu, Wenbin Jiang, Yang Ma, Xuanhua Shi, Bingsheng He, and Shaofeng Zhao. 2018 · 2018
Cited alongside, same era.
Pipe-SGD: A Decentralized Pipelined SGD Framework for Distributed Deep Net Training. In Proceedings of the 32nd Conference on Neural Information Processing Systems (NeurIPS’18) . Montreal, Canada, 8056–8067
Youjie Li, Mingchao Yu, Songze Li, Salman Avestimehr, Nam Sung Kim, and Alexander Schwing. 2018 · 2018
Cited alongside, same era.
PyTorch Large-Model-Support, https://github.com/IBM/pytorch-large-model-support
IBM. 2020 · 2020
Later among the works it cites.
Checkmate: Breaking the Memory Wall with Optimal Tensor Rematerialization. In Proceedings of Machine Learning and Systems (MLSys’20) , Vol. 2. 497–511
Paras Jain, Ajay Jain, Aniruddha Nrusimha, Amir Gholami, Pieter Abbeel, Joseph Gonzalez, Kurt Keutzer, and Ion Stoica. 2020 · 2020
Later among the works it cites.
Scaling laws for neural language models
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei. 2020 · 2020
Later among the works it cites.
Pytorch distributed: Experiences on accelerating data parallel training
Shen Li, Yanli Zhao, Rohan Varma, Omkar Salpekar, Pieter Noordhuis, Teng Li, Adam Paszke, Jeff Smith, Brian Vaughan, Pritam Damania, et al · 2020
Later among the works it cites.
PyTorch Distributed: Experiences on Accelerating Data Parallel Training
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
NVIDIA DGX-2H The World’s Most Powerful System for The Most Complex AI Challenges, https://www.nvidia.com/content/dam/en-zz/es_em/Solutions/Data-Center/dgx-2/dgx-2h-datasheet-us-nvidia-841283-r6-web.pdf
NVIDIA. 2018 · 2018
Cited alongside, same era.
GLUE: A multi-task benchmark and analysis platform for natural language understanding
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R Bowman. 2018a · 2018
Cited alongside, same era.
High-density 4U GPU server, https://www.asus.com/us/Commercial-Servers-Workstations/ESC8000-G4
ASUS. 2019 · 2019
Cited alongside, same era.
GPipe: Efficient training of giant neural networks using pipeline parallelism. In Proceedings of the 33st International Conference on Neural Information Processing Systems (NeurIPS’19) . Vancouver, Canada, 103–112
Yanping Huang, Youlong Cheng, Ankur Bapna, Orhan Firat, Mia Xu Chen, Dehao Chen, HyoukJoong Lee, Jiquan Ngiam, Quoc V Le, Yonghui Wu, et al · 2019
Cited alongside, same era.
Evaluating modern GPU interconnect: PCIe, NVLink, NV-SLI, NVSwitch and GPUDirect
Ang Li, Shuaiwen Leon Song, Jieyang Chen, Jiajia Li, Xu Liu, Nathan R Tallent, and Kevin J Barker. 2019 · 2019
Cited alongside, same era.
PipeDream: Generalized Pipeline Parallelism for DNN training. In Proceedings of the 27th ACM Symposium on Operating Systems Principles (SOSP’19) . Huntsville, Canada, 1–15
Deepak Narayanan, Aaron Harlap, Amar Phanishayee, Vivek Seshadri, Nikhil R Devanur, Gregory R Ganger, Phillip B Gibbons, and Matei Zaharia. 2019 · 2019
Cited alongside, same era.
PyTorch: An Imperative Style, High-Performance Deep Learning Library. In Proceedings of the 25th International Conference on Neural Information Processing Systems (NeurIPS’19) . Vancouver, Canada, 8024–8035
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Kopf, Edward Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala. 2019 · 2019
Cited alongside, same era.
Language Models are Unsupervised Multitask Learners
Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019 · 2019
Cited alongside, same era.
Shen Li, Yanli Zhao, Rohan Varma, Omkar Salpekar, Pieter Noordhuis, Teng Li, Adam Paszke, Jeff Smith, Brian Vaughan, Pritam Damania, and Soumith Chintala. 2020a · 2020
Later among the works it cites.
How Do You Know a Human Wrote This?
Farhad Manjoo. 2020 · 2020
Later among the works it cites.
GPT-2 fine-tuning with ONNX Runtime, https://cloudblogs.microsoft.com/opensource/2020/08/24/pytorch-gpt-2-fine-tuning-onnx-runtime-speedup-training-time
Microsoft. 2020 · 2020
Later among the works it cites.
Capuchin: Tensor-based GPU memory management for deep learning. In Proceedings of the 25th International Conference on Architectural Support for Programming Languages and Operating Systems (ASPLOS’20) . Lausanne, Switzerland, 891–905
Xuan Peng, Xuanhua Shi, Hulin Dai, Hai Jin, Weiliang Ma, Qian Xiong, Fan Yang, and Xuehai Qian. 2020 · 2020
Later among the works it cites.
Zero: Memory optimization towards training a trillion parameter models. In Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis (SC’20) (Atlanta, Georgia). 1–16
Samyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, and Yuxiong He. 2020 · 2020
Later among the works it cites.
OpenAI Releases GPT-3, The Largest Model So Far
Ram Sagar. 2020 · 2020
Later among the works it cites.
Fine-tuning giant neural networks on commodity hardware with automatic pipeline model parallelism. In Annual Technical Conference (ATC’21) . 381–396
Saar Eliad, Ido Hakimi, Alon De Jagger, Mark Silberstein, and Assaf Schuster. 2021 · 2021
Later among the works it cites.
DAPPLE: A Pipelined Data Parallel Approach for Training Large Models. In Proceedings of the 26th ACM SIGPLAN Symposium on Principles and Practice of Parallel Programming (PPoPP ’21) . 431–445
Shiqing Fan, Yi Rong, Chen Meng, Zongyan Cao, Siyu Wang, Zhen Zheng, Chuan Wu, Guoping Long, Jun Yang, Lixue Xia, Lansong Diao, Xiaoyong Liu, and Wei Lin. 2021 · 2021
Later among the works it cites.
AI and Memory Wall
Amir Gholami, Zhewei Yao, Sehoon Kim, Michael W Mahoney, and Kurt Keutzer. 2021 · 2021
Later among the works it cites.
Transformer Examples, https://huggingface.co/transformers/v2.3.0/examples.html
Huggingface. 2021 · 2021
Later among the works it cites.
GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding. In Proceedings of the International Conference on Learning Representations (ICLR’21)
Dmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen, Orhan Firat, Yanping Huang, Maxim Krikun, Noam Shazeer, and Zhifeng Chen. 2021 · 2021
Later among the works it cites.
Doing More with Less: Training Large DNN Models on Commodity Servers for the Masses. In Proceedings of the Workshop on Hot Topics in Operating Systems (HotOS’21) . Ann Arbor, Michigan, 119–127
Youjie Li, Amar Phanishayee, Derek Murray, and Nam Sung Kim. 2021 · 2021
Later among the works it cites.
Efficient Large-Scale Language Model Training on GPU Clusters
Deepak Narayanan, Mohammad Shoeybi, Jared Casper, Patrick LeGresley, Mostofa Patwary, Vijay Korthikanti, Dmitri Vainbrand, Prethvi Kashinkunti, Julie Bernauer, Bryan Catanzaro, Amar Phanishayee, and Matei Zaharia. 2021b · 2021
Later among the works it cites.
NVIDIA TESLA GPUs, https://en.wikipedia.org/wiki/Nvidia_Tesla
NVIDIA. 2021 · 2021
Later among the works it cites.
Single Root Complex Purley 4U GPU Server for Deep Learning Applications, https://www.pny.eu/en/consumer/explore-all-products/pny-gpu-servers/\ 983-single-root-complex-purley-4u-gpu\ -server-for-deep-learning-applications
PNY. 2021 · 2021
Later among the works it cites.
ZeRO-Infinity: Breaking the GPU Memory Wall for Extreme Scale Deep Learning
Samyam Rajbhandari, Olatunji Ruwase, Jeff Rasley, Shaden Smith, and Yuxiong He. 2021 · 2021
Later among the works it cites.
ZeRO-Offload: Democratizing Billion-Scale Model Training
Jie Ren, Samyam Rajbhandari, Reza Yazdani Aminabadi, Olatunji Ruwase, Shuangyan Yang, Minjia Zhang, Dong Li, and Yuxiong He. 2021 · 2021
Later among the works it cites.