Fetching the paper…
Reading the bibliography…
In the last three years, the largest dense deep learning models have grown over 1000x to reach hundreds of billions of parameters, while the GPU memory has only grown by 5x (16 GB to 80 GB).
Adaptive subgradient methods for online learning and stochastic optimization
John Duchi, Elad Hazan, and Yoram Singer · 2011
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
vdnn: Virtualized deep neural networks for scalable, memory-efficient neural network design
Minsoo Rhu, Natalia Gimelshein, Jason Clemons, Arslan Zulfiqar, and Stephen W Keckler · 2016
Earlier work this paper cites.
Training deep nets with sublinear memory cost
Tianqi Chen, Bing Xu, Chiyuan Zhang, and Carlos Guestrin · 2016
Earlier work this paper cites.
Scaling SGD batch size to 32k for imagenet training
Yang You, Igor Gitman, and Boris Ginsburg · 2017
Earlier work this paper cites.
Mixed precision training, 2017
Paulius Micikevicius, Sharan Narang, Jonah Alben, Gregory Diamos, Erich Elsen, David Garcia, Boris Ginsburg, Michael Houston, Oleksii Kuchaiev, Ganesh Venkatesh, and Hao Wu · 2017
Earlier work this paper cites.
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Ultra-performance pascal gpu and nvlink interconnect
Denis Foley and John Danskin · 2017
Earlier work this paper cites.
BERT: pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Earlier work this paper cites.
Deep contextualized word representations
Matthew E Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer · 2018
Earlier work this paper cites.
Pipedream: Fast and efficient pipeline parallel DNN training
Aaron Harlap, Deepak Narayanan, Amar Phanishayee, Vivek Seshadri, Nikhil R. Devanur, Gregory R. Ganger, and Phillip B. Gibbons · 2018
Earlier work this paper cites.
Gpipe: Efficient training of giant neural networks using pipeline parallelism
Yanping Huang, Yonglong Cheng, Dehao Chen, HyoukJoong Lee, Jiquan Ngiam, Quoc V. Le, and Zhifeng Chen · 2018
Earlier work this paper cites.
Mesh-tensorflow: Deep learning for supercomputers
Noam Shazeer, Youlong Cheng, Niki Parmar, Dustin Tran, Ashish Vaswani, Penporn Koanantakool, Peter Hawkins, HyoukJoong Lee, Mingsheng Hong, Cliff Young, Ryan Sepassi, and Blake A. Hechtman · 2018
Earlier work this paper cites.
Layer-centric memory reuse and data migration for extreme-scale deep learning on many-core architectures
Hai Jin, Bo Liu, Wenbin Jiang, Yang Ma, Xuanhua Shi, Bingsheng He, and Shaofeng Zhao · 2018
Earlier work this paper cites.
Superneurons: Dynamic gpu memory management for training deep neural networks
Linnan Wang, Jinmian Ye, Yiyang Zhao, Wei Wu, Ang Li, Shuaiwen Leon Song, Zenglin Xu, and Tim Kraska · 2018
Earlier work this paper cites.
Gist: Efficient data encoding for deep neural network training
Animesh Jain, Amar Phanishayee, Jason Mars, Lingjia Tang, and Gennady Pekhimenko · 2018
Cited alongside, same era.
Superneurons: Dynamic GPU memory management for training deep neural networks
Linnan Wang, Jinmian Ye, Yiyang Zhao, Wei Wu, Ang Li, Shuaiwen Leon Song, Zenglin Xu, and Tim Kraska · 2018
Cited alongside, same era.
Improving language understanding by generative pre-training
Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever · 2018
Cited alongside, same era.
Language models are unsupervised multitask learners
Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever · 2019
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer, 2019
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu · 2019
Cited alongside, same era.
ZeRO: Memory Optimizations toward Training Trillion Parameter Models
Samyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, and Yuxiong He · 2020
Later among the works it cites.
DeepSpeed: Extreme-scale model training for everyone
Microsoft · 2020
Later among the works it cites.
DeepSpeed: Extreme-scale model training for everyone
DeepSpeed Team and Rangan Majumder · 2020
Later among the works it cites.
Autotm: Automatic tensor movement in heterogeneous memory systems using integer linear programming
Mark Hildebrand, Jawad Khan, Sanjeev Trika, Jason Lowe-Power, and Venkatesh Akella · 2020
Later among the works it cites.
Swapadvisor: Pushing deep learning beyond the gpu memory limit via smart swapping
Chien-Chin Huang, Gu Jin, and Jinyang Li · 2020
Later among the works it cites.
Capuchin: Tensor-based gpu memory management for deep learning
Xuan Peng, Xuanhua Shi, Hulin Dai, Hai Jin, Weiliang Ma, Qian Xiong, Fan Yang, and Xuehai Qian · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Megatron-lm: Training multi-billion parameter language models using model parallelism, 2019
Mohammad Shoeybi, Mostofa Patwary, Raul Puri, Patrick LeGresley, Jared Casper, and Bryan Catanzaro · 2019
Cited alongside, same era.
Supporting very large models using automatic dataflow graph partitioning
Minjie Wang, Chien-chin Huang, and Jinyang Li · 2019
Cited alongside, same era.
Demystifying parallel and distributed deep learning: An in-depth concurrency analysis
Tal Ben-Nun and Torsten Hoefler · 2019
Cited alongside, same era.
Checkmate: Breaking the memory wall with optimal tensor rematerialization
Paras Jain, Ajay Jain, Aniruddha Nrusimha, Amir Gholami, Pieter Abbeel, Kurt Keutzer, Ion Stoica, and Joseph E. Gonzalez · 2019
Cited alongside, same era.
Reducing BERT pre-training time from 3 days to 76 minutes
Yang You, Jing Li, Jonathan Hseu, Xiaodan Song, James Demmel, and Cho-Jui Hsieh · 2019
Cited alongside, same era.
NVIDIA DGX SuperPOD delivers world record supercomputing to any enterprise
NVIDIA · 2019
Cited alongside, same era.
Language models are few-shot learners
Tom B Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Cited alongside, same era.
Later among the works it cites.
Distributed hierarchical gpu parameter server for massive scale deep learning ads systems, 2020
Weijie Zhao, Deping Xie, Ronglai Jia, Yulei Qian, Ruiquan Ding, Mingming Sun, and Ping Li · 2020
Later among the works it cites.
Pytorch distributed: Experiences on accelerating data parallel training
Shen Li, Yanli Zhao, Rohan Varma, Omkar Salpekar, Pieter Noordhuis, Teng Li, Adam Paszke, Jeff Smith, Brian Vaughan, Pritam Damania, et al · 2020
Later among the works it cites.
ZeRO-Offload: Democratizing Billion-Scale Model Training, 2021
Jie Ren, Samyam Rajbhandari, Reza Yazdani Aminabadi, Olatunji Ruwase, Shuangyan Yang, Minjia Zhang, Dong Li, and Yuxiong He · 2021
Closest in time.
Megatron-LM: Software repository
Mohammad Shoeybi, Mostofa Patwary, Raul Puri, Patrick LeGresley, Jared Casper, and Bryan Catanzaro · 2021
Closest in time.
Ai and memory wall
Amir Gholami, Zhewei Yao, Sehoon Kim, Michael W. Mahoney, and Kurt Keutzer · 2021
Closest in time.
Sentinel: Efficient tensor migration and allocation on heterogeneous memory systems for deep learning
Jie Ren, Jiaolin Luo, Kai Wu, Minjia Zhang, Hyeran Jeon, and Dong Li · 2021
Closest in time.
https://www.nvidia.com/en-us/data-center/tensor-cores/ , 2018
NVIDIA Tensor Cores · 2021
Closest in time.
ORNL Launches Summit Supercomputer
Oak Ridge National Laboratory · 2021
Closest in time.