Fetching the paper…
Reading the bibliography…
Large-scale model training has been a playing ground for a limited few requiring complex model refactoring and access to prohibitively expensive GPU clusters.
Parallelized Stochastic Gradient Descent
M. A. Zinkevich, M.Weimer, A. Smola, and L. Li · 2010
Earlier work this paper cites.
Large scale distributed deep networks
Jeffrey Dean, Greg Corrado, Rajat Monga, Kai Chen, Matthieu Devin, Mark Mao, Marc’aurelio Ranzato, Andrew Senior, Paul Tucker, Ke Yang, Quoc V. Le, and Andrew Y. Ng · 2012
Earlier work this paper cites.
Use of simd vector operations to accelerate application code performance on low-powered arm and intel platforms
Gaurav Mitra, Beau Johnston, Alistair Rendell, Eric McCreath, and Jun Zhou · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization, 2014
Diederik P. Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
The performance impact analysis of loop unrolling
G. Velkoski, M. Gusev, and S. Ristov · 2014
Earlier work this paper cites.
Training deep nets with sublinear memory cost
Tianqi Chen, Bing Xu, Chiyuan Zhang, and Carlos Guestrin · 2016
Earlier work this paper cites.
Squad: 100, 000+ questions for machine comprehension of text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang · 2016
Earlier work this paper cites.
vdnn: Virtualized deep neural networks for scalable, memory-efficient neural network design
Minsoo Rhu, Natalia Gimelshein, Jason Clemons, Arslan Zulfiqar, and Stephen W. Keckler · 2016
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Earlier work this paper cites.
Pipedream: Fast and efficient pipeline parallel dnn training, 2018
Aaron Harlap, Deepak Narayanan, Amar Phanishayee, Vivek Seshadri, Nikhil Devanur, Greg Ganger, and Phil Gibbons · 2018
Earlier work this paper cites.
Gpipe: Efficient training of giant neural networks using pipeline parallelism, 2018
Yanping Huang, Youlong Cheng, Ankur Bapna, Orhan Firat, Mia Xu Chen, Dehao Chen, HyoukJoong Lee, Jiquan Ngiam, Quoc V. Le, Yonghui Wu, and Zhifeng Chen · 2018
Earlier work this paper cites.
Layer-centric memory reuse and data migration for extreme-scale deep learning on many-core architectures
Hai Jin, Bo Liu, Wenbin Jiang, Yang Ma, Xuanhua Shi, Bingsheng He, and Shaofeng Zhao · 2018
Cited alongside, same era.
Mesh-tensorflow: Deep learning for supercomputers
Noam Shazeer, Youlong Cheng, Niki Parmar, Dustin Tran, Ashish Vaswani, Penporn Koanantakool, Peter Hawkins, HyoukJoong Lee, Mingsheng Hong, Cliff Young, Ryan Sepassi, and Blake A. Hechtman · 2018
Cited alongside, same era.
Superneurons: Dynamic gpu memory management for training deep neural networks
Linnan Wang, Jinmian Ye, Yiyang Zhao, Wei Wu, Ang Li, Shuaiwen Leon Song, Zenglin Xu, and Tim Kraska · 2018
Cited alongside, same era.
Automatic Mixed Precision for Deep Learning
Nvidia · 2019
Cited alongside, same era.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever · 2019
Cited alongside, same era.
Swapadvisor: Pushing deep learning beyond the gpu memory limit via smart swapping
Chien-Chin Huang, Gu Jin, and Jinyang Li · 2020
Later among the works it cites.
Scaling laws for neural language models, 2020
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B. Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei · 2020
Later among the works it cites.
Pytorch distributed: Experiences on accelerating data parallel training
Shen Li, Yanli Zhao, Rohan Varma, Omkar Salpekar, Pieter Noordhuis, Teng Li, Adam Paszke, Jeff Smith, Brian Vaughan, Pritam Damania, and Soumith Chintala · 2020
Later among the works it cites.
Capuchin: Tensor-based gpu memory management for deep learning
Xuan Peng, Xuanhua Shi, Hulin Dai, Hai Jin, Weiliang Ma, Qian Xiong, Fan Yang, and Xuehai Qian · 2020
Later among the works it cites.
Training large neural networks with constant memory using a new execution algorithm
Bharadwaj Pudipeddi, Maral Mesmakhosroshahi, Jinwen Xi, and Sujeeth Bharadwaj · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Christopher J. Shallue, Jaehoon Lee, Joseph Antognini, Jascha Sohl-Dickstein, Roy Frostig, and George E. Dahl · 2019
Cited alongside, same era.
Megatron-lm: Training multi-billion parameter language models using model parallelism
Mohammad Shoeybi, Mostofa Patwary, Raul Puri, Patrick LeGresley, Jared Casper, and Bryan Catanzaro · 2019
Cited alongside, same era.
Efficient memory management for gpu-based deep learning systems
Junzhe Zhang, Sai-Ho Yeung, Yao Shu, Bingsheng He, and Wei Wang · 2019
Cited alongside, same era.
Language models are few-shot learners, 2020
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei · 2020
Cited alongside, same era.
Pretrained language models for dialogue generation with multiple input sources, 2020
Yu Cao, Wei Bi, Meng Fang, and Dacheng Tao · 2020
Cited alongside, same era.
Autotm: Automatic tensor movement in heterogeneous memory systems using integer linear programming
Mark Hildebrand, Jawad Khan, Sanjeev Trika, Jason Lowe-Power, and Venkatesh Akella · 2020
Cited alongside, same era.
https://rajpurkar.github.io/SQuAD-explorer/
The Stanford Question Answering Dataset (SQuAD) leaderboard
Cited in the paper.
Exploring the limits of transfer learning with a unified text-to-text transformer, 2020
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu · 2020
Later among the works it cites.
ZeRO: Memory Optimizations Toward Training Trillion Parameter Models
Samyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, and Yuxiong He · 2020
Later among the works it cites.
Sentinel: Efficient Tensor Migration and Allocation on Heterogeneous Memory Systems for Deep Learning
Jie Ren, Jiaolin Luo, Kai Wu, Minjia Zhang, Hyeran Jeon, and Dong Li · 2020
Later among the works it cites.
Turing-nlg: A 17-billion-parameter language model by microsoft, 2020
Corby Rosset · 2020
Later among the works it cites.
PyTorch implementation of L2L execution algorithm
Roman Tezikov · 2020
Later among the works it cites.
Buddy Compression: Enabling Larger Memory for Deep Learning and HPC Workloads on GPUs
Oreste Villa, Mark Stephenson, David Nellans, and Stephen Keckler · 2020
Later among the works it cites.