Fetching the paper…
Reading the bibliography…
Training Transformer models on long sequences in a distributed setting poses significant challenges in terms of efficiency and scalability.
Data parallel algorithms
W Daniel Hillis and Guy L Steele Jr · 1986
Earlier work this paper cites.
Adam: A method for stochastic optimization, 2017
Diederik P. Kingma and Jimmy Ba · 2017
Earlier work this paper cites.
Gpipe: Efficient training of giant neural networks using pipeline parallelism, 2018
Yanping Huang, Youlong Cheng, Ankur Bapna, Orhan Firat, Mia Xu Chen, Dehao Chen, HyoukJoong Lee, Jiquan Ngiam, Quoc V. Le, Yonghui Wu, and Zhifeng Chen · 2018
Earlier work this paper cites.
Mixed precision training, 2018
Paulius Micikevicius, Sharan Narang, Jonah Alben, Gregory Diamos, Erich Elsen, David Garcia, Boris Ginsburg, Michael Houston, Oleksii Kuchaiev, Ganesh Venkatesh, and Hao Wu · 2018
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever · 2018
Earlier work this paper cites.
Imagenet training in minutes
Yang You, Zhao Zhang, Cho-Jui Hsieh, James Demmel, and Kurt Keutzer · 2018
Earlier work this paper cites.
Pytorch: An imperative style, high-performance deep learning library, 2019
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Köpf, Edward Yang, Zach DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala · 2019
Earlier work this paper cites.
Fast transformer decoding: One write-head is all you need, 2019
Noam Shazeer · 2019
Earlier work this paper cites.
Megatron-lm: Training multi-billion parameter language models using model parallelism
Mohammad Shoeybi, Mostofa Patwary, Raul Puri, Patrick LeGresley, Jared Casper, and Bryan Catanzaro · 2019
Earlier work this paper cites.
Longformer: The long-document transformer, 2020
Iz Beltagy, Matthew E. Peters, and Arman Cohan · 2020
Earlier work this paper cites.
Dapple: A pipelined data parallel approach for training large models, 2020
Shiqing Fan, Yi Rong, Chen Meng, Zongyan Cao, Siyu Wang, Zhen Zheng, Chuan Wu, Guoping Long, Jun Yang, Lixue Xia, Lansong Diao, Xiaoyong Liu, and Wei Lin · 2020
Earlier work this paper cites.
Taming unbalanced training workloads in deep learning with partial collective operations
Shigang Li, Tal Ben-Nun, Salvatore Di Girolamo, Dan Alistarh, and Torsten Hoefler · 2020
Earlier work this paper cites.
Nvidia collective communications library, 2020
NVIDIA · 2020
Earlier work this paper cites.
Zero: Memory optimizations toward training trillion parameter models, 2020
Samyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, and Yuxiong He · 2020
Earlier work this paper cites.
Colossal-ai: A unified deep learning system for large-scale parallel training
Zhengda Bian, Hongxin Liu, Boxiang Wang, Haichen Huang, Yongbin Li, Chuanrui Wang, Fan Cui, and Yang You · 2021
Cited alongside, same era.
Highly accurate protein structure prediction with alphafold
John M. Jumper, Richard Evans, Alexander Pritzel, Tim Green, Michael Figurnov, Olaf Ronneberger, Kathryn Tunyasuvunakool, Russ Bates, Augustin Zídek, Anna Potapenko, Alex Bridgland, Clemens Meyer, Simon A A Kohl, Andy Ballard, Andrew Cowie, Bernardino Romera-Paredes, Stanislav Nikolov, Rishub Jain, Jonas Adler, Trevor Back, Stig Petersen, David A. Reiman, Ellen Clancy, Michal Zielinski, Martin Steinegger, Michalina Pacholska, Tamas Berghammer, Sebastian Bodenstein, David Silver, Oriol Vinyals, Andrew W. Senior, Koray Kavukcuoglu, Pushmeet Kohli, and Demis Hassabis · 2021
Cited alongside, same era.
Sequence parallelism: Long sequence training from system perspective
Shenggui Li, Fuzhao Xue, Yongbin Li, and Yang You · 2021
Cited alongside, same era.
Chimera: efficiently training large-scale neural networks with bidirectional pipelines
Shigang Li and Torsten Hoefler · 2021
Cited alongside, same era.
Efficient memory management for large language model serving with pagedattention, 2023
Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph E. Gonzalez, Hao Zhang, and Ion Stoica · 2023
Later among the works it cites.
Ring attention with blockwise transformers for near-infinite context, 2023
Hao Liu, Matei Zaharia, and Pieter Abbeel · 2023
Later among the works it cites.
Hanayo: Harnessing wave-like pipeline parallelism for enhanced large model training efficiency
Ziming Liu, Shenggan Cheng, Haotian Zhou, and Yang You · 2023
Later among the works it cites.
Scalable diffusion models with transformers, 2023
William Peebles and Saining Xie · 2023
Later among the works it cites.
Scalable diffusion models with transformers, 2023
William Peebles and Saining Xie · 2023
Later among the works it cites.
Attention is all you need, 2023
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Deepak Narayanan, Mohammad Shoeybi, Jared Casper, Patrick LeGresley, Mostofa Patwary, Vijay Anand Korthikanti, Dmitri Vainbrand, Prethvi Kashinkunti, Julie Bernauer, Bryan Catanzaro, Amar Phanishayee, and Matei Zaharia · 2021
Cited alongside, same era.
An efficient 2d method for training super-large deep learning models
Qifan Xu, Shenggui Li, Chaoyu Gong, and Yang You · 2021
Cited alongside, same era.
Flashattention: Fast and memory-efficient exact attention with io-awareness, 2022
Tri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra, and Christopher Ré · 2022
Cited alongside, same era.
An empirical survey on long document summarization: Datasets, models, and metrics
Huan Yee Koh, Jiaxin Ju, Ming Liu, and Shirui Pan · 2022
Cited alongside, same era.
Reducing activation recomputation in large transformer models, 2022
Vijay Korthikanti, Jared Casper, Sangkug Lym, Lawrence McAfee, Michael Andersch, Mohammad Shoeybi, and Bryan Catanzaro · 2022
Cited alongside, same era.
Self-attention does not need o ( n 2 ) o(n^{2}) memory, 2022
Markus N. Rabe and Charles Staats · 2022
Cited alongside, same era.
Tesseract: Parallelize the tensor parallelism efficiently
Boxiang Wang, Qifan Xu, Zhengda Bian, and Yang You · 2022
Cited alongside, same era.
Gpt-4 technical report, 2023
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, et al · 2023
Cited alongside, same era.
Later among the works it cites.
Tim Brooks, Bill Peebles, Connor Holmes, Will DePue, Yufei Guo, Li Jing, David Schnurr, Joe Taylor, Troy Luhman, Eric Luhman, Clarence Ng, Ricky Wang, and Aditya Ramesh, Feb 2024
2024
Closest in time.
Centauri: Enabling efficient scheduling for communication-computation overlap in large model training via communication partitioning
Chang Chen, Xiuhong Li, Qianchao Zhu, Jiangfei Duan, Peng Sun, Xingcheng Zhang, and Chao Yang · 2024
Closest in time.
Fastfold: Optimizing alphafold training and inference on gpu clusters
Shenggan Cheng, Xuanlei Zhao, Guangyang Lu, Jiarui Fang, Tian Zheng, Ruidong Wu, Xiwen Zhang, Jian Peng, and Yang You · 2024
Closest in time.
Scaling rectified flow transformers for high-resolution image synthesis, 2024
Patrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari, Jonas Müller, Harry Saini, Yam Levi, Dominik Lorenz, Axel Sauer, Frederic Boesel, Dustin Podell, Tim Dockhorn, Zion English, Kyle Lacey, Alex Goodwin, Yannik Marek, and Robin Rombach · 2024
Closest in time.
Usp: A unified sequence parallelism approach for long context generative ai, 2024
Jiarui Fang and Shangchun Zhao · 2024
Closest in time.
Distflashattn: Distributed memory-efficient attention for long-context llms training, 2024
Dacheng Li, Rulin Shao, Anze Xie, Eric P. Xing, Xuezhe Ma, Ion Stoica, Joseph E. Gonzalez, and Hao Zhang · 2024
Closest in time.
Flashattention-3: Fast and accurate attention with asynchrony and low-precision, 2024
Jay Shah, Ganesh Bikshandi, Ying Zhang, Vijay Thakkar, Pradeep Ramani, and Tri Dao · 2024
Closest in time.
Ring flash attention, 2024
Zilin Zhu et al · 2024
Closest in time.