Fetching the paper…
Reading the bibliography…
FlashAttention series has been widely applied in the inference of large language models (LLMs).
The diffusion of development
Enrico Spolaore and Romain Wacziarg · 2009
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Volta: Performance and programmability
Jack Choquette, Olivier Giroux, and Denis Foley · 2018
Earlier work this paper cites.
Nvidia tensor core programmability, performance & precision
Stefano Markidis, Steven Wei Der Chien, Erwin Laure, Ivy Bo Peng, and Jeffrey S Vetter · 2018
Earlier work this paper cites.
Online normalizer calculation for softmax
Maxim Milakov and Natalia Gimelshein · 2018
Earlier work this paper cites.
Deep learning for computer vision: A brief review
Athanasios Voulodimos, Nikolaos Doulamis, Anastasios Doulamis, and Eftychios Protopapadakis · 2018
Earlier work this paper cites.
Longformer: The long-document transformer
Iz Beltagy, Matthew E Peters, and Arman Cohan · 2020
Earlier work this paper cites.
A survey of accelerator architectures for deep neural networks
Yiran Chen, Yuan Xie, Linghao Song, Fan Chen, and Tianqi Tang · 2020
Earlier work this paper cites.
Hanguang 800 npu–the ultimate ai inference solution for data centers
Yang Jiao, Liang Han, and Xin Long · 2020
Earlier work this paper cites.
Scaling laws for neural language models
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei · 2020
Earlier work this paper cites.
Megatron-lm: Training multi-billion parameter language models using model parallelism
Mohammad Shoeybi, Mostofa Patwary, Raul Puri, Patrick LeGresley, Jared Casper, and Bryan Catanzaro · 2020
Earlier work this paper cites.
Nvidia a100 tensor core gpu: Performance and innovation
Jack Choquette, Wishwesh Gandhi, Olivier Giroux, Nick Stam, and Ronny Krashinsky · 2021
Earlier work this paper cites.
Dissecting the ampere gpu architecture via microbenchmarking
Zhe Jia and Peter Van Sandt · 2021
Earlier work this paper cites.
Ascend: a scalable and unified architecture for ubiquitous deep neural network computing: Industry track paper
Heng Liao, Jiajin Tu, Jing Xia, Hu Liu, Xiping Zhou, Honghui Yuan, and Yuxing Hu · 2021
Cited alongside, same era.
Self-attention does not need o ( n 2 ) o(n^{2}) memory
Markus N Rabe and Charles Staats · 2021
Cited alongside, same era.
Tokens-to-token vit: Training vision transformers from scratch on imagenet
Li Yuan, Yunpeng Chen, Tao Wang, Weihao Yu, Yujun Shi, Zi-Hang Jiang, Francis EH Tay, Jiashi Feng, and Shuicheng Yan · 2021
Cited alongside, same era.
Wei Zeng, Xiaozhe Ren, Teng Su, Hui Wang, Yi Liao, Zhiwei Wang, Xin Jiang, ZhenZhang Yang, Kaisheng Wang, Xiaoda Zhang, et al · 2021
Cited alongside, same era.
Ai accelerator embedded computational storage for large-scale dnn models
Flashattention-2: Faster attention with better parallelism and work partitioning
Tri Dao · 2023
Later among the works it cites.
Hnpu-v1: An adaptive dnn training processor utilizing stochastic dynamic fixed-point and active bit-precision searching
Donghyeon Han and Hoi-Jun Yoo · 2023
Later among the works it cites.
Cutlass: a collection of cuda c++ template abstractions for implementing high-performance matrix-matrix multiplication (gemm) and related computations at all levels and scales within cuda
Nvidia · 2023
Later among the works it cites.
Baolin Peng, Chunyuan Li, Pengcheng He, Michel Galley, and Jianfeng Gao · 2023
Later among the works it cites.
Pangu-{ Σ \Sigma }: Towards trillion parameter language model with sparse heterogeneous computing
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Byungmin Ahn, Jaehun Jang, Hanbyeul Na, Mankeun Seo, Hongrak Son, and Yong Ho Song · 2022
Cited alongside, same era.
Deepspeed-inference: enabling efficient inference of transformer models at unprecedented scale
Reza Yazdani Aminabadi, Samyam Rajbhandari, Ammar Ahmad Awan, Cheng Li, Du Li, Elton Zheng, Olatunji Ruwase, Shaden Smith, Minjia Zhang, Jeff Rasley, et al · 2022
Cited alongside, same era.
Nvidia hopper gpu: Scaling performance
Jack Choquette · 2022
Cited alongside, same era.
Flashattention: Fast and memory-efficient exact attention with io-awareness
Tri Dao, Dan Fu, Stefano Ermon, Atri Rudra, and Christopher Ré · 2022
Cited alongside, same era.
xformers: A modular and hackable transformer modelling library
Benjamin Lefaudeux, Francisco Massa, Diana Liskovich, Wenhan Xiong, Vittorio Caggiano, Sean Naren, Min Xu, Jieru Hu, Marta Tintore, Susan Zhang, Patrick Labatut, Daniel Haziza, Luca Wehrstedt, Jeremy Reizenstein, and Grigory Sizov · 2022
Cited alongside, same era.
Opt: Open pre-trained transformer language models, 2022
Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li, Xi Victoria Lin, Todor Mihaylov, Myle Ott, Sam Shleifer, Kurt Shuster, Daniel Simig, Punit Singh Koura, Anjali Sridhar, Tianlu Wang, and Luke Zettlemoyer · 2022
Cited alongside, same era.
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al · 2023
Cited alongside, same era.
Gqa: Training generalized multi-query transformer models from multi-head checkpoints
Joshua Ainslie, James Lee-Thorp, Michiel de Jong, Yury Zemlyanskiy, Federico Lebrón, and Sumit Sanghai · 2023
Cited alongside, same era.
Xiaozhe Ren, Pingyi Zhou, Xinfan Meng, Xinjing Huang, Yadao Wang, Weichao Wang, Pengfei Li, Xiaoda Zhang, Alexander Podolskiy, Grigory Arshinov, et al · 2023
Later among the works it cites.
Flexgen: High-throughput generative inference of large language models with a single gpu
Ying Sheng, Lianmin Zheng, Binhang Yuan, Zhuohan Li, Max Ryabinin, Beidi Chen, Christopher Liang, Ion Stoica, and Ce Zhang · 2023
Later among the works it cites.
Cambricon-r: A fully fused accelerator for real-time learning of neural scene representation
Xinkai Song, Yuanbo Wen, Xing Hu, Tianbo Liu, Haoxuan Zhou, Husheng Han, Tian Zhi, Zidong Du, Wei Li, Rui Zhang, et al · 2023
Later among the works it cites.
Leveraging the opt large language model for sentiment analysis of game reviews
Markos Viggiato and Cor-Paul Bezemer · 2023
Later among the works it cites.
Fast distributed inference serving for large language models
Bingyang Wu, Yinmin Zhong, Zili Zhang, Gang Huang, Xuanzhe Liu, and Xin Jin · 2023
Later among the works it cites.
A survey of large language models
Wayne Xin Zhao, Kun Zhou, Junyi Li, Tianyi Tang, Xiaolei Wang, Yupeng Hou, Yingqian Min, Beichen Zhang, Junjie Zhang, Zican Dong, et al · 2023
Later among the works it cites.
Nsight compute profiling: Kernel profiling guide with metric types and meaning, data collection modes and faq for common problems
Nvidia · 2024
Closest in time.
Flashattention-3: Fast and accurate attention with asynchrony and low-precision
Jay Shah, Ganesh Bikshandi, Ying Zhang, Vijay Thakkar, Pradeep Ramani, and Tri Dao · 2024
Closest in time.