Fetching the paper…
Reading the bibliography…
We provide an optimized implementation of the forward pass of FlashAttention-2, a popular memory-aware scaled dot-product attention algorithm, as a custom fused CUDA kernel targeting NVIDIA Hopper architecture and written using the open-source CUTLASS library.
DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter
1910
Earlier work this paper cites.
Language Models are Few-Shot Learners
2005
Earlier work this paper cites.
Using Shared Memory in CUDA C/C++
2013
Earlier work this paper cites.
Faster Parallel Reductions on Kepler
2014
Earlier work this paper cites.
CUTLASS: Fast Linear Algebra in CUDA C++
2017
Earlier work this paper cites.
Online normalizer calculation for softmax
2018
Cited alongside, same era.
FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness
2022
Cited alongside, same era.
xFormers: A modular and hackable Transformer modelling library
2022
Cited alongside, same era.
Developing CUDA Kernels for Accelerated Matrix Multiplication on NVIDIA Hopper Architecture using the CUTLASS Library
2023
Cited alongside, same era.
FlashAttention — Fast and memory-efficient exact attention . https://github.com/Dao-AILab/flash-attention
Cited in the paper.
FlashAttention adoption
Cited in the paper.
CUTLASS — CUDA Templates for Linear Algebra Subroutines
Cited in the paper.
CuTe Layouts
Cited in the paper.
CuTe Tensors
Cited in the paper.
CuTe’s support for Matrix Multiply-Accumulate instructions
Cited in the paper.
Efficient GEMM in CUDA
Cited in the paper.
NVIDIA H100 Tensor Core GPU Datasheet
Cited in the paper.
2023
Closest in time.
FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning
2023
Closest in time.
Setting New Records at Data Center Scale Using NVIDIA H100 GPUs and NVIDIA Quantum-2 InfiniBand. Ashraf Eassa and Sukru Burc Eryilmaz. November 8, 2023. https://developer.nvidia.com/blog/setting-new-records-at-data-center-scale-using-nvidia-h100-gpus-and-quantum-2-infiniband/
2023
Closest in time.
How Nvidia’s CUDA Monopoly In Machine Learning Is Breaking - OpenAI Triton And PyTorch 2.0
2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…