Fetching the paper…

SpargeAttention: Accurate and Training-free Sparse Attention Accelerating Any Model Inference · Around