Fetching the paper…

ChunkAttention: Efficient Self-Attention with Prefix-Aware KV Cache and Two-Phase Partition · Around