Fetching the paper…

RazorAttention: Efficient KV Cache Compression Through Retrieval Heads · Around