Fetching the paper…

Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity · Around