Fetching the paper…

CentroidKV: Efficient Long-Context LLM Inference via KV Cache Clustering · Around