Fetching the paper…

RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval · Around