Fetching the paper…

XC-Cache: Cross-Attending to Cached Context for Efficient LLM Inference · Around