Fetching the paper…

Keyformer: KV Cache Reduction through Key Tokens Selection for Efficient Generative Inference · Around