Fetching the paper…

Dynamic Memory Compression: Retrofitting LLMs for Accelerated Inference · Around