Fetching the paper…

DiffKV: Differentiated Memory Management for Large Language Models with Parallel KV Compaction · Around