Fetching the paper…

LayerKV: Optimizing Large Language Model Serving with Layer-wise KV Cache Management · Around