Fetching the paper…

CaraServe: CPU-Assisted and Rank-Aware LoRA Serving for Generative LLM Inference · Around