Fetching the paper…

PARD: Accelerating LLM Inference with Low-Cost PARallel Draft Model Adaptation · Around