Fetching the paper…

Speculative Decoding via Early-exiting for Faster LLM Inference with Thompson Sampling Control Mechanism · Around