Fetching the paper…

Accelerating LLM Inference with Lossless Speculative Decoding Algorithms for Heterogeneous Vocabularies · Around