Fetching the paper…

Lossless Acceleration of Large Language Model via Adaptive N-gram Parallel Decoding · Around