Fetching the paper…

Think Big, Generate Quick: LLM-to-SLM for Fast Autoregressive Decoding · Around