Fetching the paper…

Accelerating LLaMA Inference by Enabling Intermediate Layer Decoding via Instruction Tuning with LITE · Around