Fetching the paper…

Accelerating Inference in Large Language Models with a Unified Layer Skipping Strategy · Around