Fetching the paper…

Accelerating Training of Transformer-Based Language Models with Progressive Layer Dropping · Around