Fetching the paper…

Efficient Transformer-based Large Scale Language Representations using Hardware-friendly Block Structured Pruning · Around