Fetching the paper…

Improving Transformers with Dynamically Composable Multi-Head Attention · Around