Fetching the paper…

Bi-Directional Block Self-Attention for Fast and Memory-Efficient Sequence Modeling · Around