Fetching the paper…

Transformer-XL: Attentive Language Models Beyond a Fixed-Length Context · Around