Fetching the paper…

Attention Alignment and Flexible Positional Embeddings Improve Transformer Length Extrapolation · Around