Fetching the paper…
Reading the bibliography…
We propose a novel self-attention mechanism that can learn its optimal attention span.
Transformer-xl: Attentive language models beyond a fixed-length context
Zihang Dai, Zhilin Yang, Yiming Yang, Jaime G. Carbonell, Quoc V. Le, and Ruslan Salakhutdinov. 2019 · 1901
Earlier work this paper cites.
Large text compression benchmark
Matt Mahoney. 2011 · 2011
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. 2015 · 2015
Earlier work this paper cites.
Effective approaches to attention-based neural machine translation
Thang Luong, Hieu Pham, and Christopher D. Manning. 2015 · 2015
Earlier work this paper cites.
End-to-end memory networks
Sainbayar Sukhbaatar, Arthur Szlam, Jason Weston, and Rob Fergus. 2015 · 2015
Cited alongside, same era.
Adaptive computation time for recurrent neural networks
Alex Graves. 2016 · 2016
Cited alongside, same era.
Variable computation in recurrent neural networks
Yacine Jernite, Edouard Grave, Armand Joulin, and Tomas Mikolov. 2017 · 2017
Cited alongside, same era.
An empirical study of adequate vision span for attention-based neural machine translation
Raphael Shu and Hideki Nakayama. 2017 · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Later among the works it cites.
Self-attention with relative position representations
Peter Shaw, Jakob Uszkoreit, and Ashish Vaswani. 2018 · 2018
Later among the works it cites.
Character-level language modeling with deeper self-attention
Rami Al-Rfou, Dokook Choe, Noah Constant, Mandy Guo, and Llion Jones. 2019 · 2019
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…