Fetching the paper…
Reading the bibliography…
In sequence to sequence learning, the self-attention mechanism proves to be highly effective, and achieves significant improvements in many tasks.
Pay less attention with lightweight and dynamic convolutions
Felix Wu, Angela Fan, Alexei Baevski, Yann N Dauphin, and Michael Auli · 1901
Earlier work this paper cites.
Sequence to sequence learning with neural networks, 2014
Ilya Sutskever, Oriol Vinyals, and Quoc V. Le · 2014
Earlier work this paper cites.
Weighted transformer network for machine translation, 2017
Karim Ahmed, Nitish Shirish Keskar, and Richard Socher · 2017
Earlier work this paper cites.
Xception: Deep learning with depthwise separable convolutions
Francois Chollet · 2017
Earlier work this paper cites.
Convolutional sequence to sequence learning
Jonas Gehring, Michael Auli, David Grangier, Denis Yarats, and Yann N Dauphin · 2017
Earlier work this paper cites.
Towards neural phrase-based machine translation
Po-Sen Huang, Chong Wang, Sitao Huang, Dengyong Zhou, and Li Deng · 2017
Earlier work this paper cites.
Depthwise separable convolutions for neural machine translation, 2017
Lukasz Kaiser, Aidan N. Gomez, and Francois Chollet · 2017
Earlier work this paper cites.
Attention is all you need, 2017
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Universal transformers, 2018
Mostafa Dehghani, Stephan Gouws, Oriol Vinyals, Jakob Uszkoreit, and Łukasz Kaiser · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding, 2018
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Cited alongside, same era.
Neural phrase-to-phrase machine translation, 2018
Jiangtao Feng, Lingpeng Kong, Po-Sen Huang, Chong Wang, Da Huang, Jiayuan Mao, Kan Qiao, and Dengyong Zhou · 2018
Cited alongside, same era.
Layer-wise coordination between encoder and decoder for neural machine translation
Tianyu He, Xu Tan, Yingce Xia, Di He, Tao Qin, Zhibo Chen, and Tie-Yan Liu · 2018
Cited alongside, same era.
Learning when to concentrate or divert attention: Self-adaptive attention temperature for neural machine translation
Junyang Lin, Xu Sun, Xuancheng Ren, Muyu Li, and Qi Su · 2018
Cited alongside, same era.
Scaling neural machine translation
Myle Ott, Sergey Edunov, David Grangier, and Michael Auli · 2018
Cited alongside, same era.
Joint source-target self attention with locality constraints, 2019
José A. R. Fonollosa, Noe Casas, and Marta R. Costa-jussà · 2019
Closest in time.
Understanding and improving transformer from a multi-particle dynamic system point of view, 2019
Yiping Lu, Zhuohan Li, Di He, Zhiqing Sun, Bin Dong, Tao Qin, Liwei Wang, and Tie-Yan Liu · 2019
Closest in time.
fairseq: A fast, extensible toolkit for sequence modeling
Myle Ott, Sergey Edunov, Alexei Baevski, Angela Fan, Sam Gross, Nathan Ng, David Grangier, and Michael Auli · 2019
Closest in time.
Stand-alone self-attention in vision models, 2019
Prajit Ramachandran, Niki Parmar, Ashish Vaswani, Irwan Bello, Anselm Levskaya, and Jonathon Shlens · 2019
Closest in time.
David R So, Chen Liang, and Quoc V Le · 2019
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Why self-attention? A targeted evaluation of neural machine translation architectures
Gongbo Tang, Mathias Müller, Annette Rios, and Rico Sennrich · 2018
Cited alongside, same era.
Qanet: Combining local convolution with global self-attention for reading comprehension, 2018
Adams Wei Yu, David Dohan, Minh-Thang Luong, Rui Zhao, Kai Chen, Mohammad Norouzi, and Quoc V. Le · 2018
Cited alongside, same era.
Generating long sequences with sparse transformers, 2019
Rewon Child, Scott Gray, Alec Radford, and Ilya Sutskever · 2019
Cited alongside, same era.
Transformer-xl: Attentive language models beyond a fixed-length context
Zihang Dai, Zhilin Yang, Yiming Yang, Jaime Carbonell, Quoc Le, and Ruslan Salakhutdinov · 2019
Cited alongside, same era.
Adaptive attention span in transformers
Sainbayar Sukhbaatar, Edouard Grave, Piotr Bojanowski, and Armand Joulin
Cited in the paper.
Augmenting self-attention with persistent memory, 2019b
Sainbayar Sukhbaatar, Edouard Grave, Guillaume Lample, Herve Jegou, and Armand Joulin
Cited in the paper.
Depth growing for neural machine translation
Lijun Wu, Yiren Wang, Yingce Xia, Fei Tian, Fei Gao, Tao Qin, Jianhuang Lai, and Tie-Yan Liu
Cited in the paper.
Xlnet: Generalized autoregressive pretraining for language understanding, 2019
Zhilin Yang, Zihang Dai, Yiming Yang, Jaime Carbonell, Ruslan Salakhutdinov, and Quoc V. Le · 2019
Closest in time.
Fixup initialization: Residual learning without normalization
Hongyi Zhang, Yann N Dauphin, and Tengyu Ma · 2019
Closest in time.
Bag-of-words as target for neural machine translation
Shuming Ma, Xu Sun, Yizhong Wang, and Junyang Lin · 2053
Closest in time.
Self-attention with relative position representations
Peter Shaw, Jakob Uszkoreit, and Ashish Vaswani · 2074
Closest in time.