Fetching the paper…
Reading the bibliography…
Self-attention networks (SANs) have drawn increasing interest due to their high parallelization in computation and flexibility in modeling dependencies.
Statistical Significance Tests for Machine Translation Evaluation
Philipp Koehn. 2004 · 2004
Earlier work this paper cites.
Multimodal Deep Learning
Jiquan Ngiam, Aditya Khosla, Mingyu Kim, Juhan Nam, Honglak Lee, and Andrew Y Ng. 2011 · 2011
Earlier work this paper cites.
Convolutional Neural Networks for Sentence Classification
Yoon Kim. 2014 · 2014
Earlier work this paper cites.
Network in Network
Min Lin, Qiang Chen, and Shuicheng Yan. 2014 · 2014
Earlier work this paper cites.
Neural Machine Translation by Jointly Learning to Align and Translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. 2015 · 2015
Earlier work this paper cites.
Effective Approaches to Attention-based Neural Machine Translation
Thang Luong, Hieu Pham, and Christopher D. Manning. 2015 · 2015
Earlier work this paper cites.
Multimodal Compact Bilinear Pooling for Visual Question Answering and Visual Grounding
Akira Fukui, Dong Huk Park, Daylen Yang, Anna Rohrbach, Trevor Darrell, and Marcus Rohrbach. 2016 · 2016
Earlier work this paper cites.
A Decomposable Attention Model for Natural Language Inference
Ankur Parikh, Oscar Täckström, Dipanjan Das, and Jakob Uszkoreit. 2016 · 2016
Earlier work this paper cites.
Neural Machine Translation of Rare Words with Subword Units
Rico Sennrich, Barry Haddow, and Alexandra Birch. 2016 · 2016
Earlier work this paper cites.
A Structured Self-Aattentive Sentence Embedding
Zhouhan Lin, Minwei Feng, Cicero Nogueira dos Santos, Mo Yu, Bing Xiang, Bowen Zhou, and Yoshua Bengio. 2017 · 2017
Earlier work this paper cites.
NTT Neural Machine Translation Systems at WAT 2017
Makoto Morishita, Jun Suzuki, and Masaaki Nagata. 2017 · 2017
Cited alongside, same era.
Attention is All You Need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
Towards Bidirectional Hierarchical Representations for Attention-based Neural Machine Translation
Baosong Yang, Derek F Wong, Tong Xiao, Lidia S Chao, and Jingbo Zhu. 2017 · 2017
Cited alongside, same era.
The best of both worlds: Combining recent advances in neural machine translation
Mia Xu Chen, Orhan Firat, Ankur Bapna, Melvin Johnson, Wolfgang Macherey, George Foster, Llion Jones, Mike Schuster, Noam Shazeer, Niki Parmar, Ashish Vaswani, Jakob Uszkoreit, Lukasz Kaiser, Zhifeng Chen, Yonghui Wu, and Macduff Hughes. 2018 · 2018
Cited alongside, same era.
Exploiting Deep Representations for Neural Machine Translation
Ziyi Dou, Zhaopeng Tu, Xing Wang, Shuming Shi, and Tong Zhang. 2018 · 2018
Cited alongside, same era.
Linguistically-Informed Self-Attention for Semantic Role Labeling
Emma Strubell, Patrick Verga, Daniel Andor, David Weiss, and Andrew McCallum. 2018 · 2018
Later among the works it cites.
Phrase-level Self-Attention Networks for Universal Sentence Encoding
Wei Wu, Houfeng Wang, Tianyu Liu, and Shuming Ma. 2018 · 2018
Later among the works it cites.
Yuxin Wu and Kaiming He. 2018 · 2018
Later among the works it cites.
Modeling Localness for Self-Attention Networks
Baosong Yang, Zhaopeng Tu, Derek F. Wong, Fandong Meng, Lidia S. Chao, and Tong Zhang. 2018 · 2018
Later among the works it cites.
QANet: Combining Local Convolution with Global Self-attention for Reading Comprehension
Adams Wei Yu, David Dohan, Minh-Thang Luong, Rui Zhao, Kai Chen, Mohammad Norouzi, and Quoc V Le. 2018 · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Multi-Head Attention with Disagreement Regularization
Jian Li, Zhaopeng Tu, Baosong Yang, Michael R. Lyu, and Tong Zhang. 2018 · 2018
Cited alongside, same era.
Deep Contextualized Word Representations
Matthew E Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer. 2018 · 2018
Cited alongside, same era.
An Analysis of Encoder Representations in Transformer-Based Machine Translation
Alessandro Raganato and Jörg Tiedemann. 2018 · 2018
Cited alongside, same era.
Self-Attention with Relative Position Representations
Peter Shaw, Jakob Uszkoreit, and Ashish Vaswani. 2018 · 2018
Cited alongside, same era.
Self-Attentional Acoustic Models
Matthias Sperber, Jan Niehues, Graham Neubig, Sebastian Stüker, and Alex Waibel. 2018 · 2018
Cited alongside, same era.
DiSAN: Directional Self-Attention Network for RNN/CNN-Free Language Understanding
Tao Shen, Tianyi Zhou, Guodong Long, Jing Jiang, Shirui Pan, and Chengqi Zhang. 2018a
Cited in the paper.
Bi-Directional Block Self-Attention for Fast and Memory-Efficient Sequence Modeling
Tao Shen, Tianyi Zhou, Guodong Long, Jing Jiang, and Chengqi Zhang. 2018b
Cited in the paper.
Ziyi Dou, Zhaopeng Tu, Xing Wang, Longyue Wang, Shuming Shi, and Tong Zhang. 2019 · 2019
Closest in time.
Gaussian Transformer: A Lightweight Approach for Natural Language Inference
Maosheng Guo, Yu Zhang, and Ting Liu. 2019 · 2019
Closest in time.
Modeling Recurrence for Transformer
Jie Hao, Xing Wang, Baosong Yang, Longyue Wang, Jinfeng Zhang, and Zhaopeng Tu. 2019 · 2019
Closest in time.
Information Aggregation for Multi-Head Attention with Routing-by-Agreement
Jian Li, Baosong Yang, Zi-Yi Dou, Xing Wang, Michael R. Lyu, and Zhaopeng Tu. 2019 · 2019
Closest in time.
Context-Aware Self-Attention Networks
Baosong Yang, Jian Li, Derek F. Wong, Lidia S. Chao, Xing Wang, and Zhaopeng Tu. 2019 · 2019
Closest in time.