Fetching the paper…
Reading the bibliography…
Self-attention networks have proven to be of profound value for its strength of capturing global dependencies.
BLEU: A Method for Automatic Evaluation of Machine Translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
Statistical Significance Tests for Machine Translation Evaluation
Philipp Koehn. 2004 · 2004
Earlier work this paper cites.
Neural Machine Translation by Jointly Learning to Align and Translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. 2015 · 2015
Earlier work this paper cites.
Effective Approaches to Attention-based Neural Machine Translation
Thang Luong, Hieu Pham, and Christopher D. Manning. 2015 · 2015
Earlier work this paper cites.
Show, Attend and Tell: Neural Image Caption Generation with Visual Attention
Kelvin Xu, Jimmy Ba, Ryan Kiros, Kyunghyun Cho, Aaron C Courville, Ruslan Salakhutdinov, Richard S Zemel, and Yoshua Bengio. 2015 · 2015
Earlier work this paper cites.
Neural Machine Translation of Rare Words with Subword Units
Rico Sennrich, Barry Haddow, and Alexandra Birch. 2016 · 2016
Earlier work this paper cites.
Does String-based Neural MT Learn Source Syntax?
Xing Shi, Inkit Padhi, and Kevin Knight. 2016 · 2016
Earlier work this paper cites.
Google’s Neural Machine Translation System: Bridging the Gap Between Human and Machine Translation
Yonghui Wu, Mike Schuster, Zhifeng Chen, Quoc V Le, Mohammad Norouzi, Wolfgang Macherey, Maxim Krikun, Yuan Cao, Qin Gao, Klaus Macherey, et al. 2016 · 2016
Earlier work this paper cites.
Frustratingly short attention spans in neural language modeling
Michał Daniluk, Tim Rocktäschel, Johannes Welbl, and Sebastian Riedel. 2017 · 2017
Earlier work this paper cites.
Convolutional Sequence to Sequence Learning
Jonas Gehring, Michael Auli, David Grangier, Denis Yarats, and Yann N. Dauphin. 2017 · 2017
Cited alongside, same era.
Distance-based Self-Attention Network for Natural Language Inference
Jinbae Im and Sungzoon Cho. 2017 · 2017
Cited alongside, same era.
Attention is All You Need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
Translating Phrases in Neural Machine Translation
Xing Wang, Zhaopeng Tu, Deyi Xiong, and Min Zhang. 2017 · 2017
Cited alongside, same era.
Towards Bidirectional Hierarchical Representations for Attention-based Neural Machine Translation
Baosong Yang, Derek F Wong, Tong Xiao, Lidia S Chao, and Jingbo Zhu. 2017 · 2017
Cited alongside, same era.
THUMT: An Open Source Toolkit for Neural Machine Translation
Achieving Human Parity on Automatic Chinese to English News Translation
Hany Hassan, Anthony Aue, Chang Chen, Vishal Chowdhary, Jonathan Clark, Christian Federmann, Xuedong Huang, Marcin Junczys-Dowmunt, William Lewis, Mu Li, et al. 2018 · 2018
Closest in time.
Towards Neural Phrase-based Machine Translation
Po Sen Huang, Chong Wang, Sitao Huang, Dengyong Zhou, and Li Deng. 2018 · 2018
Closest in time.
Multi-Head Attention with Disagreement Regularization
Jian Li, Zhaopeng Tu, Baosong Yang, Michael R. Lyu, and Tong Zhang. 2018 · 2018
Closest in time.
Deep Contextualized Word Representations
Matthew E. Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer. 2018 · 2018
Closest in time.
Self-Attention with Relative Position Representations
Peter Shaw, Jakob Uszkoreit, and Ashish Vaswani. 2018 · 2018
Closest in time.
Self-Attentional Acoustic Models
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jiacheng Zhang, Yanzhuo Ding, Shiqi Shen, Yong Cheng, Maosong Sun, Huanbo Luan, and Yang Liu. 2017 · 2017
Cited alongside, same era.
Tied Multitask Learning for Neural Speech Translation
Antonios Anastasopoulos and David Chiang. 2018 · 2018
Cited alongside, same era.
Exploiting Deep Representations for Neural Machine Translation
Ziyi Dou, Zhaopeng Tu, Xing Wang, Shuming Shi, and Tong Zhang. 2018 · 2018
Cited alongside, same era.
DiSAN: Directional Self-Attention Network for RNN/CNN-Free Language Understanding
Tao Shen, Tianyi Zhou, Guodong Long, Jing Jiang, Shirui Pan, and Chengqi Zhang. 2018a
Cited in the paper.
Bi-Directional Block Self-Attention for Fast and Memory-Efficient Sequence Modeling
Tao Shen, Tianyi Zhou, Guodong Long, Jing Jiang, and Chengqi Zhang. 2018b
Cited in the paper.
Matthias Sperber, Jan Niehues, Graham Neubig, Sebastian Stüker, and Alex Waibel. 2018 · 2018
Closest in time.
Cross-Target Stance Classification with Self-Attention Networks
Chang Xu, Cecile Paris, Surya Nepal, and Ross Sparks. 2018 · 2018
Closest in time.
Accelerating Neural Transformer via an Average Attention Network
Biao Zhang, Deyi Xiong, and Jinsong Su. 2018 · 2018
Closest in time.