Fetching the paper…
Reading the bibliography…
Multi-head attention is appealing for the ability to jointly attend to information from different representation subspaces at different positions.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
Statistical phrase-based translation
Philipp Koehn, Franz Josef Och, and Daniel Marcu. 2003 · 2003
Earlier work this paper cites.
Clause restructuring for statistical machine translation
M. Collins, P. Koehn, and I. Kučerová. 2005 · 2005
Earlier work this paper cites.
Alignment by agreement
Percy Liang, Ben Taskar, and Dan Klein. 2006 · 2006
Earlier work this paper cites.
Agreement-based Learning
Percy Liang, Dan Klein, and Michael I Jordan. 2008 · 2008
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. 2015 · 2015
Earlier work this paper cites.
Attention-based Models for Speech Recognition
Jan K Chorowski, Dzmitry Bahdanau, Dmitriy Serdyuk, Kyunghyun Cho, and Yoshua Bengio. 2015 · 2015
Earlier work this paper cites.
Model invertibility regularization: Sequence alignment with or without parallel data
Tomer Levinboim, Ashish Vaswani, and David Chiang. 2015 · 2015
Earlier work this paper cites.
Effective Approaches to Attention-based Neural Machine Translation
Minh-Thang Luong, Hieu Pham, and Christopher D Manning. 2015 · 2015
Cited alongside, same era.
Show, Attend and Tell: Neural Image Caption Generation with Visual Attention
Kelvin Xu, Jimmy Ba, Ryan Kiros, Kyunghyun Cho, Aaron C Courville, Ruslan Salakhutdinov, Richard S Zemel, and Yoshua Bengio. 2015 · 2015
Cited alongside, same era.
Agreement-based joint training for bidirectional attention-based neural machine translation
Yong Cheng, Shiqi Shen, Zhongjun He, Wei He, Hua Wu, Maosong Sun, and Yang Liu. 2016 · 2016
Cited alongside, same era.
Neural machine translation of rare words with subword units
Rico Sennrich, Barry Haddow, and Alexandra Birch. 2016 · 2016
Cited alongside, same era.
Modeling coverage for neural machine translation
Zhaopeng Tu, Zhengdong Lu, Yang Liu, Xiaohua Liu, and Hang Li. 2016 · 2016
Cited alongside, same era.
Google’s neural machine translation system: Bridging the gap between human and machine translation
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Later among the works it cites.
THUMT: An Open Source Toolkit for Neural Machine Translation
Jiacheng Zhang, Yanzhuo Ding, Shiqi Shen, Yong Cheng, Maosong Sun, Huanbo Luan, and Yang Liu. 2017 · 2017
Later among the works it cites.
Exploiting Deep Representations for Neural Machine Translation
Ziyi Dou, Zhaopeng Tu, Xing Wang, Shuming Shi, and Tong Zhang. 2018 · 2018
Closest in time.
Achieving human parity on automatic chinese to english news translation
Hany Hassan, Anthony Aue, Chang Chen, Vishal Chowdhary, Jonathan Clark, Christian Federmann, Xuedong Huang, Marcin Junczys-Dowmunt, William Lewis, Mu Li, et al. 2018 · 2018
Closest in time.
Self-Attention with Relative Position Representations
Peter Shaw, Jakob Uszkoreit, and Ashish Vaswani. 2018 · 2018
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yonghui Wu, Mike Schuster, Zhifeng Chen, Quoc V Le, Mohammad Norouzi, Wolfgang Macherey, Maxim Krikun, Yuan Cao, Qin Gao, Klaus Macherey, et al. 2016 · 2016
Cited alongside, same era.
Weighted transformer network for machine translation
Karim Ahmed, Nitish Shirish Keskar, and Richard Socher. 2017 · 2017
Cited alongside, same era.
Convolutional sequence to sequence learning
Jonas Gehring, Michael Auli, David Grangier, Denis Yarats, and Yann N. Dauphin. 2017 · 2017
Cited alongside, same era.
DiSAN: Directional Self-Attention Network for RNN/CNN-free Language Understanding
Tao Shen, Tianyi Zhou, Guodong Long, Jing Jiang, Shirui Pan, and Chengqi Zhang. 2018 · 2018
Closest in time.
Modeling Localness for Self-Attention Networks
Baosong Yang, Zhaopeng Tu, Derek F. Wong, Fandong Meng, Lidia S. Chao, and Tong Zhang. 2018 · 2018
Closest in time.