Fetching the paper…
Reading the bibliography…
Current state-of-the-art neural machine translation (NMT) uses a deep multi-head self-attention network with no explicit phrase information.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
Accurate unlexicalized parsing
Dan Klein and Christopher D Manning. 2003 · 2003
Earlier work this paper cites.
Statistical phrase-based translation
Philipp Koehn, Franz Josef Och, and Daniel Marcu. 2003 · 2003
Earlier work this paper cites.
A hierarchical phrase-based model for statistical machine translation
David Chiang. 2005 · 2005
Earlier work this paper cites.
Tree-to-string alignment template for statistical machine translation
Yang Liu, Qun Liu, and Shouxun Lin. 2006 · 2006
Earlier work this paper cites.
Convolutional neural network for paraphrase identification
Wenpeng Yin and Hinrich Schütze. 2015 · 2015
Earlier work this paper cites.
Yukun Zhu, Ryan Kiros, Rich Zemel, Ruslan Salakhutdinov, Raquel Urtasun, Antonio Torralba, and Sanja Fidler. 2015 · 2015
Earlier work this paper cites.
Tree-to-sequence attentional neural machine translation
Akiko Eriguchi, Kazuma Hashimoto, and Yoshimasa Tsuruoka. 2016 · 2016
Earlier work this paper cites.
Neural machine translation of rare words with subword units
Rico Sennrich, Barry Haddow, and Alexandra Birch. 2016 · 2016
Earlier work this paper cites.
Does string-based neural mt learn source syntax
Xing Shi, Inkit Padhi, and Kevin Knight. 2016 · 2016
Earlier work this paper cites.
Multi-granularity chinese word embedding
Rongchao Yin, Quan Wang, Peng Li, Rui Li, and Bin Wang. 2016 · 2016
Earlier work this paper cites.
Graph convolutional encoders for syntax-aware neural machine translation
Joost Bastings, Ivan Titov, Wilker Aziz, Diego Marcheggiani, and Khalil Simaan. 2017 · 2017
Cited alongside, same era.
Attention is All You Need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
Translating phrases in neural machine translation
Xing Wang, Zhaopeng Tu, Deyi Xiong, and Min Zhang. 2017 · 2017
Cited alongside, same era.
What you can cram into a single $ & ! # ∗ {\$}{\&}!{\#}* vector: Probing sentence embeddings for linguistic properties
Alexis Conneau, German Kruszewski, Guillaume Lample, Loïc Barrault, and Marco Baroni. 2018 · 2018
Cited alongside, same era.
Multi-Head Attention with Disagreement Regularization
Jian Li, Zhaopeng Tu, Baosong Yang, Michael R. Lyu, and Tong Zhang. 2018 · 2018
Cited alongside, same era.
The importance of being recurrent for modeling hierarchical structure
Ke Tran, Arianna Bisazza, and Christof Monz. 2018 · 2018
Later among the works it cites.
Phrase-level self-attention networks for universal sentence encoding
Wei Wu, Houfeng Wang, Tianyu Liu, and Shuming Ma. 2018 · 2018
Later among the works it cites.
Modeling localness for self-attention networks
Baosong Yang, Zhaopeng Tu, Derek F. Wong, Fandong Meng, Lidia S. Chao, and Tong Zhang. 2018 · 2018
Later among the works it cites.
Phrase table as recommendation memory for neural machine translation
Yang Zhao, Yining Wang, Jiajun Zhang, and Chengqing Zong. 2018 · 2018
Later among the works it cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Closest in time.
Strategies for structuring story generation
Angela Fan, Mike Lewis, and Yann Dauphin. 2019 · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Phi Xuan Nguyen and Shafiq Joty. 2018 · 2018
Cited alongside, same era.
Deep contextualized word representations
Matthew E Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer. 2018 · 2018
Cited alongside, same era.
An analysis of encoder representations in transformer-based machine translation
Alessandro Raganato and Jörg Tiedemann. 2018 · 2018
Cited alongside, same era.
Bi-directional block self-attention for fast and memory-efficient sequence modeling
Tao Shen, Tianyi Zhou, Guodong Long, Jing Jiang, and Chengqi Zhang. 2018 · 2018
Cited alongside, same era.
Linguistically-Informed Self-Attention for Semantic Role Labeling
Emma Strubell, Patrick Verga, Daniel Andor, David Weiss, and Andrew McCallum. 2018 · 2018
Cited alongside, same era.
Get the point of my utterance! learning towards effective responses with multi-head attention mechanism
Chongyang Tao, Shen Gao, Mingyue Shang, Wei Wu, Dongyan Zhao, and Rui Yan. 2018 · 2018
Cited alongside, same era.
Towards better modeling hierarchical structure for self-attention with ordered neurons
Jie Hao, Xing Wang, Shuming Shi, Jinfeng Zhang, and Zhaopeng Tu. 2019a
Cited in the paper.
Closest in time.
Gaussian transformer: a lightweight approach for natural language inference
Maosheng Guo, Yu Zhang, and Ting Liu. 2019 · 2019
Closest in time.
Analyzing multi-head self-attention: Specialized heads do the heavy lifting, the rest can be pruned
Elena Voita, David Talbot, Fedor Moiseev, Rico Sennrich, and Ivan Titov. 2019 · 2019
Closest in time.
Convolutional self-attention networks
Baosong Yang, Longyue Wang, Derek Wong, Lidia S Chao, and Zhaopeng Tu. 2019 · 2019
Closest in time.
Syntax-enhanced neural machine translation with syntax-aware word representations
Meishan Zhang, Zhenghua Li, Guohong Fu, and Min Zhang. 2019 · 2019
Closest in time.