Fetching the paper…
Reading the bibliography…
Although self-attention networks (SANs) have advanced the state-of-the-art on various NLP tasks, one criticism of SANs is their ability of encoding positions of input words (Shaw et al., 2018).
Dependency grammar and dependency parsing
Joakim Nivre. 2005 · 1959
Earlier work this paper cites.
Aspects of the Theory of Syntax , volume 11
Noam Chomsky. 1965 · 1965
Earlier work this paper cites.
The cognitive basis for linguistic structures
Thomas G Bever. 1970 · 1970
Earlier work this paper cites.
A non-projective dependency parser
Pasi Tapanainen and Timo Jarvinen. 1997 · 1997
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
Head-driven statistical models for natural language parsing
Michael Collins. 2003 · 2003
Earlier work this paper cites.
Accurate unlexicalized parsing
Dan Klein and Christopher D Manning. 2003 · 2003
Earlier work this paper cites.
Statistical significance tests for machine translation evaluation
Philipp Koehn. 2004 · 2004
Earlier work this paper cites.
Dependency parsing
Sandra Kübler, Ryan McDonald, and Joakim Nivre. 2009 · 2009
Earlier work this paper cites.
Abstract meaning representation for sembanking
Laura Banarescu, Claire Bonial, Shu Cai, Madalina Georgescu, Kira Griffitt, Ulf Hermjakob, Kevin Knight, Philipp Koehn, Martha Palmer, and Nathan Schneider. 2013 · 2013
Earlier work this paper cites.
Improved semantic representations from tree-structured long short-term memory networks
Kai Sheng Tai, Richard Socher, and Christopher D Manning. 2015 · 2015
Earlier work this paper cites.
Recurrent neural network grammars
Chris Dyer, Adhiguna Kuncoro, Miguel Ballesteros, and Noah A Smith. 2016 · 2016
Earlier work this paper cites.
A Decomposable Attention Model for Natural Language Inference
Ankur Parikh, Oscar Täckström, Dipanjan Das, and Jakob Uszkoreit. 2016 · 2016
Cited alongside, same era.
Learning to parse and translate improves neural machine translation
Akiko Eriguchi, Yoshimasa Tsuruoka, and Kyunghyun Cho. 2017 · 2017
Cited alongside, same era.
Structured attention networks
Yoon Kim, Carl Denton, Luong Hoang, and Alexander M Rush. 2017 · 2017
Cited alongside, same era.
A Structured Self-attentive Sentence Embedding
Zhouhan Lin, Minwei Feng, Cicero Nogueira dos Santos, Mo Yu, Bing Xiang, Bowen Zhou, and Yoshua Bengio. 2017 · 2017
Cited alongside, same era.
Attention is All You Need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
Neural machine translation advised by statistical machine translation
Xing Wang, Zhengdong Lu, Zhaopeng Tu, Hang Li, Deyi Xiong, and Min Zhang. 2017 · 2017
Phrase-level self-attention networks for universal sentence encoding
Wei Wu, Houfeng Wang, Tianyu Liu, and Shuming Ma. 2018 · 2018
Later among the works it cites.
Modeling localness for self-attention networks
Baosong Yang, Zhaopeng Tu, Derek F. Wong, Fandong Meng, Lidia S. Chao, and Tong Zhang. 2018 · 2018
Later among the works it cites.
Attention augmented convolutional networks
Irwan Bello, Barret Zoph, Ashish Vaswani, Jonathon Shlens, and Quoc V Le. 2019 · 2019
Closest in time.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Closest in time.
Dynamic layer aggregation for neural machine translation
Ziyi Dou, Zhaopeng Tu, Xing Wang, Longyue Wang, Shuming Shi, and Tong Zhang. 2019 · 2019
Closest in time.
A structural probe for finding syntax in word representations
John Hewitt and Christopher D. Manning. 2019 · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
THUMT: An Open Source Toolkit for Neural Machine Translation
Jiacheng Zhang, Yanzhuo Ding, Shiqi Shen, Yong Cheng, Maosong Sun, Huanbo Luan, and Yang Liu. 2017 · 2017
Cited alongside, same era.
What you can cram into a single vector: Probing sentence embeddings for linguistic properties
Alexis Conneau, Germán Kruszewski, Guillaume Lample, Loïc Barrault, and Marco Baroni. 2018 · 2018
Cited alongside, same era.
Exploiting deep representations for neural machine translation
Zi-Yi Dou, Zhaopeng Tu, Xing Wang, Shuming Shi, and Tong Zhang. 2018 · 2018
Cited alongside, same era.
Self-Attention with Relative Position Representations
Peter Shaw, Jakob Uszkoreit, and Ashish Vaswani. 2018 · 2018
Cited alongside, same era.
Linguistically-Informed Self-Attention for Semantic Role Labeling
Emma Strubell, Patrick Verga, Daniel Andor, David Weiss, and Andrew McCallum. 2018 · 2018
Cited alongside, same era.
Multi-layer representation fusion for neural machine translation
Qiang Wang, Fuxue Li, Tong Xiao, Yanyang Li, Yinqiao Li, and Jingbo Zhu. 2018 · 2018
Cited alongside, same era.
Closest in time.
Improving neural machine translation with neural syntactic distance
Chunpeng Ma, Akihiro Tamura, Masao Utiyama, Eiichiro Sumita, and Tiejun Zhao. 2019 · 2019
Closest in time.
Ordered neurons: Integrating tree structures into recurrent neural networks
Yikang Shen, Shawn Tan, Alessandro Sordoni, and Aaron Courville. 2019 · 2019
Closest in time.
Semantic neural machine translation using AMR
Linfeng Song, Daniel Gildea, Yue Zhang, Zhiguo Wang, and Jinsong Su. 2019 · 2019
Closest in time.
Adaptive attention span in transformers
Sainbayar Sukhbaatar, Edouard Grave, Piotr Bojanowski, and Armand Joulin. 2019 · 2019
Closest in time.
Exploiting sentential context for neural machine translation
Xing Wang, Zhaopeng Tu, Longyue Wang, and Shuming Shi. 2019 · 2019
Closest in time.