Fetching the paper…
Reading the bibliography…
Self-attention networks (SANs) with selective mechanism has produced substantial improvements in various NLP tasks by concentrating on a subset of input words.
Selective attention based graph convolutional networks for aspect-level sentiment classification
Xiaochen Hou, Jing Huang, Guangtao Wang, Kevin Huang, Xiaodong He, and Bowen Zhou. 2019 · 1910
Earlier work this paper cites.
Statistical theory of extreme values and some practical applications: a series of lectures
Emil Julius Gumbel. 1954 · 1954
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J. Williams. 1992 · 1992
Earlier work this paper cites.
The Kyoto free translation task
Graham Neubig. 2011 · 2011
Earlier work this paper cites.
Towards robust linguistic analysis using ontonotes
Sameer Pradhan, Alessandro Moschitti, Nianwen Xue, Hwee Tou Ng, Anders Björkelund, Olga Uryupina, Yuchen Zhang, and Zhi Zhong. 2013 · 2013
Earlier work this paper cites.
A* sampling
Chris J Maddison, Daniel Tarlow, and Tom Minka. 2014 · 2014
Earlier work this paper cites.
Glove: Global vectors for word representation
Jeffrey Pennington, Richard Socher, and Christopher D. Manning. 2014 · 2014
Earlier work this paper cites.
Neural Machine Translation by Jointly Learning to Align and Translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. 2015 · 2015
Earlier work this paper cites.
A large annotated corpus for learning natural language inference
Samuel R. Bowman, Gabor Angeli, Christopher Potts, and Christopher D. Manning. 2015 · 2015
Earlier work this paper cites.
Effective Approaches to Attention-based Neural Machine Translation
Thang Luong, Hieu Pham, and Christopher D. Manning. 2015 · 2015
Earlier work this paper cites.
English-chinese knowledge base translation with neural network
Xiaocheng Feng, Duyu Tang, Bing Qin, and Ting Liu. 2016 · 2016
Earlier work this paper cites.
Neural Machine Translation of Rare Words with Subword Units
Rico Sennrich, Barry Haddow, and Alexandra Birch. 2016 · 2016
Earlier work this paper cites.
Does String-based Neural MT Learn Source Syntax?
Xing Shi, Inkit Padhi, and Kevin Knight. 2016 · 2016
Cited alongside, same era.
Categorical reparameterization with gumbel-softmax
Eric Jang, Shixiang Gu, and Ben Poole. 2017 · 2017
Cited alongside, same era.
A structured self-attentive sentence embedding
Zhouhan Lin, Minwei Feng, Cicero Nogueira dos Santos, Mo Yu, Bing Xiang, Bowen Zhou, and Yoshua Bengio. 2017 · 2017
Cited alongside, same era.
Attention is All You Need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
The best of both worlds: Combining recent advances in neural machine translation
Mia Xu Chen, Orhan Firat, Ankur Bapna, Melvin Johnson, Wolfgang Macherey, George Foster, Llion Jones, Mike Schuster, Noam Shazeer, Niki Parmar, Ashish Vaswani, Jakob Uszkoreit, Lukasz Kaiser, Zhifeng Chen, Yonghui Wu, and Macduff Hughes. 2018 · 2018
Cited alongside, same era.
Self-Attention with Relative Position Representations
Peter Shaw, Jakob Uszkoreit, and Ashish Vaswani. 2018 · 2018
Later among the works it cites.
Double path networks for sequence to sequence learning
Kaitao Song, Xu Tan, Di He, Jianfeng Lu, Tao Qin, and Tie-Yan Liu. 2018 · 2018
Later among the works it cites.
Linguistically-informed self-attention for semantic role labeling
Emma Strubell, Patrick Verga, Daniel Andor, David Weiss, and Andrew McCallum. 2018 · 2018
Later among the works it cites.
Deep semantic role labeling with self-attention
Zhixing Tan, Mingxuan Wang, Jun Xie, Yidong Chen, and Xiaodong Shi. 2018 · 2018
Later among the works it cites.
Why Self-Attention? A Targeted Evaluation of Neural Machine Translation Architectures
Gongbo Tang, Mathias Müller, Annette Rios, and Rico Sennrich. 2018 · 2018
Later among the works it cites.
Phrase-level self-attention networks for universal sentence encoding
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Alexis Conneau, German Kruszewski, Guillaume Lample, Loïc Barrault, and Marco Baroni. 2018 · 2018
Cited alongside, same era.
Topic-to-essay generation with neural networks
Xiaocheng Feng, Ming Liu, Jiahao Liu, Bing Qin, Yibo Sun, and Ting Liu. 2018 · 2018
Cited alongside, same era.
Layer-wise coordination between encoder and decoder for neural machine translation
Tianyu He, Xu Tan, Yingce Xia, Di He, Tao Qin, Zhibo Chen, and Tie-Yan Liu. 2018 · 2018
Cited alongside, same era.
Multi-Head Attention with Disagreement Regularization
Jian Li, Zhaopeng Tu, Baosong Yang, Michael R. Lyu, and Tong Zhang. 2018 · 2018
Cited alongside, same era.
Extracting syntactic trees from transformer encoder self-attentions
David Mareček and Rudolf Rosa. 2018 · 2018
Cited alongside, same era.
Deep Contextualized Word Representations
Matthew E Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer. 2018 · 2018
Cited alongside, same era.
Multi-granularity self-attention for neural machine translation
Jie Hao, Xing Wang, Shuming Shi, Jinfeng Zhang, and Zhaopeng Tu. 2019a
Cited in the paper.
Wei Wu, Houfeng Wang, Tianyu Liu, and Shuming Ma. 2018 · 2018
Later among the works it cites.
Modeling localness for self-attention networks
Baosong Yang, Zhaopeng Tu, Derek F. Wong, Fandong Meng, Lidia S. Chao, and Tong Zhang. 2018 · 2018
Later among the works it cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Later among the works it cites.
Gaussian transformer: a lightweight approach for natural language inference
Maosheng Guo, Yu Zhang, and Ting Liu. 2019 · 2019
Later among the works it cites.
Towards understanding neural machine translation with word importance
Shilin He, Zhaopeng Tu, Xing Wang, Longyue Wang, Michael Lyu, and Shuming Shi. 2019 · 2019
Later among the works it cites.
Information aggregation for multi-head attention with routing-by-agreement
Jian Li, Baosong Yang, Zi-Yi Dou, Xing Wang, Michael R. Lyu, and Zhaopeng Tu. 2019 · 2019
Later among the works it cites.