Fetching the paper…
Reading the bibliography…
Multi-head attention is appealing for its ability to jointly extract different types of information from multiple representation subspaces.
BLEU: A method for Automatic Evaluation of Machine Translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
Statistical Significance Tests for Machine Translation Evaluation
Philipp Koehn. 2004 · 2004
Earlier work this paper cites.
Transforming Auto-encoders
Geoffrey E Hinton, Alex Krizhevsky, and Sida D Wang. 2011 · 2011
Earlier work this paper cites.
Neural Machine Translation by Jointly Learning to Align and Translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. 2015 · 2015
Earlier work this paper cites.
Attention-based Models for Speech Recognition
Jan K Chorowski, Dzmitry Bahdanau, Dmitriy Serdyuk, Kyunghyun Cho, and Yoshua Bengio. 2015 · 2015
Earlier work this paper cites.
Effective Approaches to Attention-based Neural Machine Translation
Minh-Thang Luong, Hieu Pham, and Christopher D Manning. 2015 · 2015
Earlier work this paper cites.
Show, Attend and Tell: Neural Image Caption Generation with Visual Attention
Kelvin Xu, Jimmy Ba, Ryan Kiros, Kyunghyun Cho, Aaron C Courville, Ruslan Salakhutdinov, Richard S Zemel, and Yoshua Bengio. 2015 · 2015
Earlier work this paper cites.
Multimodal Compact Bilinear Pooling for Visual Question Answering and Visual Grounding
Akira Fukui, Dong Huk Park, Daylen Yang, Anna Rohrbach, Trevor Darrell, and Marcus Rohrbach. 2016 · 2016
Earlier work this paper cites.
Neural Machine Translation of Rare Words with Subword Units
Rico Sennrich, Barry Haddow, and Alexandra Birch. 2016 · 2016
Earlier work this paper cites.
Does String-based Neural MT Learn Source Syntax?
Xing Shi, Inkit Padhi, and Kevin Knight. 2016 · 2016
Earlier work this paper cites.
Google’s Neural Machine Translation System: Bridging the Gap between Human and Machine Translation
Yonghui Wu, Mike Schuster, Zhifeng Chen, Quoc V Le, Mohammad Norouzi, Wolfgang Macherey, Maxim Krikun, Yuan Cao, Qin Gao, Klaus Macherey, et al. 2016 · 2016
Earlier work this paper cites.
Mutan: Multimodal Tucker Fusion for Visual Question Answering
Hedi Ben-Younes, Rémi Cadene, Matthieu Cord, and Nicolas Thome. 2017 · 2017
Earlier work this paper cites.
Convolutional Sequence to Sequence Learning
Jonas Gehring, Michael Auli, David Grangier, Denis Yarats, and Yann N. Dauphin. 2017 · 2017
Cited alongside, same era.
A Structured Self-attentive Sentence Embedding
Zhouhan Lin, Minwei Feng, Cicero Nogueira dos Santos, Mo Yu, Bing Xiang, Bowen Zhou, and Yoshua Bengio. 2017 · 2017
Cited alongside, same era.
Dynamic Routing Between Capsules
Sara Sabour, Nicholas Frosst, and Geoffrey E Hinton. 2017 · 2017
Cited alongside, same era.
Attention Is All You Need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
Capsule network performance on complex data
Edgar Xi, Selina Bing, and Yang Jin. 2017 · 2017
Cited alongside, same era.
Achieving Human Parity on Automatic Chinese to English News Translation
Hany Hassan, Anthony Aue, Chang Chen, Vishal Chowdhary, Jonathan Clark, Christian Federmann, Xuedong Huang, Marcin Junczys-Dowmunt, William Lewis, Mu Li, et al. 2018 · 2018
Later among the works it cites.
Matrix Capsules with EM Routing
Geoffrey E Hinton, Sara Sabour, and Nicholas Frosst. 2018 · 2018
Later among the works it cites.
Capsules for object segmentation
Rodney LaLonde and Ulas Bagci. 2018 · 2018
Later among the works it cites.
Multi-Head Attention with Disagreement Regularization
Jian Li, Zhaopeng Tu, Baosong Yang, Michael R. Lyu, and Tong Zhang. 2018 · 2018
Later among the works it cites.
Deep Contextualized Word Representations
Matthew E. Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer. 2018 · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Karim Ahmed, Nitish Shirish Keskar, and Richard Socher. 2018 · 2018
Cited alongside, same era.
What You Can Cram into A Single $ & ! # ∗ {\$}{\&}!{\#}* Vector: Probing Sentence Embeddings for Linguistic Properties
Alexis Conneau, German Kruszewski, Guillaume Lample, Loïc Barrault, and Marco Baroni. 2018 · 2018
Cited alongside, same era.
Mostafa Dehghani, Stephan Gouws, Oriol Vinyals, Jakob Uszkoreit, and Lukasz Kaiser. 2018 · 2018
Cited alongside, same era.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018 · 2018
Cited alongside, same era.
How Much Attention Do You Need? A Granular Analysis of Neural Machine Translation Architectures
Tobias Domhan. 2018 · 2018
Cited alongside, same era.
Exploiting deep representations for neural machine translation
Ziyi Dou, Zhaopeng Tu, Xing Wang, Shuming Shi, and Tong Zhang. 2018 · 2018
Cited alongside, same era.
Information Aggregation via Dynamic Routing for Sequence Encoding
Jingjing Gong, Xipeng Qiu, Shaojing Wang, and Xuanjing Huang. 2018 · 2018
Cited alongside, same era.
An Analysis of Encoder Representations in Transformer-Based Machine Translation
Alessandro Raganato and Jörg Tiedemann. 2018 · 2018
Later among the works it cites.
DiSAN: Directional Self-Attention Network for RNN/CNN-free Language Understanding
Tao Shen, Tianyi Zhou, Guodong Long, Jing Jiang, Shirui Pan, and Chengqi Zhang. 2018 · 2018
Later among the works it cites.
Linguistically-Informed Self-Attention for Semantic Role Labeling
Emma Strubell, Patrick Verga, Daniel Andor, David Weiss, and Andrew McCallum. 2018 · 2018
Later among the works it cites.
Why Self-Attention? A Targeted Evaluation of Neural Machine Translation Architectures
Gongbo Tang, Mathias Müller, Annette Rios, and Rico Sennrich. 2018 · 2018
Later among the works it cites.
Investigating Capsule Networks with Dynamic Routing for Text Classification
Wei Zhao, Jianbo Ye, Min Yang, Zeyang Lei, Suofei Zhang, and Zhou Zhao. 2018 · 2018
Later among the works it cites.
Dynamic layer aggregation for neural machine translation
Ziyi Dou, Zhaopeng Tu, Xing Wang, Longyue Wang, Shuming Shi, and Tong Zhang. 2019 · 2019
Closest in time.
Convolutional self-attention networks
Baosong Yang, Longyue Wang, Derek F. Wong, Lidia S. Chao, and Zhaopeng Tu. 2019 · 2019
Closest in time.