Fetching the paper…
Reading the bibliography…
Multi-head self-attention is a key component of the Transformer, a state-of-the-art architecture for neural machine translation.
Auto-encoding variational bayes
Diederik P. Kingma and Max Welling. 2014 · 2014
Earlier work this paper cites.
The Stanford CoreNLP natural language processing toolkit
Christopher D. Manning, Mihai Surdeanu, John Bauer, Jenny Finkel, Steven J. Bethard, and David McClosky. 2014 · 2014
Earlier work this paper cites.
Stochastic backpropagation and approximate inference in deep generative models
Danilo Jimenez Rezende, Shakir Mohamed, and Daan Wierstra. 2014 · 2014
Earlier work this paper cites.
On pixel-wise explanations for non-linear classifier decisions by layer-wise relevance propagation
Sebastian Bach, Alexander Binder, Grégoire Montavon, Frederick Klauschen, Klaus-Robert Müller, and Wojciech Samek. 2015 · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik Kingma and Jimmy Ba. 2015 · 2015
Earlier work this paper cites.
To prune, or not to prune: exploring the efficacy of pruning for model compression
Michael Zhu and Suyog Gupta. 2017 · 2015
Earlier work this paper cites.
Assessing the ability of lstms to learn syntax-sensitive dependencies
Tal Linzen, Emmanuel Dupoux, and Yoav Goldberg. 2016 · 2016
Earlier work this paper cites.
Neural machine translation of rare words with subword units
Rico Sennrich, Barry Haddow, and Alexandra Birch. 2016 · 2016
Earlier work this paper cites.
Does string-based neural mt learn source syntax?
Xing Shi, Inkit Padhi, and Kevin Knight. 2016 · 2016
Earlier work this paper cites.
Understanding and improving morphological learning in the neural machine translation decoder
Fahim Dalvi, Nadir Durrani, Hassan Sajjad, Yonatan Belinkov, and Stephan Vogel. 2017 · 2017
Earlier work this paper cites.
Visualizing and understanding neural machine translation
Yanzhuo Ding, Yang Liu, Huanbo Luan, and Maosong Sun. 2017 · 2017
Earlier work this paper cites.
What does attention in neural machine translation pay attention to?
Hamidreza Ghader and Christof Monz. 2017 · 2017
Cited alongside, same era.
The representational geometry of word meanings acquired by neural machine translation models
Felix Hill, Kyunghyun Cho, Sébastien Jean, and Y Bengio. 2017 · 2017
Cited alongside, same era.
Categorical reparameterization with gumbel-softmax
Eric Jang, Shixiang Gu, and Ben Poole. 2017 · 2017
Cited alongside, same era.
The concrete distribution: A continuous relaxation of discrete random variables
Chris J. Maddison, Andriy Mnih, and Yee Whye Teh. 2017 · 2017
Cited alongside, same era.
How Grammatical is Character-level Neural Machine Translation? Assessing MT Quality with Contrastive Translation Pairs
Rico Sennrich. 2017 · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Learning sparse neural networks through l_0 regularization
Christos Louizos, Max Welling, and Diederik P. Kingma. 2018 · 2018
Later among the works it cites.
The IWSLT 2018 Evaluation Campaign
Jan Niehues, Ronaldo Cattoni, Sebastian Stüker, Mauro Cettolo, Marco Turchi, and Marcello Federico. 2018 · 2018
Later among the works it cites.
Training Tips for the Transformer Model
Martin Popel and Ondrej Bojar. 2018 · 2018
Later among the works it cites.
An analysis of encoder representations in transformer-based machine translation
Alessandro Raganato and Jörg Tiedemann. 2018 · 2018
Later among the works it cites.
Why self-attention? a targeted evaluation of neural machine translation architectures
Gongbo Tang, Mathias Müller, Annette Rios, and Rico Sennrich. 2018 · 2018
Later among the works it cites.
An analysis of attention mechanisms: The case of word sense disambiguation in neural machine translation
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
The lazy encoder: A fine-grained analysis of the role of morphology in neural machine translation
Arianna Bisazza and Clara Tump. 2018 · 2018
Cited alongside, same era.
Findings of the 2018 conference on machine translation (wmt18)
Ondřej Bojar, Christian Federmann, Mark Fishel, Yvette Graham, Barry Haddow, Matthias Huck, Philipp Koehn, and Christof Monz. 2018 · 2018
Cited alongside, same era.
Colorless green recurrent networks dream hierarchically
Kristina Gulordava, Piotr Bojanowski, Edouard Grave, Tal Linzen, and Marco Baroni. 2018 · 2018
Cited alongside, same era.
Discrete autoencoders for sequence models
Łukasz Kaiser and Samy Bengio. 2018 · 2018
Cited alongside, same era.
OpenSubtitles2018: Statistical Rescoring of Sentence Alignments in Large, Noisy Parallel Corpora
Pierre Lison, Jörg Tiedemann, and Milen Kouylekov. 2018 · 2018
Cited alongside, same era.
What do neural machine translation models learn about morphology?
Yonatan Belinkov, Nadir Durrani, Fahim Dalvi, Hassan Sajjad, and James Glass. 2017a
Cited in the paper.
Gongbo Tang, Rico Sennrich, and Joakim Nivre. 2018 · 2018
Later among the works it cites.
The importance of being recurrent for modeling hierarchical structure
Ke Tran, Arianna Bisazza, and Christof Monz. 2018 · 2018
Later among the works it cites.
Context-aware neural machine translation learns anaphora resolution
Elena Voita, Pavel Serdyukov, Rico Sennrich, and Ivan Titov. 2018 · 2018
Later among the works it cites.
Identifying and controlling important neurons in neural machine translation
Anthony Bau, Yonatan Belinkov, Hassan Sajjad, Nadir Durrani, Fahim Dalvi, and James Glass. 2019 · 2019
Closest in time.
The state of sparsity in deep neural networks
Trevor Gale, Erich Elsen, and Sara Hooker. 2019 · 2019
Closest in time.