Fetching the paper…
Reading the bibliography…
There have been significant efforts to interpret the encoder of Transformer-based encoder-decoder architectures for neural machine translation (NMT); meanwhile, the decoder remains largely unexamined despite its critical role.
Auto-association by multilayer perceptrons and singular value decomposition
Hervé Bourlard and Yves Kamp. 1988 · 1988
Earlier work this paper cites.
A systematic comparison of various statistical alignment models
Franz Josef Och and Hermann Ney. 2003 · 2003
Earlier work this paper cites.
Stacked denoising autoencoders: Learning useful representations in a deep network with a local denoising criterion
Pascal Vincent, Hugo Larochelle, Isabelle Lajoie, Yoshua Bengio, and Pierre-Antoine Manzagol. 2010 · 2010
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. 2015 · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba. 2015 · 2015
Earlier work this paper cites.
Contrastive unsupervised word alignment with non-local features
Yang Liu and Maosong Sun. 2015 · 2015
Earlier work this paper cites.
Grammar as a foreign language
Oriol Vinyals, Łukasz Kaiser, Terry Koo, Slav Petrov, Ilya Sutskever, and Geoffrey Hinton. 2015 · 2015
Earlier work this paper cites.
Layer normalization
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton. 2016 · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016 · 2016
Earlier work this paper cites.
Does string-based neural MT learn source syntax?
Xing Shi, Inkit Padhi, and Kevin Knight. 2016 · 2016
Earlier work this paper cites.
Modeling coverage for neural machine translation
Zhaopeng Tu, Zhengdong Lu, Yang Liu, Xiaohua Liu, and Hang Li. 2016 · 2016
Earlier work this paper cites.
What do neural machine translation models learn about morphology?
Yonatan Belinkov, Nadir Durrani, Fahim Dalvi, Hassan Sajjad, and James Glass. 2017 · 2017
Earlier work this paper cites.
Understanding and improving morphological learning in the neural machine translation decoder
Fahim Dalvi, Nadir Durrani, Hassan Sajjad, Yonatan Belinkov, and Stephan Vogel. 2017 · 2017
Earlier work this paper cites.
Convolutional sequence to sequence learning
Jonas Gehring, Michael Auli, David Grangier, Denis Yarats, and Yann N Dauphin. 2017 · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
Exploiting cross-sentence context for neural machine translation
Longyue Wang, Zhaopeng Tu, Andy Way, and Qun Liu. 2017 · 2017
Cited alongside, same era.
Yonatan Belinkov, Lluís Màrquez, Hassan Sajjad, Nadir Durrani, Fahim Dalvi, and James Glass. 2018 · 2018
Cited alongside, same era.
The lazy encoder: A fine-grained analysis of the role of morphology in neural machine translation
Arianna Bisazza and Clara Tump. 2018 · 2018
Cited alongside, same era.
Deep rnns encode soft hierarchical syntax
Terra Blevins, Omer Levy, and Luke Zettlemoyer. 2018 · 2018
Cited alongside, same era.
Towards understanding neural machine translation with word importance
Shilin He, Zhaopeng Tu, Xing Wang, Longyue Wang, Michael Lyu, and Shuming Shi. 2019 · 2019
Later among the works it cites.
Attention is not explanation
Sarthak Jain and Byron C. Wallace. 2019 · 2019
Later among the works it cites.
Neural machine translation with adequacy-oriented learning
Xiang Kong, Zhaopeng Tu, Shuming Shi, Eduard Hovy, and Tong Zhang. 2019 · 2019
Later among the works it cites.
On the word alignment from neural machine translation
Xintong Li, Guanlin Li, Lemao Liu, Max Meng, and Shuming Shi. 2019 · 2019
Later among the works it cites.
Are sixteen heads really better than one?
Paul Michel, Omer Levy, and Graham Neubig. 2019 · 2019
Later among the works it cites.
fairseq: A fast, extensible toolkit for sequence modeling
Myle Ott, Sergey Edunov, Alexei Baevski, Angela Fan, Sam Gross, Nathan Ng, David Grangier, and Michael Auli. 2019 · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
The best of both worlds: Combining recent advances in neural machine translation
Mia Xu Chen, Orhan Firat, Ankur Bapna, Melvin Johnson, Wolfgang Macherey, George Foster, Llion Jones, Niki Parmar, Mike Schuster, Zhifeng Chen, et al. 2018 · 2018
Cited alongside, same era.
What You Can Cram into A Single $ & ! # ∗ {\$}{\&}!{\#}* Vector: Probing Sentence Embeddings for Linguistic Properties
Alexis Conneau, German Kruszewski, Guillaume Lample, Loïc Barrault, and Marco Baroni. 2018 · 2018
Cited alongside, same era.
Exploiting deep representations for neural machine translation
Zi-Yi Dou, Zhaopeng Tu, Xing Wang, Shuming Shi, and Tong Zhang. 2018 · 2018
Cited alongside, same era.
Multi-head attention with disagreement regularization
Jian Li, Zhaopeng Tu, Baosong Yang, Michael R Lyu, and Tong Zhang. 2018 · 2018
Cited alongside, same era.
An analysis of encoder representations in transformer-based machine translation
Alessandro Raganato, Jörg Tiedemann, et al. 2018 · 2018
Cited alongside, same era.
DiSAN: Directional Self-Attention Network for RNN/CNN-Free Language Understanding
Tao Shen, Tianyi Zhou, Guodong Long, Jing Jiang, Shirui Pan, and Chengqi Zhang. 2018 · 2018
Cited alongside, same era.
Linguistically-informed self-attention for semantic role labeling
Emma Strubell, Patrick Verga, Daniel Andor, David Weiss, and Andrew McCallum. 2018 · 2018
Cited alongside, same era.
Later among the works it cites.
Language models are unsupervised multitask learners
Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019 · 2019
Later among the works it cites.
Analyzing multi-head self-attention: Specialized heads do the heavy lifting, the rest can be pruned
Elena Voita, David Talbot, Fedor Moiseev, Rico Sennrich, and Ivan Titov. 2019 · 2019
Later among the works it cites.
One model to learn both: Zero pronoun prediction and translation
Longyue Wang, Zhaopeng Tu, Xing Wang, and Shuming Shi. 2019 · 2019
Later among the works it cites.
Context-Aware Self-Attention Networks
Baosong Yang, Jian Li, Derek Wong, Lidia S. Chao, Xing Wang, and Zhaopeng Tu. 2019 · 2019
Later among the works it cites.
Assessing the ability of self-attention networks to learn word order
Baosong Yang, Longyue Wang, Derek F. Wong, Lidia S. Chao, and Zhaopeng Tu. 2019 · 2019
Later among the works it cites.
Improving deep transformer with depth-scaled initialization and merged attention
Biao Zhang, Ivan Titov, and Rico Sennrich. 2019 · 2019
Later among the works it cites.
How Does Selective Mechanism Improve Self-Attention Networks?
Xinwei Geng, Longyue Wang, Xing Wang, Bing Qin, Ting Liu, and Zhaopeng Tu. 2020 · 2020
Closest in time.
On the Inference Calibration of Neural Machine Translation
Shuo Wang, Zhaopeng Tu, Shuming Shi, and Yang Liu. 2020 · 2020
Closest in time.