Fetching the paper…
Reading the bibliography…
Multi-layer models with multiple attention heads per layer provide superior translation quality compared to simpler and shallower models, but determining what source context is most relevant to each target word is more challenging as a result.
The mathematics of statistical machine translation: Parameter estimation
Peter F Brown, Vincent J Della Pietra, Stephen A Della Pietra, and Robert L Mercer. 1993 · 1993
Earlier work this paper cites.
Improved statistical alignment models
Franz Josef Och and Hermann Ney. 2000 · 2000
Earlier work this paper cites.
An evaluation exercise for word alignment
Rada Mihalcea and Ted Pedersen. 2003 · 2003
Earlier work this paper cites.
A systematic comparison of various statistical alignment models
Franz Josef Och and Hermann Ney. 2003 · 2003
Earlier work this paper cites.
Edinburgh system description for the 2005 iwslt speech translation evaluation
Philipp Koehn, Amittai Axelrod, Alexandra Birch Mayne, Chris Callison-Burch, Miles Osborne, and David Talbot. 2005 · 2005
Earlier work this paper cites.
Parallel implementations of word alignment tool
Qin Gao and Stephan Vogel. 2008 · 2008
Earlier work this paper cites.
Visualizing higher-layer features of a deep network
Dumitru Erhan, Yoshua Bengio, Aaron Courville, and Pascal Vincent. 2009 · 2009
Earlier work this paper cites.
A simple, fast, and effective reparameterization of ibm model 2
Chris Dyer, Victor Chahuneau, and Noah A Smith. 2013 · 2013
Cited alongside, same era.
Sequence to sequence learning with neural networks
Ilya Sutskever, Oriol Vinyals, and Quoc V Le. 2014 · 2014
Cited alongside, same era.
Recurrent neural networks for word alignment model
Akihiro Tamura, Taro Watanabe, and Eiichiro Sumita. 2014 · 2014
Cited alongside, same era.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. 2015 · 2015
Cited alongside, same era.
Incorporating discrete translation lexicons into neural machine translation
Philip Arthur, Graham Neubig, and Satoshi Nakamura. 2016 · 2016
Cited alongside, same era.
Neural machine translation of rare words with subword units
Rico Sennrich, Barry Haddow, and Alexandra Birch. 2016 · 2016
Feature visualization
Chris Olah, Alexander Mordvintsev, and Ludwig Schubert. 2017 · 2017
Later among the works it cites.
Generating alignments using target foresight in attention-based neural machine translation
Jan-Thorsten Peter, Arne Nix, and Hermann Ney. 2017 · 2017
Later among the works it cites.
Visualizing Neural Machine Translation Attention and Confidence
Matiss Rikters, Mark Fishel, and Ondřej Bojar. 2017 · 2017
Later among the works it cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Later among the works it cites.
On the alignment problem in multi-head attention-based neural machine translation
Tamer Alkhouli, Gabriel Bretschner, and Hermann Ney. 2018 · 2018
Later among the works it cites.
Improving lexical choice in neural machine translation
Toan Nguyen and David Chiang. 2018 · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Later among the works it cites.