Fetching the paper…
Reading the bibliography…
In NMT, how far can we get without attention and without separate encoding and decoding? To answer that question, we introduce a recurrent neural translation model that does not use attention and does not have a separate encoder and decoder.
A learning algorithm for continually running fully recurrent neural networks
Ronald J. Williams and David Zipser. 1989 · 1989
Earlier work this paper cites.
The mathematics of statistical machine translation: Parameter estimation
Peter E. Brown, Stephen A. Della Pietra, Vincent J. Della Pietra, and Robert L. Mercer. 1993 · 1993
Earlier work this paper cites.
Recursive hetero-associative memories for translation
Mikel L. Forcada and Ramón P. Ñeco. 1997 · 1997
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber. 1997 · 1997
Earlier work this paper cites.
Learning to forget: Continual prediction with lstm
Felix A. Gers, Jürgen Schmidhuber, and Fred Cummins. 2000 · 2000
Earlier work this paper cites.
A simple, fast, and effective reparameterization of IBM model 2
Chris Dyer, Victor Chahuneau, and Noah A. Smith. 2013 · 2013
Earlier work this paper cites.
Recurrent continuous translation models
Nal Kalchbrenner and Phil Blunsom. 2013 · 2013
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. 2014 · 2014
Earlier work this paper cites.
Learning phrase representations using rnn encoder–decoder for statistical machine translation
Kyunghyun Cho, Bart van Merrienboer, Caglar Gulcehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, and Yoshua Bengio. 2014 · 2014
Earlier work this paper cites.
Don’t until the final verb wait: Reinforcement learning for simultaneous machine translation
Alvin Grissom II, He He, Jordan Boyd-Graber, John Morgan, and Hal Daumé III. 2014 · 2014
Earlier work this paper cites.
Sequence to sequence learning with neural networks
Ilya Sutskever, Oriol Vinyals, and Quoc V Le. 2014 · 2014
Earlier work this paper cites.
Recurrent neural network regularization
Wojciech Zaremba, Ilya Sutskever, and Oriol Vinyals. 2014 · 2014
Earlier work this paper cites.
Morphological inflection generation with hard monotonic attention
Roee Aharoni and Yoav Goldberg. 2017 · 2015
Cited alongside, same era.
Nal Kalchbrenner, Ivo Danihelka, and Alex Graves. 2015 · 2015
Cited alongside, same era.
Effective approaches to attention-based neural machine translation
Thang Luong, Hieu Pham, and Christopher D. Manning. 2015 · 2015
Cited alongside, same era.
Tying word vectors and word classifiers: A loss framework for language modeling
Hakan Inan, Khashayar Khosravi, and Richard Socher. 2016 · 2016
Cited alongside, same era.
Neural machine translation in linear time
Nal Kalchbrenner, Lasse Espeholt, Karen Simonyan, Aaron van den Oord, Alex Graves, and Koray Kavukcuoglu. 2016 · 2016
Cited alongside, same era.
Six challenges for neural machine translation
Philipp Koehn and Rebecca Knowles. 2017 · 2017
Later among the works it cites.
Regularizing and Optimizing LSTM Language Models
Stephen Merity, Nitish Shirish Keskar, and Richard Socher. 2017 · 2017
Later among the works it cites.
Using the output embedding to improve language models
Ofir Press and Lior Wolf. 2017 · 2017
Later among the works it cites.
Online and linear-time attention by enforcing monotonic alignments
Colin Raffel, Minh-Thang Luong, Peter J Liu, Ron J Weiss, and Douglas Eck. 2017 · 2017
Later among the works it cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Neural machine translation of rare words with subword units
Rico Sennrich, Barry Haddow, and Alexandra Birch. 2016 · 2016
Cited alongside, same era.
Google’s neural machine translation system: Bridging the gap between human and machine translation
Yonghui Wu, Mike Schuster, Zhifeng Chen, Quoc V Le, Mohammad Norouzi, Wolfgang Macherey, Maxim Krikun, Yuan Cao, Qin Gao, Klaus Macherey, et al. 2016 · 2016
Cited alongside, same era.
Online segment to segment neural transduction
Lei Yu, Jan Buys, and Phil Blunsom. 2016 · 2016
Cited alongside, same era.
Convolutional sequence to sequence learning
Jonas Gehring, Michael Auli, David Grangier, Denis Yarats, and Yann N. Dauphin. 2017 · 2017
Cited alongside, same era.
Learning to translate in real-time with neural machine translation
Jiatao Gu, Graham Neubig, Kyunghyun Cho, and Victor O.K. Li. 2017 · 2017
Cited alongside, same era.
Neural phrase-based machine translation
Po-Sen Huang, Chong Wang, Dengyong Zhou, and Li Deng. 2017 · 2017
Cited alongside, same era.
OpenNMT: Open-source toolkit for neural machine translation
Guillaume Klein, Yoon Kim, Yuntian Deng, Jean Senellart, and Alexander M. Rush. 2017 · 2017
Cited alongside, same era.
Towards two-dimensional sequence to sequence model in neural machine translation
Parnia Bahar, Christopher Brix, and Hermann Ney. 2018 · 2018
Closest in time.
The best of both worlds: Combining recent advances in neural machine translation
Mia Xu Chen, Orhan Firat, Ankur Bapna, Melvin Johnson, Wolfgang Macherey, George Foster, Llion Jones, Niki Parmar, Mike Schuster, Zhifeng Chen, Yonghui Wu, and Macduff Hughes. 2018 · 2018
Closest in time.
How much attention do you need? a granular analysis of neural machine translation architectures
Tobias Domhan. 2018 · 2018
Closest in time.
Pervasive Attention: 2D Convolutional Neural Networks for Sequence-to-Sequence Prediction
Maha Elbayad, Laurent Besacier, and Jakob Verbeek. 2018 · 2018
Closest in time.
Stacl: Simultaneous translation with integrated anticipation and controllable latency
Mingbo Ma, Liang Huang, Hao Xiong, Kaibo Liu, Chuanqiang Zhang, Zhongjun He, Hairong Liu, Xing Li, and Haifeng Wang. 2018 · 2018
Closest in time.
A call for clarity in reporting BLEU scores
Matt Post. 2018 · 2018
Closest in time.
Neural hidden markov model for machine translation
Weiyue Wang, Derui Zhu, Tamer Alkhouli, Zixuan Gan, and Hermann Ney. 2018 · 2018
Closest in time.