Fetching the paper…
Reading the bibliography…
Current state-of-the-art machine translation systems are based on encoder-decoder architectures, that first encode the input sequence, and then generate an output sequence based on the input encoding.
Long short-term memory
S. Hochreiter and J. Schmidhuber. 1997 · 1997
Earlier work this paper cites.
Bidirectional recurrent neural networks
M. Schuster and K. Paliwal. 1997 · 1997
Earlier work this paper cites.
BLEU: a method for automatic evaluation of machine translation
K. Papineni, S. Roukos, T. Ward, and W.-J. Zhu. 2002 · 2002
Earlier work this paper cites.
Moses: Open source toolkit for statistical machine translation
P. Koehn, H. Hoang, A. Birch, C. Callison-Burch, M. Federico, N. Bertoldi, B. Cowan, W. Shen, C. Moran, R. Zens, C. Dyer, O. Bojar, A. Constantin, and E. Herbst. 2007 · 2007
Earlier work this paper cites.
A unified architecture for natural language processing: Deep neural networks with multitask learning
R. Collobert and J. Weston. 2008 · 2008
Earlier work this paper cites.
Rectified linear units improve restricted Boltzmann machines
V. Nair and G. Hinton. 2010 · 2010
Earlier work this paper cites.
Sequence transduction with recurrent neural networks
A. Graves. 2012 · 2012
Earlier work this paper cites.
Recurrent continuous translation models
N. Kalchbrenner and P. Blunsom. 2013 · 2013
Earlier work this paper cites.
Report on the 11th IWSLT evaluation campaign
M. Cettolo, J. Niehues, S. Stüker, L. Bentivogli, and M. Federico. 2014 · 2014
Earlier work this paper cites.
Learning phrase representations using RNN encoder-decoder for statistical machine translation
K. Cho, B. van Merrienboer, Ç. Gülçehre, D. Bahdanau, F. Bougares, H. Schwenk, and Y. Bengio. 2014 · 2014
Earlier work this paper cites.
A convolutional neural network for modelling sentences
N. Kalchbrenner, E. Grefenstette, and P. Blunsom. 2014 · 2014
Earlier work this paper cites.
Convolutional neural networks for sentence classification
Y. Kim. 2014 · 2014
Earlier work this paper cites.
Dropout: A simple way to prevent neural networks from overfitting
N. Srivastava, G. Hinton, A. Krizhevsky, I. Sutskever, and R. Salakhutdinov. 2014 · 2014
Earlier work this paper cites.
Sequence to sequence learning with neural networks
I. Sutskever, O. Vinyals, and Q. Le. 2014 · 2014
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
D. Bahdanau, K. Cho, and Y. Bengio. 2015 · 2015
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
S. Ioffe and C. Szegedy. 2015 · 2015
Cited alongside, same era.
On using very large target vocabulary for neural machine translation
S. Jean, K. Cho, R. Memisevic, and Y. Bengio. 2015 · 2015
Cited alongside, same era.
Adam: A method for stochastic optimization
D. Kingma and J. Ba. 2015 · 2015
Cited alongside, same era.
Deep learning
Y. LeCun, Y. Bengio, and G. Hinton. 2015 · 2015
Cited alongside, same era.
Effective approaches to attention-based neural machine translation
T. Luong, H. Pham, and C. Manning. 2015 · 2015
Cited alongside, same era.
Encoding source language with convolutional neural network for machine translation
F. Meng, Z. Lu, M. Wang, H. Li, W. Jiang, and Q. Liu. 2015 · 2015
Cited alongside, same era.
Densely connected convolutional networks
G. Huang, Z. Liu, L. van der Maaten, and K. Weinberger. 2017 · 2017
Later among the works it cites.
Learning to align the source code to the compiled object code
D. Levy and L. Wolf. 2017 · 2017
Later among the works it cites.
A structured self-attentive sentence embedding
Z. Lin, M. Feng, C. dos Santos, M. Yu, B. Xiang, B. Zhou, and Y. Bengio. 2017 · 2017
Later among the works it cites.
Automatic differentiation in pytorch
A. Paszke, S. Gross, S. Chintala, G. Chanan, E. Yang, Z. DeVito, Z. Lin, A. Desmaison L. Antiga, and A. Lerer. 2017 · 2017
Later among the works it cites.
Parallel multiscale autoregressive density estimation
S. Reed, A. van den Oord, N. Kalchbrenner, S. Gómez Colmenarejo, Z. Wang, D. Belov, and N. de Freitas. 2017 · 2017
Later among the works it cites.
PixelCNN++: Improving the PixelCNN with discretized logistic mixture likelihood and other modifications
T. Salimans, A. Karpathy, X. Chen, and D. Kingma. 2017 · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Show, attend and tell: Neural image caption generation with visual attention
K. Xu, J. Ba, R. Kiros, K. Cho, A. Courville, R. Salakhutdinov, R. Zemel, and Y. Bengio. 2015 · 2015
Cited alongside, same era.
Pairwise word interaction modeling with deep neural networks for semantic similarity measurement
Hua He and Jimmy Lin. 2016 · 2016
Cited alongside, same era.
A decomposable attention model for natural language inference
A. Parikh, O. Täckström, D. Das, and J. Uszkoreit. 2016 · 2016
Cited alongside, same era.
Sequence level training with recurrent neural networks
M. Ranzato, S. Chopra, M. Auli, and W. Zaremba. 2016 · 2016
Cited alongside, same era.
Neural machine translation of rare words with subword units
R. Sennrich, B. Haddow, and A. Birch. 2016 · 2016
Cited alongside, same era.
A deep architecture for semantic matching with multiple positional sentence representations
Shengxian Wan, Yanyan Lan, Jiafeng Guo, Jun Xu, Liang Pang, and Xueqi Cheng. 2016 · 2016
Cited alongside, same era.
Later among the works it cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. Gomez, L. Kaiser, and I. Polosukhin. 2017 · 2017
Later among the works it cites.
Sequence modeling via segmentations
C. Wang, Y. Wang, P.-S. Huang, A. Mohamed, D. Zhou, and L. Deng. 2017 · 2017
Later among the works it cites.
Adversarial neural machine translation
L. Wu, Y. Xia, L. Zhao, F. Tian, T. Qin, J. Lai, and T.-Y. Liu. 2017 · 2017
Later among the works it cites.
Towards two-dimensional sequence to sequence model in neural machine translation
Parnia Bahar, Christopher Brix, and Hermann Ney. 2018 · 2018
Closest in time.
Latent alignment and variational attention
Y. Deng, Y. Kim, J. Chiu, D. Guo, and A. Rush. 2018 · 2018
Closest in time.
Classical structured prediction losses for sequence to sequence learning
S. Edunov, M. Ott, M. Auli, D. Grangier, and M. Ranzato. 2018 · 2018
Closest in time.
Towards neural phrase-based machine translation
P. Huang, C. Wang, S. Huang, D. Zhou, and L. Deng. 2018 · 2018
Closest in time.
Weaver: Deep co-encoding of questions and documents for machine reading
Martin Raison, Pierre-Emmanuel Mazaré, Rajarshi Das, and Antoine Bordes. 2018 · 2018
Closest in time.
Convolutional neural network architectures for matching natural language sentences
Baotian Hu, Zhengdong Lu, Hang Li, and Qingcai Chen. 2014 · 2050
Closest in time.