Fetching the paper…
Reading the bibliography…
Tensor2Tensor is a library for deep learning models that is well-suited for neural machine translation and includes the reference implementation of the state-of-the-art Transformer model.
Long short-term memory
Hochreiter, S. and Schmidhuber, J. (1997) · 1997
Earlier work this paper cites.
Recurrent continuous translation models
Kalchbrenner, N. and Blunsom, P. (2013) · 2013
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Bahdanau, D., Cho, K., and Bengio, Y. (2014) · 2014
Earlier work this paper cites.
Learning phrase representations using RNN encoder-decoder for statistical machine translation
Cho, K., van Merrienboer, B., Gulcehre, C., Bougares, F., Schwenk, H., and Bengio, Y. (2014) · 2014
Earlier work this paper cites.
Sequence to sequence learning with neural networks
Sutskever, I., Vinyals, O., and Le, Q. V. (2014) · 2014
Earlier work this paper cites.
Encoding source language with convolutional neural network for machine translation
Meng, F., Lu, Z., Wang, M., Li, H., Jiang, W., and Liu, Q. (2015) · 2015
Earlier work this paper cites.
Neural machine translation of rare words with subword units
Sennrich, R., Haddow, B., and Birch, A. (2015) · 2015
Earlier work this paper cites.
Tensorflow: A system for large-scale machine learning
Abadi, M., Barham, P., Chen, J., Chen, Z., Davis, A., Dean, J., Devin, M., Ghemawat, S., Irving, G., Isard, M., Kudlur, M., Levenberg, J., Monga, R., Moore, S., Murray, D. G., Steiner, B., Tucker, P., Vasudevan, V., Warden, P., Wicke, M., Yu, Y., and Zheng, X. (2016) · 2016
Cited alongside, same era.
Xception: Deep learning with depthwise separable convolutions
Chollet, F. (2016) · 2016
Cited alongside, same era.
Can active memory replace attention?
Kaiser, Ł. and Bengio, S. (2016) · 2016
Cited alongside, same era.
Neural machine translation in linear time
Kalchbrenner, N., Espeholt, L., Simonyan, K., van den Oord, A., Graves, A., and Kavukcuoglu, K. (2016) · 2016
Cited alongside, same era.
WaveNet: A generative model for raw audio
van den Oord, A., Dieleman, S., Zen, H., Simonyan, K., Vinyals, O., Graves, A., Kalchbrenner, N., Senior, A., and Kavukcuoglu, K. (2016) · 2016
Convolutional sequence to sequence learning
Gehring, J., Auli, M., Grangier, D., Yarats, D., and Dauphin, Y. N. (2017) · 2017
Later among the works it cites.
The reversible residual network: Backpropagation without storing activations
Gomez, A. N., Ren, M., Urtasun, R., and Grosse, R. B. (2017) · 2017
Later among the works it cites.
Depthwise separable convolutions for neural machine translation
Kaiser, L., Gomez, A. N., and Chollet, F. (2017) · 2017
Later among the works it cites.
Outrageously large neural networks: The sparsely-gated mixture-of-experts layer
Shazeer, N., Mirhoseini, A., Maziarz, K., Davis, A., Le, Q., Hinton, G., and Dean, J. (2017) · 2017
Later among the works it cites.
Attention is all you need
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Google’s neural machine translation system: Bridging the gap between human and machine translation
Wu, Y., Schuster, M., Chen, Z., Le, Q. V., Norouzi, M., Macherey, W., Krikun, M., Cao, Y., Gao, Q., Macherey, K., et al. (2016) · 2016
Cited alongside, same era.
Deep recurrent models with fast-forward connections for neural machine translation
Zhou, J., Cao, Y., Wang, X., Li, P., and Xu, W. (2016) · 2016
Cited alongside, same era.
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I. (2017) · 2017
Later among the works it cites.
Image Transformer
Parmar, N., Vaswani, A., Uszkoreit, J., Kaiser, Ł., Shazeer, N., and Ku, A. (2018) · 2018
Closest in time.