Fetching the paper…
Reading the bibliography…
Multilingual Neural Machine Translation has been showing great success using transformer models.
Deep encoder, shallow decoder: Reevaluating non-autoregressive machine translation
Kasai, J., Pappas, N., Peng, H., Cross, J., and Smith, N. A. (2020) · 2006
Earlier work this paper cites.
Professional CUDA c programming
Cheng, J., Grossman, M., and McKercher, T. (2014) · 2014
Earlier work this paper cites.
On symmetric and asymmetric lshs for inner product search
Neyshabur, B. and Srebro, N. (2014) · 2014
Earlier work this paper cites.
Efficient softmax approximation for gpus
Grave, E., Joulin, A., Cissé, M., Grangier, D., and Jégou, H. (2016) · 2016
Earlier work this paper cites.
Google’s multilingual neural machine translation system: Enabling zero-shot translation
Johnson, M., Schuster, M., Le, Q. V., Krikun, M., Wu, Y., Chen, Z., Thorat, N., Viégas, F., Wattenberg, M., Corrado, G., Hughes, M., and Dean, J. (2017) · 2017
Earlier work this paper cites.
Speeding up neural machine translation decoding by shrinking run-time vocabulary
Shi, X. and Knight, K. (2017) · 2017
Earlier work this paper cites.
Svd-softmax: Fast softmax approximation on large vocabulary neural networks
Shim, K., Lee, M., Choi, I., Boo, Y., and Sung, W. (2017) · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I. (2017) · 2017
Cited alongside, same era.
Machine translation human evaluation: an investigation of evaluation based on post-editing and its relation with direct assessment
Bentivogli, L., Cettolo, M., Federico, M., and Christian, F. (2018) · 2018
Cited alongside, same era.
Learning to screen for fast softmax inference on large vocabulary neural networks
Chen, P. H., Si, S., Kumar, S., Li, Y., and Hsieh, C.-J. (2018) · 2018
Cited alongside, same era.
Appraise evaluation framework for machine translation
Federmann, C. (2018) · 2018
Cited alongside, same era.
Fast locality sensitive hashing for beam search on gpu
Shi, X., Xu, S., and Knight, K. (2018) · 2018
Cited alongside, same era.
Billion-scale similarity search with GPUs
Johnson, J., Douze, M., and Jégou, H. (2019) · 2019
Later among the works it cites.
From research to production and back: Ludicrously fast neural machine translation
Kim, Y. J., Junczys-Dowmunt, M., Hassan, H., Aji, A. F., Heafield, K., Grundkiewicz, R., and Bogoychev, N. (2019) · 2019
Later among the works it cites.
On layer normalization in the transformer architecture
Xiong, R., Yang, Y., He, D., Zheng, K., Zheng, S., Xing, C., Zhang, H., Lan, Y., Wang, L., and Liu, T. (2020) · 2020
Later among the works it cites.
Findings of the 2021 conference on machine translation (WMT21)
Akhbardeh, F., Arkhangorodsky, A., Biesialska, M., Bojar, O., Chatterjee, R., Chaudhary, V., Costa-jussà, M. R., España-Bonet, C., Fan, A., Federmann, C., et al. (2021) · 2021
Later among the works it cites.
Scalable and efficient moe training for multitask multilingual models
Kim, Y. J., Awan, A. A., Muzio, A., Salinas, A. F. C., Lu, L., Hendy, A., Rajbhandari, S., He, Y., and Awadalla, H. H. (2021) · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Zhang, M., Liu, X., Wang, W., Gao, J., and He, Y. (2018) · 2018
Cited alongside, same era.
Multilingual machine translation systems from microsoft for wmt21 shared task
Yang, J., Ma, S., Huang, H., Zhang, D., Dong, L., Huang, S., Muzio, A., Singhal, S., Awadalla, H. H., Song, X., et al. (2021) · 2021
Later among the works it cites.