Fetching the paper…
Reading the bibliography…
Natural language understanding and generation models follow one of the two dominant architectural paradigms: language models (LMs) that process concatenated sequences in a single stack of layers, and encoder-decoder models (EncDec) that utilize separate layer stacks for input and output processing.
The missing ingredient in zero-shot neural machine translation
Arivazhagan, N., Bapna, A., Firat, O., Aharoni, R., Johnson, M., and Macherey, W · 1903
Earlier work this paper cites.
Massively multilingual neural machine translation in the wild: Findings and challenges
Arivazhagan, N., Bapna, A., Firat, O., Lepikhin, D., Johnson, M., Krikun, M., Chen, M. X., Cao, Y., Foster, G., Cherry, C., Macherey, W., Chen, Z., and Wu, Y · 1907
Earlier work this paper cites.
Srilm - an extensible language modeling toolkit
Stolcke, A · 2002
Earlier work this paper cites.
Statistical Machine Translation
Koehn, P · 2010
Earlier work this paper cites.
KenLM: Faster and smaller language model queries
Heafield, K · 2011
Earlier work this paper cites.
SentencePiece: A simple and language independent subword tokenizer and detokenizer for neural text processing
Kudo, T. and Richardson, J · 2012
Earlier work this paper cites.
Recurrent continuous translation models
Kalchbrenner, N. and Blunsom, P · 2013
Earlier work this paper cites.
Learning phrase representations using rnn encoder-decoder for statistical machine translation, 2014
Cho, K., van Merrienboer, B., Gulcehre, C., Bahdanau, D., Bougares, F., Schwenk, H., and Bengio, Y · 2014
Earlier work this paper cites.
Sequence to sequence learning with neural networks
Sutskever, I., Vinyals, O., and Le, Q. V · 2014
Earlier work this paper cites.
Neural Machine Translation by Jointly Learning to Align and Translate
Bahdanau, D., Cho, K., and Bengio, Y · 2015
Earlier work this paper cites.
Multi-way, multilingual neural machine translation with a shared attention mechanism
Firat, O., Cho, K., and Bengio, Y · 2016
Earlier work this paper cites.
Neural machine translation of rare words with subword units
Sennrich, R., Haddow, B., and Birch, A · 2016
Earlier work this paper cites.
Google’s multilingual neural machine translation system: Enabling zero-shot translation
Johnson, M., Schuster, M., Le, Q. V., Krikun, M., Wu, Y., Chen, Z., Thorat, N., Viégas, F., Wattenberg, M., Corrado, G., Hughes, M., and Dean, J · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L. u., and Polosukhin, I · 2017
Earlier work this paper cites.
Layer-wise coordination between encoder and decoder for neural machine translation
He, T., Tan, X., Xia, Y., He, D., Qin, T., Chen, Z., and Liu, T.-Y · 2018
Earlier work this paper cites.
A call for clarity in reporting BLEU scores
Post, M · 2018
Cited alongside, same era.
Adafactor: Adaptive learning rates with sublinear memory cost
Shazeer, N. and Stern, M · 2018
Cited alongside, same era.
Share or not? learning to schedule language-specific capacity for multilingual translation
Zhang, B., Bapna, A., Sennrich, R., and Firat, O · 2018
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2019
Cited alongside, same era.
Unified language model pre-training for natural language understanding and generation
Dong, L., Yang, N., Wang, W., Wei, F., Liu, X., Wang, Y., Gao, J., Zhou, M., and Hon, H.-W · 2019
Cited alongside, same era.
Joint source-target self attention with locality constraints
Statistical power and translationese in machine translation evaluation
Graham, Y., Haddow, B., and Koehn, P · 2020
Later among the works it cites.
Scaling laws for autoregressive generative modeling
Henighan, T. J., Kaplan, J., Katz, M., Chen, M., Hesse, C., Jackson, J., Jun, H., Brown, T. B., Dhariwal, P., Gray, S., Hallacy, C., Mann, B., Radford, A., Ramesh, A., Ryder, N., Ziegler, D. M., Schulman, J., Amodei, D., and McCandlish, S · 2020
Later among the works it cites.
Scaling laws for neural language models
Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D · 2020
Later among the works it cites.
Limits to depth efficiencies of self-attention
Levine, Y., Wies, N., Sharir, O., Bata, H., and Shashua, A · 2020
Later among the works it cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Fonollosa, J. A. R., Casas, N., and Costa-jussà, M. R · 2019
Cited alongside, same era.
Investigating multilingual nmt representations at scale, 2019
Kudugunta, S. R., Bapna, A., Caswell, I., Arivazhagan, N., and Firat, O · 2019
Cited alongside, same era.
Cross-lingual language model pretraining, 2019
Lample, G. and Conneau, A · 2019
Cited alongside, same era.
On NMT search errors and model errors: Cat got your tongue?
Stahlberg, F. and Byrne, B · 2019
Cited alongside, same era.
BERT rediscovers the classical NLP pipeline
Tenney, I., Das, D., and Pavlick, E · 2019
Cited alongside, same era.
Learning deep transformer models for machine translation
Wang, Q., Li, B., Xiao, T., Zhu, J., Li, C., Wong, D. F., and Chao, L. S · 2019
Cited alongside, same era.
FINDINGS OF THE IWSLT 2020 EVALUATION CAMPAIGN
Ansari, E., Axelrod, A., Bach, N., Bojar, O., Cattoni, R., Dalvi, F., Durrani, N., Federico, M., Federmann, C., Gu, J., Huang, F., Knight, K., Ma, X., Nagesh, A., Negri, M., Niehues, J., Pino, J., Salesky, E., Shi, X., Stüker, S., Turchi, M., Waibel, A., and Wang, C · 2020
Cited alongside, same era.
Later among the works it cites.
On negative interference in multilingual models: Findings and a meta-learning treatment
Wang, Z., Lipton, Z. C., and Tsvetkov, Y · 2020
Later among the works it cites.
Improving massively multilingual neural machine translation and zero-shot translation
Zhang, B., Williams, P., Titov, I., and Sennrich, R · 2020
Later among the works it cites.
Explaining neural scaling laws
Bahri, Y., Dyer, E., Kaplan, J., Lee, J., and Sharma, U · 2021
Later among the works it cites.
Scaling laws for neural machine translation
Ghorbani, B., Firat, O., Freitag, M., Bapna, A., Krikun, M., García, X., Chelba, C., and Cherry, C · 2021
Later among the works it cites.
Data and parameter scaling laws for neural machine translation
Gordon, M. A., Duh, K., and Kaplan, J · 2021
Later among the works it cites.
Hernandez, D., Kaplan, J., Henighan, T. J., and McCandlish, S · 2021
Later among the works it cites.
Language models are good translators
Wang, S., Tu, Z., Tan, Z., Wang, W., Sun, M., and Liu, Y · 2021
Later among the works it cites.
Language tags matter for zero-shot neural machine translation
Wu, L., Cheng, S., Wang, M., and Li, L · 2021
Later among the works it cites.
mT5: A massively multilingual pre-trained text-to-text transformer
Xue, L., Constant, N., Roberts, A., Kale, M., Al-Rfou, R., Siddhant, A., Barua, A., and Raffel, C · 2021
Later among the works it cites.
Zhai, X., Kolesnikov, A., Houlsby, N., and Beyer, L · 2021
Later among the works it cites.