Fetching the paper…
Reading the bibliography…
Multilingual machine translation (MMT) benefits from cross-lingual transfer but is a challenging multitask optimization problem.
Massively multilingual neural machine translation in the wild: Findings and challenges
Naveen Arivazhagan, Ankur Bapna, Orhan Firat, Dmitry Lepikhin, Melvin Johnson, Maxim Krikun, Mia Xu Chen, Yuan Cao, George F. Foster, Colin Cherry, Wolfgang Macherey, Zhifeng Chen, and Yonghui Wu. 2019 · 1907
Earlier work this paper cites.
Everything is all it takes: A multipronged strategy for zero-shot cross-lingual information extraction
Mahsa Yarmohammadi, Shijie Wu, Marc Marone, Haoran Xu, Seth Ebner, Guanghui Qin, Yunmo Chen, Jialiang Guo, Craig Harman, Kenton Murray, Aaron Steven White, Mark Dredze, and Benjamin Van Durme. 2021 · 1967
Earlier work this paper cites.
Google’s multilingual neural machine translation system: Enabling zero-shot translation
Melvin Johnson, Mike Schuster, Quoc V. Le, Maxim Krikun, Yonghui Wu, Zhifeng Chen, Nikhil Thorat, Fernanda Viégas, Martin Wattenberg, Greg Corrado, Macduff Hughes, and Jeffrey Dean. 2017 · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
SentencePiece: A simple and language independent subword tokenizer and detokenizer for neural text processing
Taku Kudo and John Richardson. 2018 · 2018
Earlier work this paper cites.
A call for clarity in reporting BLEU scores
Matt Post. 2018 · 2018
Earlier work this paper cites.
Massively multilingual neural machine translation
Roee Aharoni, Melvin Johnson, and Orhan Firat. 2019 · 2019
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Earlier work this paper cites.
Global sparse momentum sgd for pruning very deep neural networks
Xiaohan Ding, Xiangxin Zhou, Yuchen Guo, Jungong Han, Ji Liu, et al. 2019 · 2019
Earlier work this paper cites.
Importance estimation for neural network pruning
Pavlo Molchanov, Arun Mallya, Stephen Tyree, Iuri Frosio, and Jan Kautz. 2019 · 2019
Cited alongside, same era.
Mass: Masked sequence to sequence pre-training for language generation
Kaitao Song, Xu Tan, Tao Qin, Jianfeng Lu, and Tie-Yan Liu. 2019 · 2019
Cited alongside, same era.
Autoprune: Automatic network pruning by regularizing auxiliary parameters
Xia Xiao, Zigeng Wang, and Sanguthevar Rajasekaran. 2019 · 2019
Cited alongside, same era.
Electra: Pre-training text encoders as discriminators rather than generators
Kevin Clark, Minh-Thang Luong, Quoc V Le, and Christopher D Manning. 2020 · 2020
Cited alongside, same era.
Multilingual denoising pre-training for neural machine translation
Yinhan Liu, Jiatao Gu, Naman Goyal, Xian Li, Sergey Edunov, Marjan Ghazvininejad, Mike Lewis, and Luke Zettlemoyer. 2020 · 2020
Cited alongside, same era.
Movement pruning: Adaptive sparsity by fine-tuning
Learning language specific sub-network for multilingual machine translation
Zehui Lin, Liwei Wu, Mingxuan Wang, and Lei Li. 2021 · 2021
Later among the works it cites.
A gradient flow framework for analyzing network pruning
Ekdeep Singh Lubana and Robert Dick. 2021 · 2021
Later among the works it cites.
BERT, mBERT, or BiBERT? a study on contextualized embeddings for neural machine translation
Haoran Xu, Benjamin Van Durme, and Kenton Murray. 2021 · 2021
Later among the works it cites.
Share or not? learning to schedule language-specific capacity for multilingual translation
Biao Zhang, Ankur Bapna, Rico Sennrich, and Orhan Firat. 2021 · 2021
Later among the works it cites.
No parameters left behind: Sensitivity guided adaptive learning rate for training large transformer models
Chen Liang, Haoming Jiang, Simiao Zuo, Pengcheng He, Xiaodong Liu, Jianfeng Gao, Weizhu Chen, and Tuo Zhao. 2022 · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Victor Sanh, Thomas Wolf, and Alexander Rush. 2020 · 2020
Cited alongside, same era.
Leveraging monolingual data with self-supervision for multilingual neural machine translation
Aditya Siddhant, Ankur Bapna, Yuan Cao, Orhan Firat, Mia Chen, Sneha Kudugunta, Naveen Arivazhagan, and Yonghui Wu. 2020 · 2020
Cited alongside, same era.
Multi-task learning for multilingual neural machine translation
Yiren Wang, ChengXiang Zhai, and Hany Hassan. 2020 · 2020
Cited alongside, same era.
Scalable and efficient moe training for multitask multilingual models
Young Jin Kim, Ammar Ahmad Awan, Alexandre Muzio, Andres Felipe Cruz Salinas, Liyang Lu, Amr Hendy, Samyam Rajbhandari, Yuxiong He, and Hany Hassan Awadalla. 2021 · 2021
Cited alongside, same era.
NLLB Team, Marta R. Costa-jussà, James Cross, Onur Çelebi, Maha Elbayad, Kenneth Heafield, Kevin Heffernan, Elahe Kalbassi, Janice Lam, Daniel Licht, Jean Maillard, Anna Sun, Skyler Wang, Guillaume Wenzek, Al Youngblood, Bapi Akula, Loic Barrault, Gabriel Mejia-Gonzalez, Prangthip Hansanti, John Hoffman, Semarley Jarrett, Kaushik Ram Sadagopan, Dirk Rowe, Shannon Spruit, Chau Tran, Pierre Andrews, Necip Fazil Ayan, Shruti Bhosale, Sergey Edunov, Angela Fan, Cynthia Gao, Vedanuj Goswami, Francisco Guzmán, Philipp Koehn, Alexandre Mourachko, Christophe Ropers, Safiyyah Saleem, Holger Schwenk, and Jeff Wang. 2022 · 2022
Later among the works it cites.
Aditya Siddhant, Ankur Bapna, Orhan Firat, Yuan Cao, Mia Xu Chen, Isaac Caswell, and Xavier Garcia. 2022 · 2022
Later among the works it cites.
The importance of being parameters: An intra-distillation method for serious gains
Haoran Xu, Philipp Koehn, and Kenton Murray. 2022 · 2022
Later among the works it cites.