Fetching the paper…
Reading the bibliography…
The ability of zero-shot translation emerges when we train a multilingual model with certain translation directions; the model can then directly translate in unseen directions.
Huggingface’s Transformers: State-of-the-art natural language processing
Wolf, T.; et al. 2019 · 1910
Earlier work this paper cites.
A statistical approach to machine translation
Brown, P. F.; et al. 1990 · 1990
Earlier work this paper cites.
Stacked generalization
Wolpert, D. H. 1992 · 1992
Earlier work this paper cites.
Bagging predictors
Breiman, L. 1996 · 1996
Earlier work this paper cites.
Analyzing bagging
Bühlmann, P.; and Yu, B. 2002 · 2002
Earlier work this paper cites.
BLEU: A method for automatic evaluation of machine translation
Papineni, K.; Roukos, S.; Ward, T.; and Zhu, W.-J. 2002 · 2002
Earlier work this paper cites.
The boosting approach to machine learning: An overview
Schapire, R. E. 2003 · 2003
Earlier work this paper cites.
Europarl: A parallel corpus for statistical machine translation
Koehn, P. 2005 · 2005
Earlier work this paper cites.
A study of translation edit rate with targeted human annotation
Snover, M.; Dorr, B.; Schwartz, R.; Micciulla, L.; and Makhoul, J. 2006 · 2006
Earlier work this paper cites.
Translating from under-resourced languages: Comparing direct transfer against pivot translation
Babych, B.; Hartley, A.; and Sharoff, S. 2007 · 2007
Earlier work this paper cites.
Statistical post-editing on SYSTRAN’s rule-based translation system
Dugast, L.; Senellart, J.; and Koehn, P. 2007 · 2007
Earlier work this paper cites.
Pivot language approach for phrase-based statistical machine translation
Wu, H.; and Wang, H. 2007 · 2007
Earlier work this paper cites.
Multi-class AdaBoost
Hastie, T.; Rosset, S.; Zhu, J.; and Zou, H. 2009 · 2009
Earlier work this paper cites.
Statistical Machine Translation
Koehn, P. 2009 · 2009
Earlier work this paper cites.
Revisiting pivot language approach for machine translation
Wu, H.; and Wang, H. 2009 · 2009
Earlier work this paper cites.
Double Q-learning
Hasselt, H. 2010 · 2010
Earlier work this paper cites.
Apertium: A free/open-source platform for rule-based machine translation
Forcada, M. L.; et al. 2011 · 2011
Earlier work this paper cites.
Gradient boosting machines, a tutorial
Natekin, A.; and Knoll, A. 2013 · 2013
Earlier work this paper cites.
Ensemble triangulation for statistical machine translation
Razmara, M.; and Sarkar, A. 2013 · 2013
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Bahdanau, D.; Cho, K.; and Bengio, Y. 2015 · 2015
Earlier work this paper cites.
Mathematical Statistics: Basic Ideas and Selected Topics
Bickel, P. J.; and Doksum, K. A. 2015 · 2015
Earlier work this paper cites.
Sequence-level knowledge distillation
Kim, Y.; and Rush, A. M. 2016 · 2016
Earlier work this paper cites.
Overview of the IWSLT 2017 evaluation campaign
Cettolo, M.; et al. 2017 · 2017
Earlier work this paper cites.
MMCR4NLP: Multilingual multiway corpora repository for natural language Processing
Dabre, R.; and Kurohashi, S. 2017 · 2017
Earlier work this paper cites.
Ensemble distillation for neural machine translation
Freitag, M.; Al-Onaizan, Y.; and Sankaran, B. 2017 · 2017
Cited alongside, same era.
Google’s multilingual neural machine translation system: Enabling zero-shot translation
Johnson, M.; et al. 2017 · 2017
Cited alongside, same era.
chrF++: Words helping character n-grams
Popović, M. 2017 · 2017
Cited alongside, same era.
Attention is all you need
Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, Ł.; and Polosukhin, I. 2017 · 2017
Cited alongside, same era.
Frustratingly easy model ensemble for abstractive summarization
Kobayashi, H. 2018 · 2018
Cited alongside, same era.
A call for clarity in reporting BLEU scores
Post, M. 2018 · 2018
Cited alongside, same era.
Ensembling of distilled models from multi-task teachers for constrained resource language pairs
Hendy, A.; et al. 2021 · 2021
Later among the works it cites.
Improving zero-shot translation by disentangling positional information
Liu, D.; Niehues, J.; Cross, J.; Guzmán, F.; and Li, X. 2021 · 2021
Later among the works it cites.
Understanding the properties of minimum Bayes risk decoding in neural machine translation
Müller, M.; and Sennrich, R. 2021 · 2021
Later among the works it cites.
One teacher is enough? Pre-trained language model distillation from multiple teachers
Wu, C.; Wu, F.; and Huang, Y. 2021 · 2021
Later among the works it cites.
Multilingual machine translation with hyper-adapters
Baziotis, C.; Artetxe, M.; Cross, J.; and Bhosale, S. 2022 · 2022
Later among the works it cites.
Ensemble deep learning: A review
Ganaie, M.; Hu, M.; Malik, A.; Tanveer, M.; and Suganthan, P. 2022 · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Adaptive knowledge sharing in multi-task learning: Improving low-resource neural machine translation
Zaremoodi, P.; Buntine, W.; and Haffari, G. 2018 · 2018
Cited alongside, same era.
Improved zero-shot neural machine translation via ignoring spurious Correlations
Gu, J.; Wang, Y.; Cho, K.; and Li, V. O. 2019 · 2019
Cited alongside, same era.
The curious case of neural text degeneration
Holtzman, A.; Buys, J.; Du, L.; Forbes, M.; and Choi, Y. 2019 · 2019
Cited alongside, same era.
fairseq: A fast, extensible toolkit for sequence modeling
Ott, M.; et al. 2019 · 2019
Cited alongside, same era.
Language models are unsupervised multitask learners
Radford, A.; et al. 2019 · 2019
Cited alongside, same era.
Why do neural dialog systems generate short and meaningless replies? A comparison between dialog and translation
Wei, B.; Lu, S.; Mou, L.; Zhou, H.; Poupart, P.; Li, G.; and Jin, Z. 2019 · 2019
Cited alongside, same era.
Later among the works it cites.
Search and learn: Improving semantic coverage for data-to-text generation
Jolly, S.; Zhang, Z. X.; Dengel, A.; and Mou, L. 2022 · 2022
Later among the works it cites.
BLOOM: A 176B-parameter open-access multilingual language model
Scao, T. L.; et al. 2022 · 2022
Later among the works it cites.
NMTScore: A multilingual analysis of translation-based text similarity measures
Vamvas, J.; and Sennrich, R. 2022 · 2022
Later among the works it cites.
Parameter differentiation based multilingual neural machine translation
Wang, Q.; and Zhang, J. 2022 · 2022
Later among the works it cites.
The effects of language Token prefixing for multilingual machine translation
Wicks, R.; and Duh, K. 2022 · 2022
Later among the works it cites.
Language-family adapters for low-resource multilingual neural machine translation
Chronopoulou, A.; Stojanovski, D.; and Fraser, A. 2023 · 2023
Later among the works it cites.
Gradient-based gradual pruning for language-specific multilingual neural machine translation
He, D.; et al. 2023 · 2023
Later among the works it cites.
Scaling data-constrained language models
Muennighoff, N.; et al. 2023 · 2023
Later among the works it cites.
Neural machine translation for low-resource languages: A survey
Ranathunga, S.; Lee, E.-S. A.; Prifti Skenduli, M.; Shekhar, R.; Alam, M.; and Kaur, R. 2023 · 2023
Later among the works it cites.
Improving low resource speech translation with data augmentation and ensemble strategies
Shanbhogue, A. V. K.; Xue, R.; Saha, S.; Zhang, D.; and Ganesan, A. 2023 · 2023
Later among the works it cites.
A survey on ensemble learning under the era of deep learning
Yang, Y.; Lv, H.; and Chen, N. 2023 · 2023
Later among the works it cites.
How effective is multi-source pivoting for translation of low resource Indian languages?
Gaikwad, P.; Doshi, M.; Dabre, R.; and Bhattacharyya, P. 2024 · 2024
Closest in time.
Investigating multi-pivot ensembling with massively multilingual machine translation models
Mohammadshahi, A.; Vamvas, J.; and Sennrich, R. 2024 · 2024
Closest in time.
Ensemble distillation for unsupervised constituency parsing
Shayegh, B.; Cao, Y.; Zhu, X.; Cheung, J. C.; and Mou, L. 2024 · 2024
Closest in time.
Tree-averaging algorithms for ensemble-based unsupervised discontinuous constituency parsing
Shayegh, B.; Wen, Y.; and Mou, L. 2024 · 2024
Closest in time.
Error diversity matters: An error-resistant ensemble method for unsupervised dependency parsing
Shayegh, B.; et al. 2025 · 2025
Closest in time.
Deep reinforcement learning with double Q-learning
van Hasselt, H.; Guez, A.; and Silver, D. 2016 · 2094
Closest in time.