Fetching the paper…
Reading the bibliography…
Zero-shot translation (ZST), which is generally based on a multilingual neural machine translation model, aims to translate between unseen language pairs in training data.
1903
Earlier work this paper cites.
1907
Earlier work this paper cites.
A. Lopez, “Statistical machine translation,” ACM Computing Surveys (CSUR) , vol. 40, no. 3, pp. 1–49, 2008. [Online]. Available: https://dl.acm.org/doi/10.1145/1380584.1380586
2008
Earlier work this paper cites.
L. Van der Maaten and G. Hinton, “Visualizing data using t-sne,” JMLR , 2008. [Online]. Available: http://jmlr.org/papers/v9/vandermaaten08a.html
2008
Earlier work this paper cites.
T. Kudo and J. Richardson, “SentencePiece: A simple and language independent subword tokenizer and detokenizer for neural text processing,” in EMNLP , 2018. [Online]. Available: https://aclanthology.org/D18-2012
2012
Earlier work this paper cites.
K. Cho, B. van Merrienboer, D. Bahdanau, and Y. Bengio, “On the properties of neural machine translation: Encoder-decoder approaches,” in Eighth Workshop on Syntax, Semantics and Structure in Statistical Translation . Association for Computational Linguistics, 2014, pp. 103–111. [Online]. Available: https://aclanthology.org/W14-4012/
2014
Earlier work this paper cites.
2015
Earlier work this paper cites.
D. Dong, H. Wu, W. He, D. Yu, and H. Wang, “Multi-task learning for multiple language translation,” in Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics , 2015, pp. 1723–1732. [Online]. Available: https://aclanthology.org/P15-1166
2015
Earlier work this paper cites.
S. Bengio, O. Vinyals, N. Jaitly, and N. Shazeer, “Scheduled sampling for sequence prediction with recurrent neural networks,” Advances in neural information processing systems , 2015. [Online]. Available: https://proceedings.neurips.cc/paper/2015/hash/e995f98d56967d946471af29d7bf99f1-Abstract.html
2015
Earlier work this paper cites.
O. Firat, K. Cho, and Y. Bengio, “Multi-way, multilingual neural machine translation with a shared attention mechanism,” in Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies , 2016, pp. 866–875. [Online]. Available: https://aclanthology.org/N16-1101
2016
Earlier work this paper cites.
O. Firat, B. Sankaran, Y. Al-onaizan, F. T. Yarman Vural, and K. Cho, “Zero-resource translation with multi-lingual neural machine translation,” in Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing , 2016, pp. 268–277. [Online]. Available: https://aclanthology.org/D16-1026
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
S. Shen, Y. Cheng, Z. He, W. He, H. Wu, M. Sun, and Y. Liu, “Minimum risk training for neural machine translation,” in Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , 2016, pp. 1683–1692. [Online]. Available: https://aclanthology.org/P16-1159
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
M. Johnson, M. Schuster, Q. V. Le, M. Krikun, Y. Wu, Z. Chen, N. Thorat, F. Viégas, M. Wattenberg, G. Corrado, M. Hughes, and J. Dean, “Google’s multilingual neural machine translation system: Enabling zero-shot translation,” TACL , 2017. [Online]. Available: https://aclanthology.org/Q17-1024/
2017
Earlier work this paper cites.
D. Bahdanau, P. Brakel, K. Xu, A. Goyal, R. Lowe, J. Pineau, A. Courville, and Y. Bengio, “An actor-critic algorithm for sequence prediction,” in International Conference on Learning Representations , 2017. [Online]. Available: https://openreview.net/forum?id=SJDaqqveg
2017
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. u. Kaiser, and I. Polosukhin, “Attention is all you need,” in NeurIPS , 2017. [Online]. Available: https://proceedings.neurips.cc/paper/2017/file/3f5ee243547dee91fbd053c1c4a845aa-Paper.pdf
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
G. Blackwood, M. Ballesteros, and T. Ward, “Multilingual neural machine translation with task-specific attention,” in Proceedings of the 27th International Conference on Computational Linguistics , 2018. [Online]. Available: https://aclanthology.org/C18-1263
2018
Earlier work this paper cites.
M. Post, “A call for clarity in reporting BLEU scores,” in WMT , 2018. [Online]. Available: https://aclanthology.org/W18-6319
2018
Earlier work this paper cites.
J. Gu, Y. Wang, K. Cho, and V. O. Li, “Improved zero-shot neural machine translation via ignoring spurious correlations,” in ACL , 2019. [Online]. Available: https://aclanthology.org/P19-1121/
2019
Earlier work this paper cites.
X. Tan, Y. Ren, D. He, T. Qin, and T.-Y. Liu, “Multilingual neural machine translation with knowledge distillation,” in International Conference on Learning Representations , 2019. [Online]. Available: https://openreview.net/forum?id=S1gUsoR9YX
2019
Earlier work this paper cites.
R. Aharoni, M. Johnson, and O. Firat, “Massively multilingual neural machine translation,” in Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies , 2019, pp. 3874–3884. [Online]. Available: https://aclanthology.org/N19-1388
2019
Earlier work this paper cites.
W. Zhang, Y. Feng, F. Meng, D. You, and Q. Liu, “Bridging the gap between training and inference for neural machine translation,” in Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics , 2019. [Online]. Available: https://aclanthology.org/P19-1426
2019
Cited alongside, same era.
W. Du and Y. Ji, “An empirical comparison on imitation learning and reinforcement learning for paraphrase generation,” in Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing , 2019, pp. 6012–6018. [Online]. Available: https://aclanthology.org/D19-1619
2019
Cited alongside, same era.
F. Schmidt, “Generalization in generation: A closer look at exposure bias,” in Proceedings of the 3rd Workshop on Neural Generation and Translation . Association for Computational Linguistics, 2019, pp. 157–167. [Online]. Available: https://aclanthology.org/D19-5616
2019
Cited alongside, same era.
A. Hosseini, S. Reddy, D. Bahdanau, R. D. Hjelm, A. Sordoni, and A. Courville, “Understanding by understanding not: Modeling negation in language models,” in NAACL , 2021. [Online]. Available: https://aclanthology.org/2021.naacl-main.102
2021
Later among the works it cites.
S. Wang, Z. Tu, Z. Tan, S. Shi, M. Sun, and Y. Liu, “On the language coverage bias for neural machine translation,” in ACL , 2021. [Online]. Available: https://aclanthology.org/2021.findings-acl.422/
2021
Later among the works it cites.
Y. Yang, A. Eriguchi, A. Muzio, P. Tadepalli, S. Lee, and H. Hassan, “Improving multilingual translation by representation and gradient regularization,” in EMNLP , 2021. [Online]. Available: https://aclanthology.org/2021.emnlp-main.578/
2021
Later among the works it cites.
Y. Tang, C. Tran, X. Li, P.-J. Chen, N. Goyal, V. Chaudhary, J. Gu, and A. Fan, “Multilingual translation from denoising pre-training,” in Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021 . Association for Computational Linguistics, Aug. 2021, pp. 3450–3466. [Online]. Available: https://aclanthology.org/2021.findings-acl.304
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2019
Cited alongside, same era.
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “BERT: Pre-training of deep bidirectional transformers for language understanding,” in NAACL , 2019. [Online]. Available: https://aclanthology.org/N19-1423
2019
Cited alongside, same era.
M. Ott, S. Edunov, A. Baevski, A. Fan, S. Gross, N. Ng, D. Grangier, and M. Auli, “fairseq: A fast, extensible toolkit for sequence modeling,” in NAACL Demonstrations , 2019. [Online]. Available: https://aclanthology.org/N19-4009/
2019
Cited alongside, same era.
N. Houlsby, A. Giurgiu, S. Jastrzebski, B. Morrone, Q. De Laroussilhe, A. Gesmundo, M. Attariyan, and S. Gelly, “Parameter-efficient transfer learning for NLP,” in Proceedings of the 36th International Conference on Machine Learning , ser. Proceedings of Machine Learning Research, K. Chaudhuri and R. Salakhutdinov, Eds., vol. 97. PMLR, 09–15 Jun 2019, pp. 2790–2799. [Online]. Available: https://proceedings.mlr.press/v97/houlsby19a.html
2019
Cited alongside, same era.
B. Zhang, P. Williams, I. Titov, and R. Sennrich, “Improving massively multilingual neural machine translation and zero-shot translation,” in ACL , 2020. [Online]. Available: https://aclanthology.org/2020.acl-main.148/
2020
Cited alongside, same era.
R. Dabre, C. Chu, and A. Kunchukuttan, “A survey of multilingual neural machine translation,” ACM Comput. Surv. , 2020. [Online]. Available: https://dl.acm.org/doi/abs/10.1145/3406095
2020
Cited alongside, same era.
C. Wang and R. Sennrich, “On exposure bias, hallucination and domain shift in neural machine translation,” in Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics , 2020, pp. 3544–3552. [Online]. Available: https://aclanthology.org/2020.acl-main.326
2020
Cited alongside, same era.
M. Li, S. Roller, I. Kulikov, S. Welleck, Y.-L. Boureau, K. Cho, and J. Weston, “Don’t say that! making inconsistent dialogue unlikely with unlikelihood training,” in ACL , 2020. [Online]. Available: https://aclanthology.org/2020.acl-main.428
2020
Cited alongside, same era.
C. Nogueira dos Santos, X. Ma, R. Nallapati, Z. Huang, and B. Xiang, “Beyond [CLS] through ranking by generation,” in EMNLP , 2020. [Online]. Available: https://aclanthology.org/2020.emnlp-main.134
2020
Cited alongside, same era.
2021
Later among the works it cites.
L. Ding, L. Wang, S. Shi, D. Tao, and Z. Tu, “Redistributing low-frequency words: Making the most of monolingual data in non-autoregressive translation,” in Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics , May 2022. [Online]. Available: https://aclanthology.org/2022.acl-long.172
2022
Later among the works it cites.
C. Zan, L. Ding, L. Shen, Y. Cao, W. Liu, and D. Tao, “On the complementarity between pre-training and random-initialization for resource-rich machine translation,” in Proceedings of the 29th International Conference on Computational Linguistics , 2022, pp. 5029–5034. [Online]. Available: https://aclanthology.org/2022.coling-1.445/
2022
Later among the works it cites.
C. Zan, K. Peng, L. Ding, B. Qiu, B. Liu, S. He, Q. Lu et al. , “Vega-MT: The JD explore academy machine translation system for WMT22,” in Proceedings of the Seventh Conference on Machine Translation (WMT) . Association for Computational Linguistics, Dec. 2022, pp. 411–422. [Online]. Available: https://aclanthology.org/2022.wmt-1.37
2022
Later among the works it cites.
2022
Later among the works it cites.
R. Jin and D. Xiong, “Informative language representation learning for massively multilingual neural machine translation,” in COLING , 2022. [Online]. Available: https://aclanthology.org/2022.coling-1.458/
2022
Later among the works it cites.
S. Sun, A. Fan, J. Cross et al. , “Alternative input signals ease transfer in multilingual machine translation,” in Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics , 2022. [Online]. Available: https://aclanthology.org/2022.acl-long.363
2022
Later among the works it cites.
2022
Later among the works it cites.
K. Arora, L. El Asri, H. Bahuleyan, and J. Cheung, “Why exposure bias matters: An imitation learning perspective of error accumulation in language generation,” in Findings of the Association for Computational Linguistics: ACL 2022 . Association for Computational Linguistics, 2022, pp. 700–710. [Online]. Available: https://aclanthology.org/2022.findings-acl.58
2022
Later among the works it cites.
Z. Wang, J. Wang, and C. Jiang, “Unified multimodal model with unlikelihood training for visual dialog,” in MM , 2022. [Online]. Available: https://doi.org/10.1145/3503161.3547974
2022
Later among the works it cites.
Z. Qu and T. Watanabe, “Adapting to non-centered languages for zero-shot multilingual translation,” in COLING , 2022. [Online]. Available: https://aclanthology.org/2022.coling-1.467
2022
Later among the works it cites.
2022
Later among the works it cites.
S. He, L. Ding, D. Dong, J. Zhang, and D. Tao, “Sparseadapter: An easy approach for improving the parameter-efficiency of adapters,” in Findings of the Association for Computational Linguistics: EMNLP 2022 , 2022, pp. 2184–2190. [Online]. Available: https://aclanthology.org/2022.findings-emnlp.160/
2022
Later among the works it cites.
2022
Later among the works it cites.
K. Peng, L. Ding, Q. Zhong, Y. Ouyang, W. Rong, Z. Xiong, and D. Tao, “Token-level self-evolution training for sequence-to-sequence learning,” in Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics , 2023. [Online]. Available: https://aclanthology.org/2023.acl-short.73
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
A. Joulin, E. Grave, P. Bojanowski, and T. Mikolov, “Bag of tricks for efficient text classification,” in EACL: Short Papers , 2017. [Online]. Available: https://aclanthology.org/E17-2068/
2068
Closest in time.
Y. Qi, D. Sachan, M. Felix, S. Padmanabhan, and G. Neubig, “When and why are pre-trained word embeddings useful for neural machine translation?” in NAACL: Short Papers , 2018. [Online]. Available: https://aclanthology.org/N18-2084
2084
Closest in time.