Fetching the paper…
Reading the bibliography…
We address the task of machine translation (MT) from extremely low-resource language (ELRL) to English by leveraging cross-lingual transfer from 'closely-related' high-resource language (HRL).
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 1907
Earlier work this paper cites.
Character-based nmt with transformer
Rohit Gupta, Laurent Besacier, Marc Dymetman, and Matthias Gallé. 2019 · 1911
Earlier work this paper cites.
Improving robustness of machine translation with synthetic noise
Vaibhav Vaibhav, Sumeet Singh, Craig Stewart, and Graham Neubig. 2019 · 1920
Earlier work this paper cites.
A new algorithm for data compression
Philip Gage. 1994 · 1994
Earlier work this paper cites.
Automatic evaluation and uniform filter cascades for inducing n-best translation lexicons
I. Dan Melamed. 1995 · 1995
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
chrF: character n-gram F-score for automatic MT evaluation
Maja Popović. 2015 · 2015
Earlier work this paper cites.
A character-level decoder without explicit segmentation for neural machine translation
Junyoung Chung, Kyunghyun Cho, and Yoshua Bengio. 2016 · 2016
Earlier work this paper cites.
Transfer learning for low-resource neural machine translation
Barret Zoph, Deniz Yuret, Jonathan May, and Kevin Knight. 2016 · 2016
Earlier work this paper cites.
Byte-based neural machine translation
Marta R. Costa-jussà, Carlos Escolano, and José A. R. Fonollosa. 2017 · 2017
Earlier work this paper cites.
Google’s multilingual neural machine translation system: Enabling zero-shot translation
Melvin Johnson, Mike Schuster, Quoc V. Le, Maxim Krikun, Yonghui Wu, Zhifeng Chen, Nikhil Thorat, Fernanda Viégas, Martin Wattenberg, Greg Corrado, Macduff Hughes, and Jeffrey Dean. 2017 · 2017
Earlier work this paper cites.
Six challenges for neural machine translation
Philipp Koehn and Rebecca Knowles. 2017 · 2017
Earlier work this paper cites.
Fully Character-Level Neural Machine Translation without Explicit Segmentation
Jason Lee, Kyunghyun Cho, and Thomas Hofmann. 2017 · 2017
Earlier work this paper cites.
Transfer learning across low-resource, related languages for neural machine translation
Toan Q. Nguyen and David Chiang. 2017 · 2017
Earlier work this paper cites.
Toward robust neural machine translation for noisy input sequences
Matthias Sperber, Jan Niehues, and Alex Waibel. 2017 · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
Unsupervised neural machine translation
Mikel Artetxe, Gorka Labaka, Eneko Agirre, and Kyunghyun Cho. 2018 · 2018
Earlier work this paper cites.
How robust are character-based word embeddings in tagging and MT against wrod scramlbing or randdm nouse?
Georg Heigold, Stalin Varanasi, Günter Neumann, and Josef van Genabith. 2018 · 2018
Cited alongside, same era.
SentencePiece: A simple and language independent subword tokenizer and detokenizer for neural text processing
Taku Kudo and John Richardson. 2018 · 2018
Cited alongside, same era.
Leveraging orthographic similarity for multilingual neural transliteration
Anoop Kunchukuttan, Mitesh Khapra, Gurneet Singh, and Pushpak Bhattacharyya. 2018 · 2018
Cited alongside, same era.
Unsupervised machine translation using monolingual corpora only
Guillaume Lample, Alexis Conneau, Ludovic Denoyer, and Marc’Aurelio Ranzato. 2018 · 2018
Cited alongside, same era.
A call for clarity in reporting BLEU scores
Matt Post. 2018 · 2018
Cited alongside, same era.
COMET: A neural framework for MT evaluation
Ricardo Rei, Craig Stewart, Ana C Farinha, and Alon Lavie. 2020 · 2020
Later among the works it cites.
BLEURT: Learning robust metrics for text generation
Thibault Sellam, Dipanjan Das, and Ankur Parikh. 2020 · 2020
Later among the works it cites.
Beyond english-centric multilingual machine translation
Angela Fan, Shruti Bhosale, Holger Schwenk, Zhiyi Ma, Ahmed El-Kishky, Siddharth Goyal, Mandeep Baines, Onur Celebi, Guillaume Wenzek, Vishrav Chaudhary, et al. 2021 · 2021
Later among the works it cites.
Exploiting language relatedness for low web-resource language model adaptation: An Indic languages study
Yash Khemchandani, Sarvesh Mehtani, Vaidehi Patil, Abhijeet Awasthi, Partha Talukdar, and Sunita Sarawagi. 2021 · 2021
Later among the works it cites.
Similar language translation for Catalan, Portuguese and Spanish using Marian NMT
Reinhard Rapp. 2021 · 2021
Later among the works it cites.
Neural machine translation without embeddings
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Xinyi Wang, Hieu Pham, Zihang Dai, and Graham Neubig. 2018 · 2018
Cited alongside, same era.
Training on synthetic noise improves robustness to natural noise in machine translation
Vladimir Karpukhin, Omer Levy, Jacob Eisenstein, and Marjan Ghazvininejad. 2019 · 2019
Cited alongside, same era.
fairseq: A fast, extensible toolkit for sequence modeling
Myle Ott, Sergey Edunov, Alexei Baevski, Angela Fan, Sam Gross, Nathan Ng, David Grangier, and Michael Auli. 2019 · 2019
Cited alongside, same era.
Multilingual neural machine translation with soft decoupled encoding
Xinyi Wang, Hieu Pham, Philip Arthur, and Graham Neubig. 2019 · 2019
Cited alongside, same era.
A call for more rigor in unsupervised cross-lingual learning
Mikel Artetxe, Sebastian Ruder, Dani Yogatama, Gorka Labaka, and Eneko Agirre. 2020 · 2020
Cited alongside, same era.
Unsupervised cross-lingual representation learning at scale
Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov. 2020 · 2020
Cited alongside, same era.
The IndicNLP Library
Anoop Kunchukuttan. 2020 · 2020
Cited alongside, same era.
Uri Shaham and Omer Levy. 2021 · 2021
Later among the works it cites.
Facebook AI’s WMT21 news translation task submission
Chau Tran, Shruti Bhosale, James Cross, Philipp Koehn, Sergey Edunov, and Angela Fan. 2021 · 2021
Later among the works it cites.
Improving zero-shot cross-lingual transfer between closely related languages by injecting character-level noise
Noëmi Aepli and Rico Sennrich. 2022 · 2022
Later among the works it cites.
No language left behind: Scaling human-centered machine translation
Marta R Costa-jussà, James Cross, Onur Çelebi, Maha Elbayad, Kenneth Heafield, Kevin Heffernan, Elahe Kalbassi, Janice Lam, Daniel Licht, Jean Maillard, et al. 2022 · 2022
Later among the works it cites.
Why don’t people use character-level machine translation?
Jindřich Libovický, Helmut Schmid, and Alexander Fraser. 2022 · 2022
Later among the works it cites.
WECHSEL: Effective initialization of subword embeddings for cross-lingual transfer of monolingual language models
Benjamin Minixhofer, Fabian Paischer, and Navid Rekabsaz. 2022 · 2022
Later among the works it cites.
Overlap-based vocabulary generation improves cross-lingual transfer among related languages
Vaidehi Patil, Partha Talukdar, and Sunita Sarawagi. 2022 · 2022
Later among the works it cites.
Samanantar: The largest publicly available parallel corpora collection for 11 indic languages
Gowtham Ramesh, Sumanth Doddapaneni, Aravinth Bheemaraj, Mayank Jobanputra, Raghavan AK, Ajitesh Sharma, Sujit Sahoo, Harshita Diddee, Mahalakshmi J, Divyanshu Kakwani, Navneet Kumar, Aswin Pradeep, Srihari Nagaraj, Kumar Deepak, Vivek Raghavan, Anoop Kunchukuttan, Pratyush Kumar, and Mitesh Shantadevi Khapra. 2022 · 2022
Later among the works it cites.
Aditya Siddhant, Ankur Bapna, Orhan Firat, Yuan Cao, Mia Xu Chen, Isaac Caswell, and Xavier Garcia. 2022 · 2022
Later among the works it cites.
Does manipulating tokenization aid cross-lingual transfer? a study on POS tagging for non-standardized languages
Verena Blaschke, Hinrich Schütze, and Barbara Plank. 2023 · 2023
Closest in time.
The wordpiece algorithm in open source bert
Google-2018. 2022 · 2023
Closest in time.