Fetching the paper…
Reading the bibliography…
Despite being the seventh most widely spoken language in the world, Bengali has received much less attention in machine translation literature due to being low in resources.
Wikimatrix: Mining 135m parallel sentences in 1620 language pairs from wikipedia
Holger Schwenk, Vishrav Chaudhary, Shuo Sun, Hongyu Gong, and Francisco Guzmán. 2019 · 1907
Earlier work this paper cites.
ANGLABHARTI: a multilingual machine aided translation project on translation from English to Indian languages
RMK Sinha, K Sivaraman, Aditi Agrawal, Renu Jain, Rakesh Srivastava, and Ajai Jain. 1995 · 1995
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
Morphological analysis of Bangla words for automatic machine translation
MM Asaduzzaman and Muhammad Masroor Ali. 2003 · 2003
Earlier work this paper cites.
Statistical phrase-based translation
Philipp Koehn, Franz J. Och, and Daniel Marcu. 2003 · 2003
Earlier work this paper cites.
An optimal way of machine translation from English to Bengali
Sajib Dasgupta, Abu Wasif, and Sharmin Azam. 2004 · 2004
Earlier work this paper cites.
A phrasal EBMT system for translating English to Bengali
Sudip Kumar Naskar and Sivaji Bandyopadhyay. 2005 · 2005
Earlier work this paper cites.
A semantics-based English-Bengali EBMT system for translating news headlines
Diganta Saha and Sivaji Bandyopadhyay. 2005 · 2005
Earlier work this paper cites.
Parallel corpora for medium density languages
Dániel Varga, Péter Halácsy, András Kornai, Viktor Nagy, László Németh, and Viktor Trón. 2005 · 2005
Earlier work this paper cites.
Moses: Open source toolkit for statistical machine translation
Philipp Koehn, Hieu Hoang, Alexandra Birch, Chris Callison-Burch, Marcello Federico, Nicola Bertoldi, Brooke Cowan, Wade Shen, Christine Moran, Richard Zens, Chris Dyer, Ondřej Bojar, Alexandra Constantin, and Evan Herbst. 2007 · 2007
Earlier work this paper cites.
Active learning for statistical phrase-based machine translation
Gholamreza Haffari, Maxim Roy, and Anoop Sarkar. 2009 · 2009
Earlier work this paper cites.
A semi-supervised approach to Bengali-English phrase-based statistical machine translation
Maxim Roy. 2009 · 2009
Earlier work this paper cites.
Improved unsupervised sentence alignment for symmetrical and asymmetrical parallel corpora
Fabienne Braune and Alexander Fraser. 2010 · 2010
Earlier work this paper cites.
English to Bangla phrase-based machine translation
Md Zahurul Islam, Jörg Tiedemann, and Andreas Eisele. 2010 · 2010
Earlier work this paper cites.
MT-based sentence alignment for OCR-generated parallel texts
Rico Sennrich and Martin Volk. 2010 · 2010
Earlier work this paper cites.
Extrinsic evaluation of sentence alignment systems
Sadaf Abdul-Rauf, Mark Fishel, Patrik Lambert, Sandra Noubours, and Rico Sennrich. 2012 · 2012
Earlier work this paper cites.
Wit3: Web inventory of transcribed and translated talks
Mauro Cettolo, Christian Girardi, and Marcello Federico. 2012 · 2012
Cited alongside, same era.
SUPara: a balanced English-Bengali parallel corpus
Md Abdullah Al Mumin, Abu Awal Md Shoeb, Md Reza Selim, and Muhammed Zafar Iqbal. 2012 · 2012
Cited alongside, same era.
Constructing parallel corpora for six Indian languages via crowdsourcing
Matt Post, Chris Callison-Burch, and Miles Osborne. 2012 · 2012
Cited alongside, same era.
Parallel data, tools and interfaces in OPUS
Jörg Tiedemann. 2012 · 2012
Cited alongside, same era.
Polyglot: Distributed word representations for multilingual NLP
Rami Al-Rfou’, Bryan Perozzi, and Steven Skiena. 2013 · 2013
Cited alongside, same era.
Combining bilingual and comparable corpora for low resource machine translation
Ann Irvine and Chris Callison-Burch. 2013 · 2013
Multilingual Indian language translation system at WAT 2018: Many-to-one phrase-based SMT
Tamali Banerjee, Anoop Kunchukuttan, and Pushpak Bhattacharya. 2018 · 2018
Later among the works it cites.
Training deployable general domain MT for a low resource language pair: English–Bangla
Sandipan Dandapat and William Lewis. 2018 · 2018
Later among the works it cites.
Universal neural machine translation for extremely low resource languages
Jiatao Gu, Hany Hassan, Jacob Devlin, and Victor O.K. Li. 2018 · 2018
Later among the works it cites.
On the impact of various types of noise on neural machine translation
Huda Khayrallah and Philipp Koehn. 2018 · 2018
Later among the works it cites.
Subword regularization: Improving neural network translation models with multiple subword candidates
Taku Kudo. 2018 · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Yet another fast, robust and open source sentence aligner. time to reconsider sentence alignment
Fethi Lamraoui and Philippe Langlais. 2013 · 2013
Cited alongside, same era.
The AMARA corpus: Building parallel language resources for the educational domain
Ahmed Abdelali, Francisco Guzman, Hassan Sajjad, and Stephan Vogel. 2014 · 2014
Cited alongside, same era.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. 2015 · 2015
Cited alongside, same era.
Improving neural machine translation models with monolingual data
Rico Sennrich, Barry Haddow, and Alexandra Birch. 2016 · 2016
Cited alongside, same era.
Google’s neural machine translation system: Bridging the gap between human and machine translation
Yonghui Wu, Mike Schuster, Zhifeng Chen, Quoc V Le, Mohammad Norouzi, Wolfgang Macherey, Maxim Krikun, Yuan Cao, Qin Gao, Klaus Macherey, et al. 2016 · 2016
Cited alongside, same era.
Google’s multilingual neural machine translation system: Enabling zero-shot translation
Melvin Johnson, Mike Schuster, Quoc V. Le, Maxim Krikun, Yonghui Wu, Zhifeng Chen, Nikhil Thorat, Fernanda Viégas, Martin Wattenberg, Greg Corrado, Macduff Hughes, and Jeffrey Dean. 2017 · 2017
Cited alongside, same era.
Taku Kudo and John Richardson. 2018 · 2018
Later among the works it cites.
OpenSubtitles2018: Statistical rescoring of sentence alignments in large, noisy parallel corpora
Pierre Lison, Jörg Tiedemann, and Milen Kouylekov. 2018 · 2018
Later among the works it cites.
SUPara-benchmark: A benchmark dataset for English-Bangla machine translation
Md Abdullah Al Mumin, Md Hanif Seddiqui, Muhammed Zafar Iqbal, and Mohammed Jahirul Islam. 2018 · 2018
Later among the works it cites.
A call for clarity in reporting BLEU scores
Matt Post. 2018 · 2018
Later among the works it cites.
JW300: A wide-coverage parallel corpus for low-resource languages
Željko Agić and Ivan Vulić. 2019 · 2019
Later among the works it cites.
Margin-based parallel corpus mining with multilingual sentence embeddings
Mikel Artetxe and Holger Schwenk. 2019 · 2019
Later among the works it cites.
Low-resource corpus filtering using multilingual sentence embeddings
Vishrav Chaudhary, Yuqing Tang, Francisco Guzmán, Holger Schwenk, and Philipp Koehn. 2019 · 2019
Later among the works it cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Later among the works it cites.
The FLORES evaluation datasets for low-resource machine translation: Nepali–English and Sinhala–English
Francisco Guzmán, Peng-Jen Chen, Myle Ott, Juan Pino, Guillaume Lample, Philipp Koehn, Vishrav Chaudhary, and Marc’Aurelio Ranzato. 2019 · 2019
Later among the works it cites.
Neural machine translation for the Bangla-English language pair
Md. Arid Hasan, Firoj Alam, Shammur Absar Chowdhury, and Naira Khan. 2019 · 2019
Later among the works it cites.
Findings of the WMT 2019 shared task on parallel corpus filtering for low-resource conditions
Philipp Koehn, Francisco Guzmán, Vishrav Chaudhary, and Juan Pino. 2019 · 2019
Later among the works it cites.