Fetching the paper…
Reading the bibliography…
We explore the use of segments learnt using Byte Pair Encoding (referred to as BPE units) as basic units for statistical machine translation between related languages and compare it with orthographic syllables, which are currently the best performing basic units for this translation task.
Proposition 16
Nikolai Trubetzkoy. 1928 · 1928
Earlier work this paper cites.
India as a lingustic area
Murray B Emeneau. 1956 · 1956
Earlier work this paper cites.
A new algorithm for data compression
Philip Gage. 1994 · 1994
Earlier work this paper cites.
Automatic evaluation and uniform filter cascades for inducing n-best translation lexicons
I Dan Melamed. 1995 · 1995
Earlier work this paper cites.
Linguistic areas and language history
Sarah Thomason. 2000 · 2000
Earlier work this paper cites.
The european linguistic area: Standard average european
Martin Haspelmath. 2001 · 2001
Earlier work this paper cites.
BLEU: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
Statistical significance tests for machine translation evaluation
Philipp Koehn. 2004 · 2004
Earlier work this paper cites.
Is Japanese Related to Korean, Tungusic, Mongolic and Turkic?
Martine Irma Robbeets. 2005 · 2005
Earlier work this paper cites.
Corpus portal for search in monolingual corpora
Uwe Quasthoff, Matthias Richter, and Christian Biemann. 2006 · 2006
Earlier work this paper cites.
Moses: Open source toolkit for Statistical Machine Translation
Philipp Koehn, Hieu Hoang, Alexandra Birch, Chris Callison-Burch, Marcello Federico, Nicola Bertoldi, Brooke Cowan, Wade Shen, Christine Moran, Richard Zens, et al. 2007 · 2007
Earlier work this paper cites.
Can we translate letters?
David Vilar, Jan-T Peter, and Hermann Ney. 2007 · 2007
Earlier work this paper cites.
Urdu word segmentation
Nadir Durrani and Sarmad Hussain. 2010 · 2010
Cited alongside, same era.
Hindi-to-Urdu machine translation through transliteration
Nadir Durrani, Hassan Sajjad, Alexander Fraser, and Helmut Schmid. 2010 · 2010
Cited alongside, same era.
Korea-Japonica: A Re-Evaluation of a Common Genetic Origin
Alexander Vovin. 2010 · 2010
Cited alongside, same era.
Batch tuning strategies for statistical machine translation
Colin Cherry and George Foster. 2012 · 2012
Cited alongside, same era.
The TDIL program and the Indian Language Corpora Initiative
Girish Nath Jha. 2012 · 2012
Cited alongside, same era.
Combining word-level and character-level models for machine translation between closely-related languages
Preslav Nakov and Jörg Tiedemann. 2012 · 2012
Cited alongside, same era.
HindEnCorp – Hindi-English and Hindi-only Corpus for Machine Translation
Ondřej Bojar, Vojtěch Diatka, Pavel Rychlý, Pavel Straňák, Vít Suchomel, Aleš Tamchyna, and Daniel Zeman. 2014 · 2014
Later among the works it cites.
Urdu monolingual corpus
Bushra Jawaid, Amir Kamran, and Ondřej Bojar. 2014 · 2014
Later among the works it cites.
The IIT Bombay SMT System for ICON 2014 Tools contest
Anoop Kunchukuttan, Ratish Pudupully, Rajen Chatterjee, Abhijit Mishra, and Pushpak Bhattacharyya. 2014 · 2014
Later among the works it cites.
Morfessor 2.0: Toolkit for statistical morphological segmentation
Peter Smit, Sami Virpioja, Stig-Arne Grönroos, and Mikko Kurimo. 2014 · 2014
Later among the works it cites.
Sequence to sequence learning with neural networks
Ilya Sutskever, Oriol Vinyals, and Quoc V Le. 2014 · 2014
Later among the works it cites.
Variable-length word encodings for neural translation models
Rohan Chitnis and John DeNero. 2015 · 2015
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Morphological Processing for English-Tamil Statistical Machine Translation
Loganathan Ramasamy, Ondřej Bojar, and Zdeněk Žabokrtský. 2012 · 2012
Cited alongside, same era.
Japanese and korean voice search
Mike Schuster and Kaisuke Nakajima. 2012 · 2012
Cited alongside, same era.
South Asian Languages: A Syntactic Typology
Karumuri Subbarao. 2012 · 2012
Cited alongside, same era.
Character-based pivot translation for under-resourced languages and domains
Jörg Tiedemann. 2012 · 2012
Cited alongside, same era.
Source language adaptation for resource-poor machine translation
Pidong Wang, Preslav Nakov, and Hwee Tou Ng. 2012 · 2012
Cited alongside, same era.
Analyzing the use of character-level translation with sparse and noisy datasets
Jörg Tiedemann and Preslav Nakov. 2013 · 2013
Cited alongside, same era.
Later among the works it cites.
Brahmi-Net: A transliteration and script conversion system for languages of the Indian subcontinent
Anoop Kunchukuttan, Ratish Puduppully, and Pushpak Bhattacharyya. 2015 · 2015
Later among the works it cites.
Lebleu: N-gram-based translation evaluation score for morphologically complex languages
Sami Virpioja and Stig-Arne Grönroos. 2015 · 2015
Later among the works it cites.
Statistical machine translation between related languages
Pushpak Bhattacharyya, Mitesh Khapra, and Anoop Kunchukuttan. 2016 · 2016
Closest in time.
A character-level decoder without explicit segmentation for neural machine translation
Junyoung Chung, Kyunghyun Cho, and Yoshua Bengio. 2016 · 2016
Closest in time.
Neural machine translation of rare words with subword units
Rico Sennrich, Barry Haddow, and Alexandra Birch. 2016 · 2016
Closest in time.
Google’s neural machine translation system: Bridging the gap between human and machine translation
Y. Wu, M. Schuster, Z. Chen, Q. V. Le, and M. Norouzi. 2016 · 2016
Closest in time.