Fetching the paper…
Reading the bibliography…
An ongoing challenge in current natural language processing is how its major advancements tend to disproportionately favor resource-rich languages, leaving a significant number of under-resourced languages behind.
The mathematics of statistical machine translation: Parameter estimation
Peter F Brown, Stephen A Della Pietra, Vincent J Della Pietra, and Robert L Mercer. 1993 · 1993
Earlier work this paper cites.
Hmm-based word alignment in statistical translation
Stephan Vogel, Hermann Ney, and Christoph Tillmann. 1996 · 1996
Earlier work this paper cites.
Improved statistical alignment models
Franz Josef Och and Hermann Ney. 2000 · 2000
Earlier work this paper cites.
An evaluation exercise for word alignment
Rada Mihalcea and Ted Pedersen. 2003 · 2003
Earlier work this paper cites.
Minimum error rate training in statistical machine translation
Franz Josef Och. 2003 · 2003
Earlier work this paper cites.
A systematic comparison of various statistical alignment models
Franz Josef Och and Hermann Ney. 2003 · 2003
Earlier work this paper cites.
End-to-end neural word alignment outperforms GIZA++
Thomas Zenkel, Joern Wuebker, and John DeNero. 2020 · 2003
Earlier work this paper cites.
Europarl: A parallel corpus for statistical machine translation
Philipp Koehn. 2005 · 2005
Earlier work this paper cites.
Edinburgh system description for the 2005 IWSLT speech translation evaluation
Philipp Koehn, Amittai Axelrod, Alexandra Birch Mayne, Chris Callison-Burch, Miles Osborne, and David Talbot. 2005 · 2005
Earlier work this paper cites.
AER: Do we need to “improve” our alignments?
David Vilar, Maja Popović, and Hermann Ney. 2006 · 2006
Earlier work this paper cites.
Measuring word alignment quality for statistical machine translation
Alexander Fraser and Daniel Marcu. 2007 · 2007
Earlier work this paper cites.
Parallel corpora for medium density languages
Dániel Varga, Péter Halácsy, András Kornai, Viktor Nagy, László Németh, and Viktor Trón. 2007 · 2007
Earlier work this paper cites.
Parallel implementations of word alignment tool
Qin Gao and Stephan Vogel. 2008 · 2008
Earlier work this paper cites.
Improved statistical machine translation using monolingually-derived paraphrases
Yuval Marton, Chris Callison-Burch, and Philip Resnik. 2009 · 2009
Earlier work this paper cites.
Diversify and combine: Improving word alignment for machine translation on low-resource languages
Bing Xiang, Yonggang Deng, and Bowen Zhou. 2010a · 2010
Earlier work this paper cites.
Diversify and combine: Improving word alignment for machine translation on low-resource languages
Bing Xiang, Yonggang Deng, and Bowen Zhou. 2010b · 2010
Cited alongside, same era.
A simple, fast, and effective reparameterization of ibm model 2
Chris Dyer, Victor Chahuneau, and Noah A Smith. 2013 · 2013
Cited alongside, same era.
Multi-task word alignment triangulation for low-resource languages
Tomer Levinboim and David Chiang. 2015 · 2015
Cited alongside, same era.
Machine reading the primeros libros
Hannah Alpert-Abrams. 2016 · 2016
Cited alongside, same era.
Improving word alignment for low resource languages using English monolingual SRL
Meriem Beloucif, Markus Saers, and Dekai Wu. 2016a · 2016
Cited alongside, same era.
Incorporating structural alignment biases into an attentional neural translation model
Trevor Cohn, Cong Duy Vu Hoang, Ekaterina Vymolova, Kaisheng Yao, Chris Dyer, and Gholamreza Haffari. 2016 · 2016
Jointly learning to align and translate with transformer models
Sarthak Garg, Stephan Peitz, Udhyakumar Nallasamy, and Matthias Paulik. 2019 · 2019
Later among the works it cites.
ICDAR 2019 competition on post-OCR text correction
Christophe Rigaud, Antoine Doucet, Mickaël Coustaty, and Jean-Philippe Moreux. 2019 · 2019
Later among the works it cites.
Sparse transcription
Steven Bird. 2020 · 2020
Later among the works it cites.
No data to crawl? monolingual corpus creation from pdf files of truly low-resource languages in peru
Gina Bustamante, Arturo Oncevay, and Roberto Zariquiey. 2020 · 2020
Later among the works it cites.
Accurate word alignment induction from neural machine translation
Yun Chen, Yang Liu, Guanhua Chen, Xin Jiang, and Qun Liu. 2020 · 2020
Later among the works it cites.
A supervised word alignment method based on cross-language span prediction using multilingual BERT
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Phrase-based SMT for Finnish with more data, better models and alternative alignment and translation tools
Jörg Tiedemann, Fabienne Cap, Jenna Kanerva, Filip Ginter, Sara Stymne, Robert Östling, and Marion Weller-Di Marco. 2016 · 2016
Cited alongside, same era.
Supervised ocr error detection and correction using statistical and neural machine translation methods
Chantal Amrhein and Simon Clematide. 2018 · 2018
Cited alongside, same era.
Leveraging translations for speech transcription in low-resource settings
Antonios Anastasopoulos and David Chiang. 2018 · 2018
Cited alongside, same era.
Part-of-speech tagging on an endangered language: a parallel griko-italian resource
Antonios Anastasopoulos, Marika Lekakou, Josep Quer, Eleni Zimianiti, Justin DeBenedetto, and David Chiang. 2018 · 2018
Cited alongside, same era.
Using word vectors to improve word alignments for low resource machine translation
Nima Pourdamghani, Marjan Ghazvininejad, and Kevin Knight. 2018 · 2018
Cited alongside, same era.
Massively multilingual sentence embeddings for zero-shot cross-lingual transfer and beyond
Mikel Artetxe and Holger Schwenk. 2019 · 2019
Cited alongside, same era.
Masaaki Nagata, Katsuki Chousa, and Masaaki Nishino. 2020 · 2020
Later among the works it cites.
OCR Post Correction for Endangered Language Texts
Shruti Rijhwani, Antonios Anastasopoulos, and Graham Neubig. 2020 · 2020
Later among the works it cites.
Assessing the impact of OCR quality on downstream NLP tasks
Daniel Van Strien, Kaspar Beelen, Mariona Coll Ardanuy, Kasra Hosseini, Barbara McGillivray, and Giovanni Colavizza. 2020 · 2020
Later among the works it cites.
Mask-align: Self-supervised neural word alignment
Chi Chen, Maosong Sun, and Yang Liu. 2021 · 2021
Later among the works it cites.
Word alignment by fine-tuning embeddings on parallel corpora
Zi-Yi Dou and Graham Neubig. 2021 · 2021
Later among the works it cites.
When Being Unseen from mBERT is just the Beginning: Handling New Languages with Multilingual Language Models
Benjamin Muller, Antonios Anastasopoulos, Benoît Sagot, and Djamé Seddah. 2021 · 2021
Later among the works it cites.
Lexically aware semi-supervised learning for ocr post-correction
Shruti Rijhwani, Daisy Rosenblum, Antonios Anastasopoulos, and Graham Neubig. 2021 · 2021
Later among the works it cites.
OCR Improves Machine Translation for Low-Resource Languages
Oana Ignat, Jean Maillard, Vishrav Chaudhary, and Francisco Guzmán. 2022 · 2022
Later among the works it cites.
Mirroralign: A super lightweight unsupervised word alignment model via cross-lingual contrastive learning
Di Wu, Liang Ding, Shuo Yang, and Mingyang Li. 2022 · 2022
Later among the works it cites.