Fetching the paper…
Reading the bibliography…
Mining high-quality bitexts for low-resource languages is challenging.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
Toward accurate dynamic time warping in linear time and space
Stan Salvador and Philip Chan. 2007 · 2007
Earlier work this paper cites.
Parallel corpora for medium density languages
Dániel Varga, Péter Halácsy, András Kornai, Nagy Viktor, Nagy Laszlo, Németh László, and Tron Viktor. 2007 · 2007
Earlier work this paper cites.
Hubs in space: Popular nearest neighbors in high-dimensional data
Miloš Radovanović, Alexandros Nanopoulos, and Mirjana Ivanović. 2010 · 2010
Earlier work this paper cites.
MT-based sentence alignment for OCR-generated parallel texts
Rico Sennrich and Martin Volk. 2010 · 2010
Earlier work this paper cites.
Hubness and pollution: Delving into cross-space mapping for zero-shot learning
Angeliki Lazaridou, Georgiana Dinu, and Marco Baroni. 2015 · 2015
Earlier work this paper cites.
Findings of the WMT 2016 bilingual document alignment shared task
Christian Buck and Philipp Koehn. 2016 · 2016
Earlier work this paper cites.
Efficient natural language response suggestion for smart reply
Matthew Henderson, Rami Al-Rfou, Brian Strope, Yun-hsuan Sung, Laszlo Lukacs, Ruiqi Guo, Sanjiv Kumar, Balint Miklos, and Ray Kurzweil. 2017 · 2017
Earlier work this paper cites.
Overview of the second BUCC shared task: Spotting parallel sentences in comparable corpora
Pierre Zweigenbaum, Serge Sharoff, and Reinhard Rapp. 2017 · 2017
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018 · 2018
Cited alongside, same era.
Dual conditional cross-entropy filtering of noisy parallel corpora
Marcin Junczys-Dowmunt. 2018 · 2018
Cited alongside, same era.
Findings of the WMT 2018 shared task on parallel corpus filtering
Philipp Koehn, Huda Khayrallah, Kenneth Heafield, and Mikel L. Forcada. 2018 · 2018
Cited alongside, same era.
Overview of the Third BUCC Shared Task: Spotting Parallel Sentences in Comparable Corpora
Pierre Zweigenbaum, Serge Sharoff, and Reinhard Rapp. 2018 · 2018
Cited alongside, same era.
Unsupervised cross-lingual representation learning at scale
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 2019
Later among the works it cites.
Sentence-BERT: Sentence embeddings using Siamese BERT-networks
Nils Reimers and Iryna Gurevych. 2019 · 2019
Later among the works it cites.
Vecalign: Improved sentence alignment in linear time and space
Brian Thompson and Philipp Koehn. 2019 · 2019
Later among the works it cites.
Filtering noisy parallel corpus using transformers with proxy task learning
Haluk Açarçiçek, Talha Çolakoğlu, Pınar Ece Aktan Hatipoğlu, Chong Hsuan Huang, and Wei Peng. 2020 · 2020
Later among the works it cites.
A simple framework for contrastive learning of visual representations
Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton. 2020 · 2020
Later among the works it cites.
Findings of the WMT 2020 shared task on parallel corpus filtering and alignment
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 2019
Cited alongside, same era.
Billion-scale similarity search with GPUs
Jeff Johnson, Matthijs Douze, and Hervé Jégou. 2019 · 2019
Cited alongside, same era.
Findings of the WMT 2019 shared task on parallel corpus filtering for low-resource conditions
Philipp Koehn, Francisco Guzmán, Vishrav Chaudhary, and Juan Pino. 2019 · 2019
Cited alongside, same era.
Margin-based parallel corpus mining with multilingual sentence embeddings
Mikel Artetxe and Holger Schwenk. 2019a
Cited in the paper.
Massively multilingual sentence embeddings for zero-shot cross-lingual transfer and beyond
Mikel Artetxe and Holger Schwenk. 2019b
Cited in the paper.
Philipp Koehn, Vishrav Chaudhary, Ahmed El-Kishky, Naman Goyal, Peng-Jen Chen, and Francisco Guzmán. 2020 · 2020
Later among the works it cites.
Alibaba submission to the WMT20 parallel corpus filtering task
Jun Lu, Xin Ge, Yangbin Shi, and Yuqi Zhang. 2020 · 2020
Later among the works it cites.