Fetching the paper…
Reading the bibliography…
Document alignment aims to identify pairs of documents in two distinct languages that are of comparable content or translations of each other.
Cross-lingual language model pretraining
Guillaume Lample and Alexis Conneau. 2019 · 1901
Earlier work this paper cites.
A massive collection of cross-lingual web-document pairs
Ahmed El-Kishky, Vishrav Chaudhary, Francisco Guzman, and Philipp Koehn. 2019 · 1911
Earlier work this paper cites.
A new measure of rank correlation
Maurice G Kendall. 1938 · 1938
Earlier work this paper cites.
Algorithms for the assignment and transportation problems
James Munkres. 1957 · 1957
Earlier work this paper cites.
An ir approach for translating new words from nonparallel, comparable texts
Pascale Fung and Lo Yuen Yee. 1998 · 1998
Earlier work this paper cites.
A metric for distributions with applications to image databases
Yossi Rubner, Carlo Tomasi, and Leonidas J Guibas. 1998 · 1998
Earlier work this paper cites.
Automatic identification of word translations from unrelated english and german corpora
Reinhard Rapp. 1999 · 1999
Earlier work this paper cites.
Mining the web for bilingual text
Philip Resnik. 1999 · 1999
Earlier work this paper cites.
Parallel web text mining for cross-language ir
Jiang Chen and Jian-Yun Nie. 2000 · 2000
Earlier work this paper cites.
Europarl: A multilingual corpus for evaluation of machine translation
Philipp Koehn et al. 2002 · 2002
Earlier work this paper cites.
Using tf-idf to determine word relevance in document queries
Juan Ramos et al. 2003 · 2003
Earlier work this paper cites.
The web as a parallel corpus
Philip Resnik and Noah A Smith. 2003 · 2003
Earlier work this paper cites.
Understanding inverse document frequency: on theoretical arguments for idf
Stephen Robertson. 2004 · 2004
Cited alongside, same era.
Europarl: A parallel corpus for statistical machine translation
Philipp Koehn. 2005 · 2005
Cited alongside, same era.
Improving machine translation performance by exploiting non-parallel corpora
Dragos Stefan Munteanu and Daniel Marcu. 2005 · 2005
Cited alongside, same era.
Extracting parallel sub-sentential fragments from non-parallel corpora
Dragos Stefan Munteanu and Daniel Marcu. 2006 · 2006
Cited alongside, same era.
On the use of comparable corpora to improve smt performance
Sadaf AbduI-Rauf and Holger Schwenk. 2009 · 2009
Cited alongside, same era.
Mining a comparable text corpus for a vietnamese-french statistical machine translation system
Thi-Ngoc-Diep Do, Viet-Bac Le, Brigitte Bigi, Laurent Besacier, and Eric Castelli. 2009 · 2009
First steps towards coverage-based document alignment
Luís Gomes and Gabriel Pereira Lopes. 2016 · 2016
Later among the works it cites.
Supervised word mover’s distance
Gao Huang, Chuan Guo, Matt J Kusner, Yu Sun, Fei Sha, and Kilian Q Weinberger. 2016 · 2016
Later among the works it cites.
Bad luc@ wmt 2016: a bilingual document alignment platform based on lucene
Laurent Jakubina and Phillippe Langlais. 2016 · 2016
Later among the works it cites.
English-french document alignment based on keywords and statistical translation
Marek Medveď, Miloš Jakubícek, and Vojtech Kovár. 2016 · 2016
Later among the works it cites.
The ilsp/arc submission to the wmt 2016 bilingual document alignment shared task
Vassilis Papavassiliou, Prokopis Prokopidis, and Stelios Piperidis. 2016 · 2016
Later among the works it cites.
Word clustering approach to bilingual document alignment (wmt 2016 shared task)
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
United nations general assembly resolutions: A six-language parallel corpus
Alexandre Rafalovitch, Robert Dale, et al. 2009 · 2009
Cited alongside, same era.
Mint: A method for effective and scalable mining of named entity transliterations from large comparable corpora
Raghavendra Udupa, K Saravanan, A Kumaran, and Jagadeesh Jagarlamudi. 2009 · 2009
Cited alongside, same era.
From word embeddings to document distances
Matt Kusner, Yu Sun, Nicholas Kolkin, and Kilian Weinberger. 2015 · 2015
Cited alongside, same era.
Findings of the wmt 2016 bilingual document alignment shared task
Christian Buck and Philipp Koehn. 2016a · 2016
Cited alongside, same era.
Yoda system for wmt16 shared task: Bilingual document alignment
Aswarth Abhilash Dara and Yiu-Chang Lin. 2016 · 2016
Cited alongside, same era.
Bitextor’s participation in wmt’16: shared task on document alignment
Miquel Esplà-Gomis, Mikel Forcada, Sergio Ortiz Rojas, and Jorge Ferrández-Tordera. 2016 · 2016
Cited alongside, same era.
Vadim Shchukin, Dmitry Khristich, and Irina Galinskaya. 2016 · 2016
Later among the works it cites.
The United Nations parallel corpus v1. 0
Michał Ziemski, Marcin Junczys-Dowmunt, and Bruno Pouliquen. 2016 · 2016
Later among the works it cites.
Linear-complexity relaxed word mover’s distance with gpu acceleration
Kubilay Atasu, Thomas Parnell, Celestine Dünner, Manolis Sifalakis, Haralampos Pozidis, Vasileios Vasileiadis, Michail Vlachos, Cesar Berrospi, and Abdel Labbi. 2017 · 2017
Later among the works it cites.
Cross-lingual document retrieval using regularized wasserstein distance
Georgios Balikas, Charlotte Laclau, Ievgen Redko, and Massih-Reza Amini. 2018 · 2018
Later among the works it cites.
Massively multilingual sentence embeddings for zero-shot cross-lingual transfer and beyond
Mikel Artetxe and Holger Schwenk. 2019 · 2019
Later among the works it cites.
Sentence mover’s similarity: Automatic evaluation for multi-sentence texts
Elizabeth Clark, Asli Celikyilmaz, and Noah A Smith. 2019 · 2019
Later among the works it cites.
Hierarchical document encoder for parallel corpus mining
Mandy Guo, Yinfei Yang, Keith Stevens, Daniel Cer, Heming Ge, Yun-hsuan Sung, Brian Strope, and Ray Kurzweil. 2019 · 2019
Later among the works it cites.