Fetching the paper…
Reading the bibliography…
Data availability limits the scope of any given task.
A program for aligning sentences in bilingual corpora
William A. Gale and Kenneth W. Church. 1993 · 1993
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
Longformer: The long-document transformer
Iz Beltagy, Matthew E. Peters, and Arman Cohan. 2020 · 2004
Earlier work this paper cites.
Statistical significance tests for machine translation evaluation
Philipp Koehn. 2004 · 2004
Earlier work this paper cites.
Europarl: A parallel corpus for statistical machine translation
Philipp Koehn. 2005 · 2005
Earlier work this paper cites.
Parallel corpora for medium density languages
Dániel Varga Varga, Péter Halácsy, András Kornai, Viktor Nagy, László Németh, and Viktor Trón. 2005 · 2005
Earlier work this paper cites.
N-gram-based machine translation
José Mariño, Rafael E. Banchs, Josep M. Crego, Adrià de Gispert, Patrik Lambert, José A. R. Fonollosa, and Marta R. Costa-jussà. 2006 · 2006
Earlier work this paper cites.
Moses: Open source toolkit for statistical machine translation
Philipp Koehn, Hieu Hoang, Alexandra Birch, Chris Callison-Burch, Marcello Federico, Nicola Bertoldi, Brooke Cowan, Wade Shen, Christine Moran, Richard Zens, Chris Dyer, Ondřej Bojar, Alexandra Constantin, and Evan Herbst. 2007 · 2007
Earlier work this paper cites.
News from OPUS - A Collection of Multilingual Parallel Corpora with Tools and Interfaces , volume V, pages 237–248
Jörg Tiedemann. 2009 · 2009
Earlier work this paper cites.
Beyond english-centric multilingual machine translation
Angela Fan, Shruti Bhosale, Holger Schwenk, Zhiyi Ma, Ahmed El-Kishky, Siddharth Goyal, Mandeep Baines, Onur Celebi, Guillaume Wenzek, Vishrav Chaudhary, Naman Goyal, Tom Birch, Vitaliy Liptchinsky, Sergey Edunov, Edouard Grave, Michael Auli, and Armand Joulin. 2020 · 2010
Earlier work this paper cites.
Iterative, MT-based sentence alignment of parallel texts
Rico Sennrich and Martin Volk. 2011 · 2011
Earlier work this paper cites.
OpenSubtitles2016: Extracting large parallel corpora from movie and TV subtitles
Pierre Lison and Jörg Tiedemann. 2016 · 2016
Earlier work this paper cites.
Learning joint multilingual sentence representations with neural machine translation
Holger Schwenk and Matthijs Douze. 2017 · 2017
Earlier work this paper cites.
Neural machine translation with extended context
Jörg Tiedemann and Yves Scherrer. 2017 · 2017
Earlier work this paper cites.
Evaluating discourse phenomena in neural machine translation
Rachel Bawden, Rico Sennrich, Alexandra Birch, and Barry Haddow. 2018 · 2018
Earlier work this paper cites.
Marian: Fast neural machine translation in C++
Marcin Junczys-Dowmunt, Roman Grundkiewicz, Tomasz Dwojak, Hieu Hoang, Kenneth Heafield, Tom Neckermann, Frank Seide, Ulrich Germann, Alham Fikri Aji, Nikolay Bogoychev, André F. T. Martins, and Alexandra Birch. 2018 · 2018
Earlier work this paper cites.
SentencePiece: A simple and language independent subword tokenizer and detokenizer for neural text processing
Taku Kudo and John Richardson. 2018 · 2018
Earlier work this paper cites.
Document-level neural machine translation with hierarchical attention networks
Lesly Miculicich, Dhananjay Ram, Nikolaos Pappas, and James Henderson. 2018 · 2018
Cited alongside, same era.
A large-scale test set for the evaluation of context-aware pronoun translation in neural machine translation
Mathias Müller, Annette Rios, Elena Voita, and Rico Sennrich. 2018 · 2018
Cited alongside, same era.
A call for clarity in reporting BLEU scores
Matt Post. 2018 · 2018
Cited alongside, same era.
Context-aware neural machine translation learns anaphora resolution
Elena Voita, Pavel Serdyukov, Rico Sennrich, and Ivan Titov. 2018 · 2018
Cited alongside, same era.
Two new evaluation datasets for low-resource machine translation: Nepali-english and sinhala-english
Francisco Guzmán, Peng-Jen Chen, Myle Ott, Juan Pino, Guillaume Lample, Philipp Koehn, Vishrav Chaudhary, and Marc’Aurelio Ranzato. 2019 · 2019
Cited alongside, same era.
BlonDe: An automatic evaluation metric for document-level machine translation
Yuchen Jiang, Tianyu Liu, Shuming Ma, Dongdong Zhang, Jian Yang, Haoyang Huang, Rico Sennrich, Ryan Cotterell, Mrinmaya Sachan, and Ming Zhou. 2022 · 2022
Later among the works it cites.
No language left behind: Scaling human-centered machine translation
James Cross Onur Çelebi Maha Elbayad Kenneth Heafield Kevin Heffernan Elahe Kalbassi Janice Lam Daniel Licht Jean Maillard Anna Sun Skyler Wang Guillaume Wenzek Al Youngblood Bapi Akula Loic Barrault Gabriel Mejia Gonzalez Prangthip Hansanti John Hoffman Semarley Jarrett Kaushik Ram Sadagopan Dirk Rowe Shannon Spruit Chau Tran Pierre Andrews Necip Fazil Ayan Shruti Bhosale Sergey Edunov Angela Fan Cynthia Gao Vedanuj Goswami Francisco Guzmán Philipp Koehn Alexandre Mourachko Christophe Ropers Safiyyah Saleem Holger Schwenk Jeff Wang NLLB Team, Marta R. Costa-jussà. 2022 · 2022
Later among the works it cites.
CometKiwi: IST-unbabel 2022 submission for the quality estimation shared task
Ricardo Rei, Marcos Treviso, Nuno M. Guerreiro, Chrysoula Zerva, Ana C Farinha, Christine Maroti, José G. C. de Souza, Taisiya Glushkova, Duarte Alves, Luisa Coheur, Alon Lavie, and André F. T. Martins. 2022 · 2022
Later among the works it cites.
Rethinking document-level neural machine translation
Zewei Sun, Mingxuan Wang, Hao Zhou, Chengqi Zhao, Shujian Huang, Jiajun Chen, and Lei Li. 2022 · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Hierarchical modeling of global context for document-level neural machine translation
Xin Tan, Longyin Zhang, Deyi Xiong, and Guodong Zhou. 2019 · 2019
Cited alongside, same era.
Vecalign: Improved sentence alignment in linear time and space
Brian Thompson and Philipp Koehn. 2019 · 2019
Cited alongside, same era.
When a good translation is wrong in context: Context-aware machine translation improves on deixis, ellipsis, and lexical cohesion
Elena Voita, Rico Sennrich, and Ivan Titov. 2019 · 2019
Cited alongside, same era.
ParaCrawl: Web-scale acquisition of parallel corpora
Marta Bañón, Pinzhen Chen, Barry Haddow, Kenneth Heafield, Hieu Hoang, Miquel Esplà-Gomis, Mikel L. Forcada, Amir Kamran, Faheem Kirefu, Philipp Koehn, Sergio Ortiz Rojas, Leopoldo Pla Sempere, Gema Ramírez-Sánchez, Elsa Sarrías, Marek Strelec, Brian Thompson, William Waites, Dion Wiggins, and Jaume Zaragoza. 2020 · 2020
Cited alongside, same era.
CCAligned: A massive collection of cross-lingual web-document pairs
Ahmed El-Kishky, Vishrav Chaudhary, Francisco Guzmán, and Philipp Koehn. 2020a · 2020
Cited alongside, same era.
CCAligned: A massive collection of cross-lingual web-document pairs
Ahmed El-Kishky, Vishrav Chaudhary, Francisco Guzmán, and Philipp Koehn. 2020b · 2020
Cited alongside, same era.
Document-level neural MT: A systematic comparison
António Lopes, M. Amin Farajian, Rachel Bawden, Michael Zhang, and André F. T. Martins. 2020 · 2020
Cited alongside, same era.
Embarrassingly easy document-level MT metrics: How to convert any pretrained metric into a document-level metric
Giorgos Vernikos, Brian Thompson, Prashant Mathur, and Marcello Federico. 2022 · 2022
Later among the works it cites.
Exploring paracrawl for document-level neural machine translation
Yusser Al Ghussin, Jingyi Zhang, and Josef van Genabith. 2023 · 2023
Later among the works it cites.
Findings of the 2023 conference on machine translation (WMT23): LLMs are here but not quite there yet
Tom Kocmi, Eleftherios Avramidis, Rachel Bawden, Ondřej Bojar, Anton Dvorkovich, Christian Federmann, Mark Fishel, Markus Freitag, Thamme Gowda, Roman Grundkiewicz, Barry Haddow, Philipp Koehn, Benjamin Marie, Christof Monz, Makoto Morishita, Kenton Murray, Makoto Nagata, Toshiaki Nakazawa, Martin Popel, Maja Popović, and Mariya Shmatova. 2023 · 2023
Later among the works it cites.
There’s no data like better data: Using QE metrics for MT data filtering
Jan-Thorsten Peter, David Vilar, Daniel Deutsch, Mara Finkelstein, Juraj Juraska, and Markus Freitag. 2023 · 2023
Later among the works it cites.
Document-level language models for machine translation
Frithjof Petrick, Christian Herold, Pavel Petrushkov, Shahram Khadivi, and Hermann Ney. 2023 · 2023
Later among the works it cites.
Sotastream: A streaming approach to machine translation training
Matt Post, Thamme Gowda, Roman Grundkiewicz, Huda Khayrallah, Rohit Jain, and Marcin Junczys-Dowmunt. 2023 · 2023
Later among the works it cites.
Escaping the sentence-level paradigm in machine translation
Matt Post and Marcin Junczys-Dowmunt. 2023 · 2023
Later among the works it cites.
Slide: Reference-free evaluation for machine translation using a sliding document window
Vikas Raunak, Tom Kocmi, and Matt Post. 2023 · 2023
Later among the works it cites.
ChatGPT MT: Competitive for high- (but not low-) resource languages
Nathaniel Robinson, Perez Ogayo, David R. Mortensen, and Graham Neubig. 2023 · 2023
Later among the works it cites.
Document-level machine translation with large language models
Longyue Wang, Chenyang Lyu, Tianbo Ji, Zhirui Zhang, Dian Yu, Shuming Shi, and Zhaopeng Tu. 2023 · 2023
Later among the works it cites.
Identifying context-dependent translations for evaluation set production
Rachel Wicks and Matt Post. 2023 · 2023
Later among the works it cites.
A shocking amount of the web is machine translated: Insights from multi-way parallelism
Brian Thompson, Mehak Preet Dhaliwal, Peter Frisch, Tobias Domhan, and Marcello Federico. 2024 · 2024
Closest in time.