Fetching the paper…
Reading the bibliography…
We present the HPLT (High Performance Language Technologies) language resources, a new massive multilingual dataset including both monolingual and bilingual corpora extracted from CommonCrawl and previously unused web crawls from the Internet Archive.
Identifying and filtering near-duplicate documents
Andrei Z Broder. 2000 · 2000
Earlier work this paper cites.
Introduction to the Special Issue on the Web as Corpus
Adam Kilgarriff and Gregory Grefenstette. 2003 · 2003
Earlier work this paper cites.
Mt-based sentence alignment for ocr-generated parallel texts
Rico Sennrich and Martin Volk. 2010 · 2010
Earlier work this paper cites.
Parallel data, tools and interfaces in OPUS
Jörg Tiedemann. 2012 · 2012
Earlier work this paper cites.
Bitextor’s participation in WMT’16: shared task on document alignment
Miquel Esplà-Gomis, Mikel Forcada, Sergio Ortiz-Rojas, and Jorge Ferrández-Tordera. 2016 · 2016
Earlier work this paper cites.
Sequence-level knowledge distillation
Yoon Kim and Alexander M. Rush. 2016 · 2016
Earlier work this paper cites.
Marian: Fast neural machine translation in C++
Marcin Junczys-Dowmunt, Roman Grundkiewicz, Tomasz Dwojak, Hieu Hoang, Kenneth Heafield, Tom Neckermann, Frank Seide, Ulrich Germann, Alham Fikri Aji, Nikolay Bogoychev, et al. 2018 · 2018
Earlier work this paper cites.
Asynchronous pipeline for processing huge corpora on medium to low resource infrastructures
Pedro Javier Ortiz Suárez, Benoît Sagot, and Laurent Romary. 2019 · 2019
Earlier work this paper cites.
Paracrawl: Web-scale acquisition of parallel corpora
Marta Bañón, Pinzhen Chen, Barry Haddow, Kenneth Heafield, Hieu Hoang, Miquel Esplà-Gomis, Mikel L Forcada, Amir Kamran, Faheem Kirefu, Philipp Koehn, et al. 2020 · 2020
Earlier work this paper cites.
Bifixer and bicleaner: two open-source tools to clean your parallel data
Gema Ramírez-Sánchez, Jaume Zaragoza-Bernabeu, Marta Bañón, and Sergio Ortiz-Rojas. 2020 · 2020
Earlier work this paper cites.
The pile: An 800gb dataset of diverse text for language modeling
Leo Gao, Stella Biderman, Sid Black, Laurence Golding, Travis Hoppe, Charles Foster, Jason Phang, Horace He, Anish Thite, Noa Nabeshima, et al. 2020 · 2020
Cited alongside, same era.
Documenting large webtext corpora: A case study on the colossal clean crawled corpus
Jesse Dodge, Maarten Sap, Ana Marasović, William Agnew, Gabriel Ilharco, Dirk Groeneveld, Margaret Mitchell, and Matt Gardner. 2021 · 2021
Cited alongside, same era.
Ccmatrix: Mining billions of high-quality parallel sentences on the web
Holger Schwenk, Guillaume Wenzek, Sergey Edunov, Édouard Grave, Armand Joulin, and Angela Fan. 2021 · 2021
Cited alongside, same era.
mt5: A massively multilingual pre-trained text-to-text transformer
Linting Xue, Noah Constant, Adam Roberts, Mihir Kale, Rami Al-Rfou, Aditya Siddhant, Aditya Barua, and Colin Raffel. 2021 · 2021
Cited alongside, same era.
No language left behind: Scaling human-centered machine translation
Marta R Costa-jussà, James Cross, Onur Çelebi, Maha Elbayad, Kenneth Heafield, Kevin Heffernan, Elahe Kalbassi, Janice Lam, Daniel Licht, Jean Maillard, et al. 2022 · 2022
Later among the works it cites.
The bigscience roots corpus: A 1.6 tb composite multilingual dataset
Hugo Laurençon, Lucile Saulnier, Thomas Wang, Christopher Akiki, Albert Villanova del Moral, Teven Le Scao, Leandro Von Werra, Chenghao Mou, Eduardo González Ponferrada, Huu Nguyen, et al. 2022 · 2022
Later among the works it cites.
HPLT: High performance language technologies
Mikko Aulamo, Nikolay Bogoychev, Shaoxiong Ji, Graeme Nail, Gema Ramírez-Sánchez, Jörg Tiedemann, Jelmer van der Linde, and Jaume Zaragoza. 2023 · 2023
Later among the works it cites.
Get to know your parallel data: Performing english variety and genre classification over macocu corpora
Taja Kuzman, Peter Rupnik, and Nikola Ljubešić. 2023 · 2023
Later among the works it cites.
Gnu parallel 20230122 (’bolsonaristas’)
Ole Tange. 2023 · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
WuDaoCorpora: A super large-scale Chinese corpora for pre-training language models
Sha Yuan and Hanyu Zhao and Zhengxiao Du and Ming Ding and Xiao Liu and Yukuo Cen and Xu Zou and Zhilin Yang and Jie Tang. 2021 · 2021
Cited alongside, same era.
Towards a cleaner document-oriented multilingual crawled corpus
Julien Abadji, Pedro Ortiz Suarez, Laurent Romary, and Benoît Sagot. 2022 · 2022
Cited alongside, same era.
Building machine translation systems for the next thousand languages
Ankur Bapna, Isaac Caswell, Julia Kreutzer, Orhan Firat, Daan van Esch, Aditya Siddhant, Mengmeng Niu, Pallavi Baljekar, Xavier Garcia, Wolfgang Macherey, et al. 2022 · 2022
Cited alongside, same era.
Quality at a glance: An audit of web-crawled multilingual datasets
Julia Kreutzer, Isaac Caswell, Lisa Wang, Ahsan Wahab, Daan van Esch, Nasanbayar Ulzii-Orshikh, Allahsera Tapo, Nishant Subramani, Artem Sokolov, Claytone Sikasote, et al. 2022 · 2022
Cited alongside, same era.
Towards better structured and less noisy web data: Oscar with register annotations
Veronika Laippala, Anna Salmela, Samuel Rönnqvist, Alham Fikri Aji, Li-Hsin Chang, Asma Dhifallah, Larissa Goulart, Henna Kortelainen, Marc Pàmies, Deise Prina Dutra, et al. 2022 · 2022
Cited alongside, same era.
Bicleaner AI: Bicleaner goes neural
Jaume Zaragoza-Bernabeu, Gema Ramírez-Sánchez, Marta Bañón, and Sergio Ortiz Rojas. 2022 · 2022
Cited alongside, same era.
Later among the works it cites.
SERENGETI: Massively Multilingual Language Models for Africa
Adebara, Ife and Elmadany, AbdelRahim and Abdul-Mageed, Muhammad and Alcoba Inciarte, Alcides. 2023 · 2023
Later among the works it cites.
Glot500: Scaling Multilingual Corpora and Language Models to 500 Languages
ImaniGooghari, Ayyoob and Lin, Peiqin and Kargaran, Amir Hossein and Severini, Silvia and Jalili Sabet, Masoud and Kassner, Nora and Ma, Chunlan and Schmid, Helmut and Martins, André and Yvon, François and Schütze, Hinrich. 2023 · 2023
Later among the works it cites.
MADLAD-400: A Multilingual And Document-Level Large Audited Dataset
Kudugunta, Sneha and Caswell, Isaac and Zhang, Biao and Garcia, Xavier and Choquette-Choo, Christopher A and Lee, Katherine and Xin, Derrick and Kusupati, Aditya and Stella, Romi and Bapna, Ankur and others. 2023 · 2023
Later among the works it cites.
CulturaX: A Cleaned, Enormous, and Multilingual Dataset for Large Language Models in 167 Languages
Nguyen, Thuat and Van Nguyen, Chien and Lai, Viet Dac and Man, Hieu and Ngo, Nghia Trung and Dernoncourt, Franck and Rossi, Ryan A and Nguyen, Thien Huu. 2023 · 2023
Later among the works it cites.
Fastspell: the Langid Magic Spell
Marta Bañón, Jaume Zaragoza-Bernabeu, Gema Ramírez-Sánchez, and Sergio Ortiz-Rojas. 2024 · 2024
Closest in time.