Fetching the paper…
Reading the bibliography…
This paper presents a comprehensive overview of the data preparation pipeline developed for the OpenGPT-X project, a large-scale initiative aimed at creating open and high-performance multilingual large language models (LLMs).
Space/time trade-offs in hash coding with allowable errors
Burton H. Bloom · 1970
Earlier work this paper cites.
Suffix arrays: A new method for on-line string searches
Udi Manber and Gene Myers · 1993
Earlier work this paper cites.
On the resemblance and containment of documents
Andrei Broder · 1997
Earlier work this paper cites.
Similarity estimation techniques from rounding algorithms
Moses S Charikar · 2002
Earlier work this paper cites.
Europarl: A parallel corpus for statistical machine translation
P. Koehn · 2005
Earlier work this paper cites.
Detecting near-duplicates for web crawling
Gurmeet Singh Manku, Arvind Jain, and Anish Das Sarma · 2007
Earlier work this paper cites.
The wacky wide web: A collection of very large linguistically processed web-crawled corpora
M. Baroni, S. Bernardini, A. Ferraresi, and E. Zanchetta · 2009
Earlier work this paper cites.
Policy-aware content reuse on the web
Oshani Seneviratne, Lalana Kagal, and Tim Berners-Lee · 2009
Earlier work this paper cites.
Multiun: A multilingual corpus from united nation documents
Andreas Eisele and Yu Chen · 2010
Earlier work this paper cites.
Kenlm: Faster and smaller language model queries
Kenneth Heafield · 2011
Earlier work this paper cites.
Building corpora for the philological study of swiss legal texts
Stefan Höfler and Michael Piotrowski · 2011
Earlier work this paper cites.
hrwac and slwac: Compiling web corpora for croatian and slovene
Nikola Ljubešić and Tomaž Erjavec · 2011
Earlier work this paper cites.
Parallel data, tools and interfaces in opus
J. Tiedemann · 2012
Earlier work this paper cites.
Parallel data, tools and interfaces in OPUS
Jörg Tiedemann · 2012
Earlier work this paper cites.
The PAISA corpus of Italian web texts
V. Lyding, E. Stemle, C. Borghetti, M. Brunello, S. Castagnoli, F. Dell’Orletta, H. Dittmann, A. Lenci, and V. Pirrelli · 2014
Earlier work this paper cites.
Dcep-digital corpus of the european parliament
Hajlaoui Najeh, Kolovratnik David, Vaeyrynen Jaakko, Steinberger Ralf, and Varga Dániel · 2014
Earlier work this paper cites.
C4corpus: Multilingual web-size corpus with free license
Ivan Habernal, Omnia Zayed, and Iryna Gurevych · 2016
Earlier work this paper cites.
FastText.zip: Compressing text classification models
Armand Joulin, Edouard Grave, Piotr Bojanowski, Matthijs Douze, Hérve Jégou, and Tomas Mikolov · 2016
Earlier work this paper cites.
Bag of tricks for efficient text classification
Armand Joulin, Edouard Grave, Piotr Bojanowski, and Tomas Mikolov · 2016
Earlier work this paper cites.
OpenSubtitles2016: Extracting large parallel corpora from movie and TV subtitles
Pierre Lison and Jörg Tiedemann · 2016
Earlier work this paper cites.
Learning word vectors for 157 languages
Edouard Grave, Piotr Bojanowski, Prakhar Gupta, Armand Joulin, and Tomas Mikolov · 2018
Earlier work this paper cites.
WikiHow: A large scale text summarization dataset, 2018
Mahnaz Koupaee and William Yang Wang · 2018
Earlier work this paper cites.
Subword regularization: Improving neural network translation models with multiple subword candidates
Taku Kudo · 2018
Earlier work this paper cites.
Polish parliamentary corpus
Maciej Ogrodniczuk · 2018
Earlier work this paper cites.
Openwebtext corpus
Aaron Gokaslan and Vanya Cohen · 2019
Earlier work this paper cites.
Modelling large parallel corpora: The Zurich parallel corpus collection
Johannes Graën, Tannon Kew, Anastassia Shaitarova, and Martin Volk · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al · 2019
Earlier work this paper cites.
CCNet: Extracting high quality monolingual datasets from web crawl data
Guillaume Wenzek, Marie-Anne Lachaux, Alexis Conneau, Vishrav Chaudhary, Francisco Guzmán, Armand Joulin, and Edouard Grave · 2019
Earlier work this paper cites.
German political speeches corpus, January 2020
Adrien Barbaresi · 2020
Earlier work this paper cites.
Language id in the wild: Unexpected challenges on the path to a thousand-language web text corpus
Isaac Caswell, Theresa Breiner, Daan van Esch, and Ankur Bapna · 2020
Earlier work this paper cites.
Corpus der Entscheidungen des Bundesarbeitsgerichts (CE-BAG), September 2020
Sean Fobbe · 2020
Earlier work this paper cites.
The Pile: An 800gb dataset of diverse text for language modeling
Leo Gao, Stella Biderman, Sid Black, Laurence Golding, Travis Hoppe, Charles Foster, Jason Phang, Horace He, Anish Thite, Noa Nabeshima, Shawn Presser, and Connor Leahy · 2020
Cited alongside, same era.
Data augmentation using pre-trained transformer models
Varun Kumar, Ashutosh Choudhary, and Eunah Cho · 2020
Cited alongside, same era.
The CEPS Eurlex dataset collection. TRIGGER project, centre for european policy studies, 2020
Moritz Laurer and Camille Borrett · 2020
Cited alongside, same era.
A monolingual approach to contextualized word embeddings for mid-resource languages
Pedro Javier Ortiz Suárez, Laurent Romary, and Benoît Sagot · 2020
Cited alongside, same era.
Towards an open platform for legal information
Malte Ostendorff, Till Blume, and Saskia Ostendorff · 2020
Cited alongside, same era.
Pre-training data quality and quantity for a low-resource language: New corpus and BERT models for Maltese
Kurt Micallef, Albert Gatt, Marc Tanti, Lonneke van der Plas, and Claudia Borg · 2022
Later among the works it cites.
An empirical study on cross-x transfer for legal judgment prediction, 2022
Joel Niklaus, Matthias Stürmer, and Ilias Chalkidis · 2022
Later among the works it cites.
The elephant in the room: Analyzing the presence of big tech in natural language processing research
Mohamed Abdalla, Jan Philip Wahle, Terry Ruas, Aurelie Neveol, Fanny Ducel, Saif Mohammad, and Karen Fort · 2023
Later among the works it cites.
Redpajama: an open dataset for training large language models, October 2023
Together Computer · 2023
Later among the works it cites.
Redpajama: an open dataset for training large language models, 2023
Together Computer · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu · 2020
Cited alongside, same era.
Ungoliant: An optimized pipeline for the generation of a very large-scale multilingual web corpus
Julien Abadji, Pedro Javier Ortiz Suárez, Laurent Romary, and Benoît Sagot · 2021
Cited alongside, same era.
Trafilatura: A web scraping library and command-line tool for text discovery and extraction
Adrien Barbaresi · 2021
Cited alongside, same era.
The danish gigaword corpus
Leon Derczynski, Manuel R. Ciosici, Rebekah Baglini, Morten H. Christiansen, Jacob Aarup Dalsgaard, Riccardo Fusaroli, Peter Juel Henrichsen, Rasmus Hvingelby, Andreas Kirkedal, Alex Speed Kjeldsen, Claus Ladefoged, Finn Årup Nielsen, Jens Madsen, Malte Lau Petersen, Jonathan Hvithamar Rystrøm, and Daniel Varab · 2021
Cited alongside, same era.
Corpus der Drucksachen des Deutschen Bundestages (CDRS-BT), April 2021
Sean Fobbe · 2021
Cited alongside, same era.
Corpus der Plenarprotokolle des Deutschen Bundestages (CPP-BT), February 2021
Sean Fobbe · 2021
Cited alongside, same era.
The Pile: An 800GB dataset of diverse text for language modeling
Leo Gao, Stella Biderman, Sid Black, Laurence Golding, Travis Hoppe, Charles Foster, Jason Phang, Horace He, Anish Thite, Noa Nabeshima, Shawn Presser, and Connor Leahy · 2021
Cited alongside, same era.
Corpus der Entscheidungen des Bundesgerichtshofs (CE-BGH), March 2023
Sean Fobbe · 2023
Later among the works it cites.
Corpus der Entscheidungen des Bundespatentgerichts (CE-BPatG), April 2023
Sean Fobbe · 2023
Later among the works it cites.
Corpus der Entscheidungen des Bundesverwaltungsgerichts (CE-BVerwG), March 2023
Sean Fobbe · 2023
Later among the works it cites.
Corpus of Decisions: International Court of Justice (CD-ICJ), October 2023
Sean Fobbe · 2023
Later among the works it cites.
Corpus der Entscheidungen des Bundesfinanzhofs (CE-BFH), October 2023
Seán Fobbe · 2023
Later among the works it cites.
Gptscore: Evaluate as you desire
Jinlan Fu, See-Kiong Ng, Zhengbao Jiang, and Pengfei Liu · 2023
Later among the works it cites.
Suriya Gunasekar, Yi Zhang, Jyoti Aneja, Caio César Teodoro Mendes, Allie Del Giorno, Sivakanth Gopi, Mojan Javaheripi, Piero Kauffmann, Gustavo de Rosa, Olli Saarikivi, et al · 2023
Later among the works it cites.
Glot500: Scaling multilingual corpora and language models to 500 languages
Ayyoob ImaniGooghari, Peiqin Lin, Amir Hossein Kargaran, Silvia Severini, Masoud Jalili Sabet, Nora Kassner, Chunlan Ma, Helmut Schmid, André Martins, François Yvon, and Hinrich Schütze · 2023
Later among the works it cites.
Code for individual languages and language groups
ISO 639-2 · 2023
Later among the works it cites.
Large language models are state-of-the-art evaluators of translation quality
Tom Kocmi and Christian Federmann · 2023
Later among the works it cites.
Openwebmath: An open dataset of high-quality mathematical web text, 2023
Keiran Paster, Marco Dos Santos, Zhangir Azerbayev, and Jimmy Ba · 2023
Later among the works it cites.
Guilherme Penedo, Quentin Malartic, Daniel Hesslow, Ruxandra Cojocaru, Alessandro Cappelli, Hamza Alobeidli, Baptiste Pannier, Ebtesam Almazrouei, and Julien Launay · 2023
Later among the works it cites.
peS2o (pretraining efficiently on S2ORC) dataset
Luca Soldaini and Kyle Lo · 2023
Later among the works it cites.
Is ChatGPT a good NLG evaluator? a preliminary study
Jiaan Wang, Yunlong Liang, Fandong Meng, Zengkui Sun, Haoxiang Shi, Zhixu Li, Jinan Xu, Jianfeng Qu, and Jie Zhou · 2023
Later among the works it cites.
Tokenizer choice for LLM training: Negligible or crucial?
Mehdi Ali, Michael Fromm, Klaudia Thellmann, Richard Rutmann, Max Lübbering, Johannes Leveling, Katrin Klug, Jan Ebert, Niclas Doll, Jasper Schulze Buschhoff, Charvi Jain, Alexander Arno Weber, Lena Jurkschat, Hammam Abdelwahab, Chelsea John, Pedro Ortiz Suarez, Malte Ostendorff, Samuel Weinbach, Rafet Sifa, Stefan Kesselheim, and Nicolas Flores-Herr · 2024
Closest in time.
FastSpell: The LangId Magic Spell
Marta Bañón, Gema Ramírez-Sánchez, Jaume Zaragoza-Bernabeu, and Sergio Ortíz-Rojas · 2024
Closest in time.
A New Massive Multilingual Dataset for High-Performance Language Technologies
Ona de Gibert, Graeme Nail, Nikolay Arefyev, Marta Bañón, Jelmer van der Linde, Shaoxiong Ji, Jaume Zaragoza-Bernabeu, Mikko Aulamo, Gema Ramírez-Sánchez, Andrey Kutuzov, Sampo Pyysalo, Stephan Oepen, and Jörg Tiedemann · 2024
Closest in time.
Medical mt5: An open-source multilingual text-to-text llm for the medical domain, 2024
Iker García-Ferrero, Rodrigo Agerri, Aitziber Atutxa Salazar, Elena Cabrio, Iker de la Iglesia, Alberto Lavelli, Bernardo Magnini, Benjamin Molinet, Johana Ramirez-Romero, German Rigau, Jose Maria Villa-Gonzalez, Serena Villata, and Andrea Zaninello · 2024
Closest in time.
Mohammed Hassanin and Nour Moustafa · 2024
Closest in time.
Madlad-400: A multilingual and document-level large audited dataset
Sneha Kudugunta, Isaac Caswell, Biao Zhang, Xavier Garcia, Derrick Xin, Aditya Kusupati, Romi Stella, Ankur Bapna, and Orhan Firat · 2024
Closest in time.
Evaluation of Document Deduplication Algorithms for Large Text Corpora (to appear)
Johannes Leveling, Lennard Helmer, Benny Stein, Dennis Wegener, Zoha Sheikh, Elanton Fernandes, and Hammam Abdelwahab · 2024
Closest in time.
Progress report: Towards European LLMs, October 2024
OpenGPT-X · 2024
Closest in time.
The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale
Guilherme Penedo, Hynek Kydlíček, Loubna Ben allal, Anton Lozhkov, Margaret Mitchell, Colin Raffel, Leandro Von Werra, and Thomas Wolf · 2024
Closest in time.
Dolma: An open corpus of three trillion tokens for language model pretraining research
Luca Soldaini, Rodney Kinney, Akshita Bhagia, Dustin Schwenk, David Atkinson, Russell Authur, Ben Bogin, Khyathi Chandu, Jennifer Dumas, Yanai Elazar, et al · 2024
Closest in time.
Finetuned multimodal language models are high-quality image-text data filters
Weizhi Wang, Khalil Mrini, Linjie Yang, Sateesh Kumar, Yu Tian, Xifeng Yan, and Heng Wang · 2024
Closest in time.