Fetching the paper…
Reading the bibliography…
As language models grow ever larger, the need for large-scale high-quality text datasets has never been more pressing, especially in multilingual settings.
On the use of arxiv as a dataset
Clement, C. B., M. Bierbaum, K. P. O’Keeffe, and A. A. Alemi (2019) · 1905
Earlier work this paper cites.
The royal society corpus: From uncharted data to corpus
Kermes, H., S. Degaetano-Ortlieb, A. Khamis, J. Knappen, and E. Teich (2016, May) · 1931
Earlier work this paper cites.
Suffix arrays: A new method for on-line string searches
Manber, U. and G. Myers (1993) · 1993
Earlier work this paper cites.
Scaling laws for neural language models
Kaplan, J., S. McCandlish, T. Henighan, T. B. Brown, B. Chess, R. Child, S. Gray, A. Radford, J. Wu, and D. Amodei (2020) · 2001
Earlier work this paper cites.
Similarity estimation techniques from rounding algorithms
Charikar, M. S. (2002) · 2002
Earlier work this paper cites.
Improving massively multilingual neural machine translation and zero-shot translation
Zhang, B., P. Williams, I. Titov, and R. Sennrich (2020) · 2004
Earlier work this paper cites.
Ai4bharat-indicnlp corpus: Monolingual corpora and word embeddings for indic languages
Kunchukuttan, A., D. Kakwani, S. Golla, G. N.C., A. Bhattacharyya, M. M. Khapra, and P. Kumar (2020) · 2005
Earlier work this paper cites.
Corpus description of the ESTER evaluation campaign for the rich transcription of French broadcast news
Galliano, S., E. Geoffrois, G. Gravier, J.-F. Bonastre, D. Mostefa, and K. Choukri (2006, May) · 2006
Earlier work this paper cites.
Detecting near-duplicates for web crawling
Manku, G. S., A. Jain, and A. Das Sarma (2007) · 2007
Earlier work this paper cites.
Participation is not a design fix for machine learning
Sloane, M., E. Moss, O. Awomolo, and L. Forlano (2020) · 2007
Earlier work this paper cites.
The warc file format 1.0 (iso 28500)
Mohr, G., J. Kunze, and M. Stack (2008) · 2008
Earlier work this paper cites.
Development of Indonesian large vocabulary continuous speech recognition system within a-STAR project
Sakti, S., E. Kelana, H. Riza, S. Sakai, K. Markov, and S. Nakamura (2008) · 2008
Earlier work this paper cites.
Natural Language Processing with Python: Analyzing Text with the Natural Language Toolkit
Bird, S., E. Klein, and E. Loper (2009) · 2009
Earlier work this paper cites.
Resource report: Building parallel text corpora for multi-domain translation system
Budiono, H. Riza, and C. Hakim (2009, August) · 2009
Earlier work this paper cites.
Probabilistic part-of-speech tagging for bahasa indonesia
Pisceldo, F., R. Manurung, and M. Adriani (2009) · 2009
Earlier work this paper cites.
mt5: A massively multilingual pre-trained text-to-text transformer
Xue, L., N. Constant, A. Roberts, M. Kale, R. Al-Rfou, A. Siddhant, A. Barua, and C. Raffel (2020) · 2010
Earlier work this paper cites.
Behavioral use licensing for responsible ai
Contractor, D., D. McDuff, J. Haines, J. Lee, C. Hines, and B. Hecht (2020) · 2011
Earlier work this paper cites.
KenLM: Faster and smaller language model queries
Heafield, K. (2011, July) · 2011
Earlier work this paper cites.
WIT3: Web inventory of transcribed and translated talks
Cettolo, M., C. Girardi, and M. Federico (2012, May 28–30) · 2012
Earlier work this paper cites.
MultiUN v2: UN documents with multilingual alignments
Chen, Y. and A. Eisele (2012, May) · 2012
Earlier work this paper cites.
Building large monolingual dictionaries at the Leipzig corpora collection: From 100 to 200 languages
Goldhahn, D., T. Eckart, and U. Quasthoff (2012, May) · 2012
Earlier work this paper cites.
Usage of indonesian possessive verbal predicates: A statistical analysis based on questionnaire and storytelling surveys
Moeljadi, D. (2012) · 2012
Earlier work this paper cites.
The design and construction of the 50 million words ksucca king saud university corpus of classical arabic
Alrabiah, M., A. Alsalman, and E. Atwell (2013, 01) · 2013
Earlier work this paper cites.
LABR: A large scale Arabic book reviews dataset
Aly, M. and A. Atiya (2013, August) · 2013
Earlier work this paper cites.
Kalimat a multipurpose arabic corpus
El-Haj, M. and R. Koulali (2013) · 2013
Earlier work this paper cites.
The amara corpus: Building parallel language resources for the educational domain
Abdelali, A., F. Guzman, H. Sajjad, and S. Vogel (2014, may) · 2014
Earlier work this paper cites.
Urdu monolingual corpus
Jawaid, B., A. Kamran, and O. Bojar (2014) · 2014
Earlier work this paper cites.
1.5 billion words arabic corpus
El-Khair, I. A. (2016) · 2016
Earlier work this paper cites.
Opensubtitles2016: Extracting large parallel corpora from movie and tv subtitles
Lison, P. and J. Tiedemann (2016) · 2016
Earlier work this paper cites.
The United Nations parallel corpus v1.0
Ziemski, M., M. Junczys-Dowmunt, and B. Pouliquen (2016, May) · 2016
Earlier work this paper cites.
Indonesian news articles published at 2017
Ashari, A. (2018) · 2017
Earlier work this paper cites.
Bag of tricks for efficient text classification
Joulin, A., E. Grave, P. Bojanowski, and T. Mikolov (2017, April) · 2017
Earlier work this paper cites.
Déjàvu: a map of code duplicates on github
Lopes, C. V., P. Maj, P. Martins, V. Saini, D. Yang, J. Zitny, H. Sajnani, and J. Vitek (2017) · 2017
Earlier work this paper cites.
Do artifacts have politics?
Winner, L. (2017) · 2017
Earlier work this paper cites.
Tashkeela: Novel corpus of arabic vocalized texts, data for auto-diacritization systems
Zerrouki, T. and A. Balla (2017) · 2017
Earlier work this paper cites.
BERT: pre-training of deep bidirectional transformers for language understanding
Devlin, J., M. Chang, K. Lee, and K. Toutanova (2018) · 2018
Earlier work this paper cites.
An annotated huge dataset for standard and colloquial arabic reviews for subjective sentiment analysis
Elnagar, A., L. Lulu, and O. Einea (2018) · 2018
Earlier work this paper cites.
Dureader: a chinese machine reading comprehension dataset from real-world applications
He, W., K. Liu, J. Liu, Y. Lyu, S. Zhao, X. Xiao, Y. Liu, Y. Wang, H. Wu, Q. She, X. Liu, T. Wu, and H. Wang (2018) · 2018
Earlier work this paper cites.
Fine-tuned language models for text classification
Howard, J. and S. Ruder (2018) · 2018
Earlier work this paper cites.
Subword regularization: Improving neural network translation models with multiple subword candidates
Kudo, T. (2018, July) · 2018
Earlier work this paper cites.
The IIT Bombay English-Hindi parallel corpus
Kunchukuttan, A., P. Mehta, and P. Bhattacharyya (2018, May) · 2018
Cited alongside, same era.
Indosum: A new benchmark dataset for indonesian text summarization
Kurniawan, K. and S. Louvan (2018) · 2018
Cited alongside, same era.
Uit-vsfc: Vietnamese students’ feedback corpus for sentiment analysis
Nguyen, K. V., V. D. Nguyen, P. X. V. Nguyen, T. T. H. Truong, and N. L.-T. Nguyen (2018) · 2018
Cited alongside, same era.
Tufs asian language parallel corpus (talpco)
Nomoto, H., K. Okano, D. Moeljadi, and H. Sawada (2018) · 2018
Cited alongside, same era.
Datasets for aspect-based sentiment analysis in bangla and its baseline evaluation
Rahman, M., E. Kumar Dey, et al. (2018) · 2018
Cited alongside, same era.
Indonesian news corpus
Rahutomo, F. and A. Miqdad Muadz Muzad (2018) · 2018
Cited alongside, same era.
The values encoded in machine learning research
Birhane, A., P. Kalluri, D. Card, W. Agnew, R. Dotan, and M. Bao (2021) · 2021
Later among the works it cites.
Envisioning communities: A participatory approach towards AI for social good
Bondi, E., L. Xu, D. Acosta-Navas, and J. A. Killian (2021) · 2021
Later among the works it cites.
Binhvq news corpus
Bình, V. Q. (2021) · 2021
Later among the works it cites.
Tecla: Text classification catalan dataset
Carrino, C. P., C. G. Rodriguez-Penagos, and C. Armentano-Oller (2021, March) · 2021
Later among the works it cites.
Guiding principles for participatory design-inspired natural language processing
Caselli, T., R. Cibin, C. Conforti, E. Encinas, and M. Teli (2021) · 2021
Later among the works it cites.
Documenting large webtext corpora: A case study on the colossal clean crawled corpus
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ensuring fairness in machine learning to advance health equity
Rajkomar, A., M. Hardt, M. D. Howell, G. Corrado, and M. H. Chin (2018) · 2018
Cited alongside, same era.
JW300: A wide-coverage parallel corpus for low-resource languages
Agić, Ž. and I. Vulić (2019, July) · 2019
Cited alongside, same era.
The adverse effects of code duplication in machine learning models of code
Allamanis, M. (2019) · 2019
Cited alongside, same era.
Studying the history of the arabic language: language technology and a large-scale historical corpus
Belinkov, Y., A. Magidow, A. Barrón-Cedeño, A. Shmidman, and M. Romanov (2019) · 2019
Cited alongside, same era.
The secret sharer: Evaluating and testing unintended memorization in neural networks
Carlini, N., C. Liu, Ú. Erlingsson, J. Kos, and D. Song (2019) · 2019
Cited alongside, same era.
Sanad: Single-label arabic news articles dataset for automatic text categorization
Einea, O., A. Elnagar, and R. Al Debsi (2019) · 2019
Cited alongside, same era.
Dodge, J., M. Sap, A. Marasović, W. Agnew, G. Ilharco, D. Groeneveld, M. Mitchell, and M. Gardner (2021) · 2021
Later among the works it cites.
Beyond english-centric multilingual machine translation
Fan, A., S. Bhosale, H. Schwenk, Z. Ma, A. El-Kishky, S. Goyal, M. Baines, O. Celebi, G. Wenzek, V. Chaudhary, et al. (2021) · 2021
Later among the works it cites.
Data governance in the age of large-scale data-driven language technology
Jernite, Y., H. Nguyen, S. Biderman, A. Rogers, M. Masoud, V. Danchev, S. Tan, A. S. Luccioni, N. Subramani, G. Dupont, J. Dodge, K. Lo, Z. Talat, D. Radev, A. Gokaslan, S. Nikpoor, P. Henderson, R. Bommasani, and M. Mitchell (2022) · 2021
Later among the works it cites.
Banglalm: Bangla corpus for language model research
Kowsher, M., M. Uddin, A. Tahabilder, M. Ruhul Amin, M. F. Shahriar, and M. S. I. Sobuj (2021, September) · 2021
Later among the works it cites.
ParlamentParla - Speech corpus of Catalan Parliamentary sessions
Külebi, B. (2021, October) · 2021
Later among the works it cites.
The hard problem of aligning AI to human values
Leahy, C. and S. Biderman (2021) · 2021
Later among the works it cites.
Datasets: A community library for natural language processing
Lhoest, Q., A. Villanova del Moral, Y. Jernite, A. Thakur, P. von Platen, S. Patil, J. Chaumond, M. Drame, J. Plu, L. Tunstall, J. Davison, M. Šaško, G. Chhablani, B. Malik, S. Brandeis, T. Le Scao, V. Sanh, C. Xu, N. Patry, A. McMillan-Major, P. Schmid, S. Gugger, C. Delangue, T. Matussière, L. Debut, S. Bekman, P. Cistac, T. Goehringer, V. Mustar, F. Lagunas, A. Rush, and T. Wolf (2021, November) · 2021
Later among the works it cites.
Jurassic-1: Technical details and evaluation
Lieber, O., O. Sharir, B. Lenz, and Y. Shoham (2021) · 2021
Later among the works it cites.
What’s in the box? a preliminary analysis of undesirable content in the common crawl corpus
Luccioni, A. S. and J. D. Viviano (2021) · 2021
Later among the works it cites.
Indonli: A natural language inference dataset for indonesian
Mahendra, R., A. F. Aji, S. Louvan, F. Rahman, and C. Vania (2021) · 2021
Later among the works it cites.
Styled augmented translation (sat)
Ngo, C. and T. H. Trinh (2021) · 2021
Later among the works it cites.
Vietnamese poem generator
Nguyen, T., H. Pham, M. Truong, H. Duc, and P. Tan (2021) · 2021
Later among the works it cites.
Scaling language models: Methods, analysis & insights from training gopher
Rae, J. W., S. Borgeaud, T. Cai, K. Millican, J. Hoffmann, H. F. Song, J. Aslanides, S. Henderson, R. Ring, S. Young, E. Rutherford, T. Hennigan, J. Menick, A. Cassirer, R. Powell, G. van den Driessche, L. A. Hendricks, M. Rauh, P. Huang, A. Glaese, J. Welbl, S. Dathathri, S. Huang, J. Uesato, J. Mellor, I. Higgins, A. Creswell, N. McAleese, A. Wu, E. Elsen, S. M. Jayakumar, E. Buchatskaya, D. Budden, E. Sutherland, K. Simonyan, M. Paganini, L. Sifre, L. Martens, X. L. Li, A. Kuncoro, A. Nematzadeh, E. Gribovskaya, D. Donato, A. Lazaridou, A. Mensch, J. Lespiau, M. Tsimpoukelli, N. Grigorev, D. Fritz, T. Sottiaux, M. Pajarskas, T. Pohlen, Z. Gong, D. Toyama, C. de Masson d’Autume, Y. Li, T. Terzi, V. Mikulik, I. Babuschkin, A. Clark, D. de Las Casas, A. Guy, C. Jones, J. Bradbury, M. Johnson, B. A. Hechtman, L. Weidinger, I. Gabriel, W. S. Isaac, E. Lockhart, S. Osindero, L. Rimell, C. Dyer, O. Vinyals, K. Ayoub, J. Stanway, L. Bennett, D. Hassabis, K. Kavukcuoglu, and G. Irving (2021) · 2021
Later among the works it cites.
Samanantar: The largest publicly available parallel corpora collection for 11 indic languages
Ramesh, G., S. Doddapaneni, A. Bheemaraj, M. Jobanputra, R. AK, A. Sharma, S. Sahoo, H. Diddee, M. J, D. Kakwani, N. Kumar, A. Pradeep, S. Nagaraj, K. Deepak, V. Raghavan, A. Kunchukuttan, P. Kumar, and M. S. Khapra (2021) · 2021
Later among the works it cites.
Changing the world by changing the data
Rogers, A. (2021, August) · 2021
Later among the works it cites.
Do datasets have politics? disciplinary values in computer vision dataset development
Scheuerman, M. K., A. Hanna, and E. Denton (2021) · 2021
Later among the works it cites.
An ai-enabled approach in analyzing media data: An example from data on covid-19 news coverage in vietnam
Vuong, Q.-H., V.-P. La, T.-H. T. Nguyen, M.-H. Nguyen, T.-T. Le, and M.-T. Ho (2021) · 2021
Later among the works it cites.
Wudaocorpora: A super large-scale chinese corpora for pre-training language models
Yuan, S., H. Zhao, Z. Du, M. Ding, X. Liu, Y. Cen, X. Zou, Z. Yang, and J. Tang (2021) · 2021
Later among the works it cites.
Counterfactual memorization in neural language models
Zhang, C., D. Ippolito, K. Lee, M. Jagielski, F. Tramèr, and N. Carlini (2021) · 2021
Later among the works it cites.
HTLM: Hyper-text pre-training and prompting of language models
Aghajanyan, A., D. Okhonko, M. Lewis, M. Joshi, H. Xu, G. Ghosh, and L. Zettlemoyer (2022) · 2022
Later among the works it cites.
Does corpus quality really matter for low-resource languages?
Artetxe, M., I. Aldabe, R. Agerri, O. Perez-de Viñaspre, and A. Soroa (2022) · 2022
Later among the works it cites.
Biderman, S., K. Bicheno, and L. Gao (2022) · 2022
Later among the works it cites.
Bloom (revision 4ab0472)
BigScience Workshop (2022) · 2022
Later among the works it cites.
Gpt-neox-20b: An open-source autoregressive language model
Black, S., S. Biderman, E. Hallahan, Q. Anthony, L. Gao, L. Golding, H. He, C. Leahy, K. McDonell, J. Phang, et al. (2022) · 2022
Later among the works it cites.
Quantifying memorization across neural language models
Carlini, N., D. Ippolito, M. Jagielski, K. Lee, F. Tramer, and C. Zhang (2022) · 2022
Later among the works it cites.
BERTIN: Efficient pre-training of a Spanish language model using perplexity sampling
De la Rosa, J., E. G. Ponferrada, M. Romero, P. Villegas, P. González de Prado Salas, and M. Grandury (2022) · 2022
Later among the works it cites.
Training compute-optimal large language models
Hoffmann, J., S. Borgeaud, A. Mensch, E. Buchatskaya, T. Cai, E. Rutherford, D. d. L. Casas, L. A. Hendricks, J. Welbl, A. Clark, T. Hennigan, E. Noland, K. Millican, G. v. d. Driessche, B. Damoc, A. Guy, S. Osindero, K. Simonyan, E. Elsen, J. W. Rae, O. Vinyals, and L. Sifre (2022) · 2022
Later among the works it cites.
Deduplicating training data mitigates privacy risks in language models
Kandpal, N., E. Wallace, and C. Raffel (2022) · 2022
Later among the works it cites.
Quality at a glance: An audit of web-crawled multilingual datasets
Kreutzer, J., I. Caswell, L. Wang, A. Wahab, D. van Esch, N. Ulzii-Orshikh, A. Tapo, N. Subramani, A. Sokolov, C. Sikasote, M. Setyawan, S. Sarin, S. Samb, B. Sagot, C. Rivera, A. Rios, I. Papadimitriou, S. Osei, P. O. Suarez, I. Orife, K. Ogueji, R. Niyongabo, T. Nguyen, M. Müller, A. Müller, S. Muhammad, N. Muhammad, A. Mnyakeni, J. Mirzakhalov, T. Matangira, C. Leong, N. Lawson, S. Kudugunta, Y. Jernite, M. Jenny, O. Firat, B. Dossou, S. Dlamini, N. de Silva, S. Çabuk Ballı, S. Biderman, A. Battisti, A. Baruwa, A. Bapna, P. Baljekar, I. Azime, A. Awokoya, D. Ataman, O. Ahia, O. Ahia, S. Agrawal, and M. Adeyemi (2022) · 2022
Later among the works it cites.
Deduplicating training data makes language models better
Lee, K., D. Ippolito, A. Nystrom, C. Zhang, D. Eck, C. Callison-Burch, and N. Carlini (2022) · 2022
Later among the works it cites.
Competition-level code generation with alphacode
Li, Y., D. Choi, J. Chung, N. Kushman, J. Schrittwieser, R. Leblond, T. Eccles, J. Keeling, F. Gimeno, A. D. Lago, T. Hubert, P. Choy, C. d. M. d’Autume, I. Babuschkin, X. Chen, P.-S. Huang, J. Welbl, S. Gowal, A. Cherepanov, J. Molloy, D. J. Mankowitz, E. S. Robson, P. Kohli, N. de Freitas, K. Kavukcuoglu, and O. Vinyals (2022) · 2022
Later among the works it cites.
Documenting geographically and contextually diverse data sources: The bigscience catalogue of language data and resources
McMillan-Major, A., Z. Alyafeai, S. Biderman, K. Chen, F. De Toni, G. Dupont, H. Elsahar, C. Emezue, A. F. Aji, S. Ilić, N. Khamis, C. Leong, M. Masoud, A. Soroa, P. O. Suarez, Z. Talat, D. van Strien, and Y. Jernite (2022) · 2022
Later among the works it cites.
Using deepspeed and megatron to train megatron-turing nlg 530b, a large-scale generative language model
Smith, S., M. Patwary, B. Norick, P. LeGresley, S. Rajbhandari, J. Casper, Z. Liu, S. Prabhumoye, G. Zerveas, V. Korthikanti, E. Zhang, R. Child, R. Y. Aminabadi, J. Bernauer, X. Song, M. Shoeybi, Y. He, M. Houston, S. Tiwary, and B. Catanzaro (2022) · 2022
Later among the works it cites.
You reap what you sow: On the challenges of bias evaluation under multilingual settings
Talat, Z., A. Névéol, S. Biderman, M. Clinciu, M. Dey, S. Longpre, S. Luccioni, M. Masoud, M. Mitchell, D. Radev, S. Sharma, A. Subramonian, J. Tae, S. Tan, D. Tunuguntla, and O. van der Wal (2022) · 2022
Later among the works it cites.
Opt: Open pre-trained transformer language models
Zhang, S., S. Roller, N. Goyal, M. Artetxe, M. Chen, S. Chen, C. Dewan, M. Diab, X. Li, X. V. Lin, et al. (2022) · 2022
Later among the works it cites.