Fetching the paper…
Reading the bibliography…
We introduce MADLAD-400, a manually audited, general domain 3T token monolingual dataset based on CommonCrawl, spanning 419 languages.
Europarl: A parallel corpus for statistical machine translation
P. Koehn · 2005
Earlier work this paper cites.
Introducing the autshumato integrated translation environment
H. J. Groenewald and W. Fourie · 2009
Earlier work this paper cites.
I. Caswell, T. Breiner, D. van Esch, and A. Bapna · 2010
Earlier work this paper cites.
Complete multilingual neural machine translation
M. Freitag and O. Firat · 2010
Earlier work this paper cites.
The Kyoto free translation task
G. Neubig · 2011
Earlier work this paper cites.
Parallel data, tools and interfaces in opus
J. Tiedemann · 2012
Earlier work this paper cites.
One billion word benchmark for measuring progress in statistical language modeling
C. Chelba, T. Mikolov, M. Schuster, Q. Ge, T. Brants, P. Koehn, and T. Robinson · 2013
Earlier work this paper cites.
Aligning books and movies: Towards story-like visual explanations by watching movies and reading books
Y. Zhu, R. Kiros, R. Zemel, R. Salakhutdinov, R. Urtasun, A. Torralba, and S. Fidler · 2015
Earlier work this paper cites.
Aspec: Asian scientific paper excerpt corpus
T. Nakazawa, M. Yaguchi, K. Uchimoto, M. Utiyama, E. Sumita, S. Kurohashi, and H. Isahara · 2016
Earlier work this paper cites.
Improving neural machine translation models with monolingual data
R. Sennrich, B. Haddow, and A. Birch · 2016
Earlier work this paper cites.
The united nations parallel corpus v1. 0
M. Ziemski, M. Junczys-Dowmunt, and B. Pouliquen · 2016
Earlier work this paper cites.
Google’s multilingual neural machine translation system: Enabling zero-shot translation
M. Johnson, M. Schuster, Q. V. Le, M. Krikun, Y. Wu, Z. Chen, N. Thorat, F. Viégas, M. Wattenberg, G. Corrado, et al · 2017
Earlier work this paper cites.
Jesc: Japanese-english subtitle corpus
R. Pryzant, Y. Chung, D. Jurafsky, and D. Britz · 2017
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. u. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova · 2018
Earlier work this paper cites.
T. Kudo and J. Richardson · 2018
Earlier work this paper cites.
A call for clarity in reporting BLEU scores
M. Post · 2018
Earlier work this paper cites.
When and why are pre-trained word embeddings useful for neural machine translation
Q. Ye, S. Devendra, F. Matthieu, P. Sarguna, and N. Graham · 2018
Earlier work this paper cites.
Jw300: A wide-coverage parallel corpus for low-resource languages
Ž. Agic and I. Vulic · 2019
Earlier work this paper cites.
Unsupervised cross-lingual representation learning at scale
A. Conneau, K. Khandelwal, N. Goyal, V. Chaudhary, G. Wenzek, F. Guzmán, E. Grave, M. Ott, L. Zettlemoyer, and V. Stoyanov · 2019
Earlier work this paper cites.
Paracrawl: Web-scale parallel corpora for the languages of the eu
M. Esplà-Gomis, M. L. Forcada, G. Ramírez-Sánchez, and H. Hoang · 2019
Earlier work this paper cites.
A baseline neural machine translation system for indian languages
J. Philip, V. P. Namboodiri, and C. Jawahar · 2019
Earlier work this paper cites.
Wikimatrix: Mining 135m parallel sentences in 1620 language pairs from wikipedia
H. Schwenk, V. Chaudhary, S. Sun, H. Gong, and F. Guzmán · 2019
Earlier work this paper cites.
Mass: Masked sequence to sequence pre-training for language generation
K. Song, X. Tan, T. Qin, J. Lu, and T.-Y. Liu · 2019
Earlier work this paper cites.
Language models are few-shot learners
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al · 2020
Earlier work this paper cites.
A. Fernando, S. Ranathunga, and G. Dias · 2020
Cited alongside, same era.
Pmindia–a collection of parallel corpora of languages of india
B. Haddow and F. Kirefu · 2020
Cited alongside, same era.
Xtreme: A massively multilingual multi-task benchmark for evaluating cross-lingual generalisation
J. Hu, S. Ruder, A. Siddhant, G. Neubig, O. Firat, and M. Johnson · 2020
Cited alongside, same era.
The nunavut hansard inuktitut–english parallel corpus 3.0 with preliminary machine translation results
E. Joanis, R. Knowles, R. Kuhn, S. Larkin, P. Littell, C.-k. Lo, D. Stewart, and J. Micher · 2020
Cited alongside, same era.
The state and fate of linguistic diversity and inclusion in the nlp world
Quantifying memorization across neural language models
N. Carlini, D. Ippolito, M. Jagielski, K. Lee, F. Tramer, and C. Zhang · 2022
Later among the works it cites.
Palm: Scaling language modeling with pathways
A. Chowdhery, S. Narang, J. Devlin, M. Bosma, G. Mishra, A. Roberts, P. Barham, H. W. Chung, C. Sutton, S. Gehrmann, et al · 2022
Later among the works it cites.
NTREX-128 – news test references for MT evaluation of 128 languages
C. Federmann, T. Kocmi, and Y. Xin · 2022
Later among the works it cites.
Preventing verbatim memorization in language models gives a false sense of privacy
D. Ippolito, F. Tramèr, M. Nasr, C. Zhang, M. Jagielski, K. Lee, C. A. Choquette-Choo, and N. Carlini · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
P. Joshi, S. Santy, A. Budhiraja, K. Bali, and M. Choudhury · 2020
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu · 2020
Cited alongside, same era.
mt5: A massively multilingual pre-trained text-to-text transformer
L. Xue, N. Constant, A. Roberts, M. Kale, R. Al-Rfou, A. Siddhant, A. Barua, and C. Raffel · 2020
Cited alongside, same era.
Improving massively multilingual neural machine translation and zero-shot translation
B. Zhang, P. Williams, I. Titov, and R. Sennrich · 2020
Cited alongside, same era.
J. Bandy and N. Vincent · 2021
Cited alongside, same era.
On the opportunities and risks of foundation models
R. Bommasani, D. A. Hudson, E. Adeli, R. Altman, S. Arora, S. von Arx, M. S. Bernstein, J. Bohg, A. Bosselut, E. Brunskill, et al · 2021
Cited alongside, same era.
Extracting training data from large language models
N. Carlini, F. Tramer, E. Wallace, M. Jagielski, A. Herbert-Voss, K. Lee, A. Roberts, T. Brown, D. Song, U. Erlingsson, et al · 2021
Cited alongside, same era.
Documenting large webtext corpora: A case study on the colossal clean crawled corpus
J. Dodge, M. Sap, A. Marasović, W. Agnew, G. Ilharco, D. Groeneveld, M. Mitchell, and M. Gardner · 2021
Cited alongside, same era.
M. Jagielski, O. Thakkar, F. Tramer, D. Ippolito, K. Lee, N. Carlini, E. Wallace, S. Song, A. Thakurta, N. Papernot, et al · 2022
Later among the works it cites.
Quality at a glance: An audit of web-crawled multilingual datasets
J. Kreutzer, I. Caswell, L. Wang, A. Wahab, D. van Esch, N. Ulzii-Orshikh, A. Tapo, N. Subramani, A. Sokolov, C. Sikasote, M. Setyawan, S. Sarin, S. Samb, B. Sagot, C. Rivera, A. Rios, I. Papadimitriou, S. Osei, P. O. Suarez, I. Orife, K. Ogueji, A. N. Rubungo, T. Q. Nguyen, M. Müller, A. Müller, S. H. Muhammad, N. Muhammad, A. Mnyakeni, J. Mirzakhalov, T. Matangira, C. Leong, N. Lawson, S. Kudugunta, Y. Jernite, M. Jenny, O. Firat, B. F. P. Dossou, S. Dlamini, N. de Silva, S. Çabuk Ballı, S. Biderman, A. Battisti, A. Baruwa, A. Bapna, P. Baljekar, I. A. Azime, A. Awokoya, D. Ataman, O. Ahia, O. Ahia, S. Agrawal, and M. Adeyemi · 2022
Later among the works it cites.
The bigscience roots corpus: A 1.6 tb composite multilingual dataset
H. Laurençon, L. Saulnier, T. Wang, C. Akiki, A. Villanova del Moral, T. Le Scao, L. Von Werra, C. Mou, E. González Ponferrada, H. Nguyen, et al · 2022
Later among the works it cites.
Bloom library: Multimodal datasets in 300+ languages for a variety of downstream tasks
C. Leong, J. Nemecek, J. Mansdorfer, A. Filighera, A. Owodunni, and D. Whitenack · 2022
Later among the works it cites.
Opportunities for human-centered evaluation of machine translation systems
D. Liebling, K. Heller, S. Robertson, and W. Deng · 2022
Later among the works it cites.
No language left behind: Scaling human-centered machine translation
NLLBTeam, M. R. Costa-jussà, J. Cross, O. Çelebi, M. Elbayad, K. Heafield, K. Heffernan, E. Kalbassi, J. Lam, D. Licht, J. Maillard, A. Sun, S. Wang, G. Wenzek, A. Youngblood, B. Akula, L. Barrault, G. M. Gonzalez, P. Hansanti, J. Hoffman, S. Jarrett, K. R. Sadagopan, D. Rowe, S. Spruit, C. Tran, P. Andrews, N. F. Ayan, S. Bhosale, S. Edunov, A. Fan, C. Gao, V. Goswami, F. Guzmán, P. Koehn, A. Mourachko, C. Ropers, S. Saleem, H. Schwenk, and J. Wang · 2022
Later among the works it cites.
A. Siddhant, A. Bapna, O. Firat, Y. Cao, M. X. Chen, I. Caswell, and X. Garcia · 2022
Later among the works it cites.
Unifying language learning paradigms
Y. Tay, M. Dehghani, V. Q. Tran, X. Garcia, D. Bahri, T. Schuster, H. S. Zheng, N. Houlsby, and D. Metzler · 2022
Later among the works it cites.
Prompting palm for translation: Assessing strategies and performance
D. Vilar, M. Freitag, C. Cherry, J. Luo, V. Ratnakar, and G. Foster · 2022
Later among the works it cites.
Mega: Multilingual evaluation of generative ai
K. Ahuja, R. Hada, M. Ochieng, P. Jain, H. Diddee, S. Maina, T. Ganu, S. Segal, M. Axmed, K. Bali, et al · 2023
Closest in time.
R. Anil, A. M. Dai, O. Firat, M. Johnson, D. Lepikhin, A. Passos, S. Shakeri, E. Taropa, P. Bailey, Z. Chen, et al · 2023
Closest in time.
Buffet: Benchmarking large language models for few-shot cross-lingual transfer
A. Asai, S. Kudugunta, X. V. Yu, T. Blevins, H. Gonen, M. Reid, Y. Tsvetkov, S. Ruder, and H. Hajishirzi · 2023
Closest in time.
When does monolingual data help multilingual translation: The role of domain and model scale
C. Baziotis, B. Zhang, A. Birch, and B. Haddow · 2023
Closest in time.
Unimax: Fairer and more effective language sampling for large-scale multilingual pretraining
H. W. Chung, N. Constant, X. Garcia, A. Roberts, Y. Tay, S. Narang, and O. Firat · 2023
Closest in time.
J. Gala, P. A. Chitale, R. AK, S. Doddapaneni, V. Gumma, A. Kumar, J. Nawale, A. Sujatha, R. Puduppully, V. Raghavan, et al · 2023
Closest in time.
The unreasonable effectiveness of few-shot learning for machine translation
X. Garcia, Y. Bansal, C. Cherry, G. Foster, M. Krikun, F. Feng, M. Johnson, and O. Firat · 2023
Closest in time.
Glot500: Scaling multilingual corpora and language models to 500 languages
A. ImaniGooghari, P. Lin, A. H. Kargaran, S. Severini, M. J. Sabet, N. Kassner, C. Ma, H. Schmid, A. F. Martins, F. Yvon, et al · 2023
Closest in time.
Bilex rx: Lexical data augmentation for massively multilingual machine translation, 2023
A. Jones, I. Caswell, I. Saxena, and O. Firat · 2023
Closest in time.
Measuring the impact of programming language distribution
G. Orlanski, K. Xiao, X. Garcia, J. Hui, J. Howland, J. Malmaud, J. Austin, R. Singh, and M. Catasta · 2023
Closest in time.
Prompting large language model for machine translation: A case study
B. Zhang, B. Haddow, and A. Birch · 2023
Closest in time.