Fetching the paper…
Reading the bibliography…
With the success of large-scale pre-training and multilingual modeling in Natural Language Processing (NLP), recent years have seen a proliferation of large, web-mined text datasets covering hundreds of languages.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020 · 1901
Earlier work this paper cites.
Massively multilingual neural machine translation in the wild: Findings and challenges
Naveen Arivazhagan, Ankur Bapna, Orhan Firat, Dmitry Lepikhin, Melvin Johnson, Maxim Krikun, Mia Xu Chen, Yuan Cao, George F. Foster, Colin Cherry, Wolfgang Macherey, Zhifeng Chen, and Yonghui Wu. 2019 · 1907
Earlier work this paper cites.
Does automation bias decision-making?
Linda J. Skitka, Kathleen L. Mosier, and Mark Burdick. 1999 · 1999
Earlier work this paper cites.
Tags for Identifying Languages
Addison Phillips and Mark Davis. 2005 · 2005
Earlier work this paper cites.
Beyond english-centric multilingual machine translation
Angela Fan, Shruti Bhosale, Holger Schwenk, Zhiyi Ma, Ahmed El-Kishky, Siddharth Goyal, Mandeep Baines, Onur Celebi, Guillaume Wenzek, Vishrav Chaudhary, Naman Goyal, Tom Birch, Vitaliy Liptchinsky, Sergey Edunov, Edouard Grave, Michael Auli, and Armand Joulin. 2020 · 2010
Earlier work this paper cites.
Intelligent selection of language model training data
Robert C. Moore and William Lewis. 2010 · 2010
Earlier work this paper cites.
Domain adaptation via pseudo in-domain data selection
Amittai Axelrod, Xiaodong He, and Jianfeng Gao. 2011 · 2011
Earlier work this paper cites.
Pitfalls in machine learning research: Reexamining the development cycle
Stella Biderman and Walter J Scheirer. 2020 · 2011
Earlier work this paper cites.
Mt detection in web-scraped parallel corpora
Spencer Rarrick, Chris Quirk, and Will Lewis. 2011 · 2011
Earlier work this paper cites.
PanLex: Building a resource for panlingual lexical translation
David Kamholz, Jonathan Pool, and Susan Colowick. 2014 · 2014
Earlier work this paper cites.
Fasttext.zip: Compressing text classification models
Armand Joulin, Edouard Grave, Piotr Bojanowski, Matthijs Douze, Hervé Jégou, and Tomás Mikolov. 2016 · 2016
Earlier work this paper cites.
Natural language processing with small feed-forward networks
Jan A. Botha, Emily Pitler, Ji Ma, Anton Bakalov, Alex Salcianu, David Weiss, Ryan McDonald, and Slav Petrov. 2017 · 2017
Earlier work this paper cites.
Bag of tricks for efficient text classification
Armand Joulin, Edouard Grave, Piotr Bojanowski, and Tomas Mikolov. 2017 · 2017
Earlier work this paper cites.
Zipporah: a fast and scalable data cleaning system for noisy web-crawled parallel corpora
Hainan Xu and Philipp Koehn. 2017 · 2017
Earlier work this paper cites.
Data statements for natural language processing: Toward mitigating system bias and enabling better science
Emily M. Bender and Batya Friedman. 2018 · 2018
Earlier work this paper cites.
XNLI: Evaluating cross-lingual sentence representations
Alexis Conneau, Ruty Rinott, Guillaume Lample, Adina Williams, Samuel Bowman, Holger Schwenk, and Veselin Stoyanov. 2018 · 2018
Earlier work this paper cites.
Timnit Gebru, Jamie Morgenstern, Briana Vecchione, Jennifer Wortman Vaughan, Hanna Wallach, Hal Daumé III, and Kate Crawford. 2018 · 2018
Earlier work this paper cites.
Learning word vectors for 157 languages
Edouard Grave, Piotr Bojanowski, Prakhar Gupta, Armand Joulin, and Tomas Mikolov. 2018 · 2018
Earlier work this paper cites.
The dataset nutrition label: A framework to drive higher data quality standards
Sarah Holland, Ahmed Hosny, Sarah Newman, Joshua Joseph, and Kasia Chmielinski. 2018 · 2018
Earlier work this paper cites.
Dual conditional cross-entropy filtering of noisy parallel corpora
Marcin Junczys-Dowmunt. 2018 · 2018
Earlier work this paper cites.
Information nutrition labels: A plugin for online news evaluation
Vincentius Kevin, Birte Högden, Claudia Schwenger, Ali Şahan, Neelu Madan, Piush Aggarwal, Anusha Bangaru, Farid Muradov, and Ahmet Aker. 2018 · 2018
Cited alongside, same era.
On the impact of various types of noise on neural machine translation
Huda Khayrallah and Philipp Koehn. 2018 · 2018
Cited alongside, same era.
When and why are pre-trained word embeddings useful for neural machine translation?
Ye Qi, Devendra Sachan, Matthieu Felix, Sarguna Padmanabhan, and Graham Neubig. 2018 · 2018
Cited alongside, same era.
Denoising neural machine translation training with trusted data and online data selection
Wei Wang, Taro Watanabe, Macduff Hughes, Tetsuji Nakagawa, and Ciprian Chelba. 2018 · 2018
Cited alongside, same era.
JW300: A wide-coverage parallel corpus for low-resource languages
Željko Agić and Ivan Vulić. 2019 · 2019
Cited alongside, same era.
XTREME: A massively multilingual multi-task benchmark for evaluating cross-lingual generalisation
Junjie Hu, Sebastian Ruder, Aditya Siddhant, Graham Neubig, Orhan Firat, and Melvin Johnson. 2020 · 2020
Later among the works it cites.
IndicNLPSuite: Monolingual corpora, evaluation benchmarks and pre-trained multilingual language models for Indian languages
Divyanshu Kakwani, Anoop Kunchukuttan, Satish Golla, Gokul N.C., Avik Bhattacharyya, Mitesh M. Khapra, and Pratyush Kumar. 2020 · 2020
Later among the works it cites.
Findings of the WMT 2020 shared task on parallel corpus filtering and alignment
Philipp Koehn, Vishrav Chaudhary, Ahmed El-Kishky, Naman Goyal, Peng-Jen Chen, and Francisco Guzmán. 2020 · 2020
Later among the works it cites.
Greek-bert: The greeks visiting sesame street
John Koutsikakis, Ilias Chalkidis, Prodromos Malakasiotis, and Ion Androutsopoulos. 2020 · 2020
Later among the works it cites.
CamemBERT: a tasty French language model
Louis Martin, Benjamin Muller, Pedro Javier Ortiz Suárez, Yoann Dupont, Laurent Romary, Éric de la Clergerie, Djamé Seddah, and Benoît Sagot. 2020 · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Massively multilingual sentence embeddings for zero-shot cross-lingual transfer and beyond
Mikel Artetxe and Holger Schwenk. 2019 · 2019
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Cited alongside, same era.
ParaCrawl: Web-scale parallel corpora for the languages of the EU
Miquel Esplà, Mikel Forcada, Gema Ramírez-Sánchez, and Hieu Hoang. 2019 · 2019
Cited alongside, same era.
Microsoft translator at WMT 2019: Towards large-scale document-level neural machine translation
Marcin Junczys-Dowmunt. 2019 · 2019
Cited alongside, same era.
Asynchronous pipelines for processing huge corpora on medium to low resource infrastructures
Pedro Javier Ortiz Suárez, Benoît Sagot, and Laurent Romary. 2019 · 2019
Cited alongside, same era.
Mithralabel: Flexible dataset nutritional labels for responsible data science
Chenkai Sun, Abolfazl Asudeh, H. V. Jagadish, Bill Howe, and Julia Stoyanovich. 2019 · 2019
Cited alongside, same era.
ParaCrawl: Web-scale acquisition of parallel corpora
Marta Bañón, Pinzhen Chen, Barry Haddow, Kenneth Heafield, Hieu Hoang, Miquel Esplà-Gomis, Mikel L. Forcada, Amir Kamran, Faheem Kirefu, Philipp Koehn, Sergio Ortiz Rojas, Leopoldo Pla Sempere, Gema Ramírez-Sánchez, Elsa Sarrías, Marek Strelec, Brian Thompson, William Waites, Dion Wiggins, and Jaume Zaragoza. 2020 · 2020
Cited alongside, same era.
RoBERT – a Romanian BERT model
Mihai Masala, Stefan Ruseti, and Mihai Dascalu. 2020 · 2020
Later among the works it cites.
A monolingual approach to contextualized word embeddings for mid-resource languages
Pedro Javier Ortiz Suárez, Laurent Romary, and Benoît Sagot. 2020 · 2020
Later among the works it cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2020 · 2020
Later among the works it cites.
IndoNLU: Benchmark and resources for evaluating Indonesian natural language understanding
Bryan Wilie, Karissa Vincentio, Genta Indra Winata, Samuel Cahyawijaya, Xiaohong Li, Zhi Yuan Lim, Sidik Soleman, Rahmad Mahendra, Pascale Fung, Syafri Bahar, and Ayu Purwarianti. 2020 · 2020
Later among the works it cites.
AraELECTRA: Pre-training text discriminators for Arabic language understanding
Wissam Antoun, Fady Baly, and Hazem Hajj. 2021 · 2021
Closest in time.
On the dangers of stochastic parrots: Can language models be too big?
Emily M. Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell. 2021 · 2021
Closest in time.
Large image datasets: A pyrrhic win for computer vision?
Abeba Birhane and Vinay Uday Prabhu. 2021 · 2021
Closest in time.
HeBERT & HebEMO: a Hebrew BERT Model and a Tool for Polarity Analysis and Emotion Recognition
Avihay Chriqui and Inbal Yahav. 2021 · 2021
Closest in time.
Documenting the english colossal clean crawled corpus
Jesse Dodge, Maarten Sap, Ana Marasovic, William Agnew, Gabriel Ilharco, Dirk Groeneveld, and Matt Gardner. 2021 · 2021
Closest in time.
The FLORES-101 evaluation benchmark for low-resource and multilingual machine translation
Naman Goyal, Cynthia Gao, Vishrav Chaudhary, Peng-Jen Chen, Guillaume Wenzek, Da Ju, Sanjana Krishnan, Marc’Aurelio Ranzato, Francisco Guzmán, and Angela Fan. 2021 · 2021
Closest in time.
What’s in the box? an analysis of undesirable content in the common crawl corpus
Alexandra Sasha Luccioni and Joseph D. Viviano. 2021 · 2021
Closest in time.
WikiMatrix: Mining 135M parallel sentences in 1620 language pairs from Wikipedia
Holger Schwenk, Vishrav Chaudhary, Shuo Sun, Hongyu Gong, and Francisco Guzmán. 2021 · 2021
Closest in time.
AlephBERT:A Hebrew Large Pre-Trained Language Model to Start-off your Hebrew NLP Application With
Amit Seker, Elron Bandel, Dan Bareket, Idan Brusilovsky, Refael Shaked Greenfeld, and Reut Tsarfaty. 2021 · 2021
Closest in time.
mT5: A massively multilingual pre-trained text-to-text transformer
Linting Xue, Noah Constant, Adam Roberts, Mihir Kale, Rami Al-Rfou, Aditya Siddhant, Aditya Barua, and Colin Raffel. 2021 · 2021
Closest in time.