Fetching the paper…
Reading the bibliography…
We present Belebele, a multiple-choice machine reading comprehension (MRC) dataset spanning 122 language variants.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020 · 1901
Earlier work this paper cites.
Cross-lingual name tagging and linking for 282 languages
Xiaoman Pan, Boliang Zhang, Jonathan May, Joel Nothman, Kevin Knight, and Heng Ji. 2017 · 1958
Earlier work this paper cites.
Scaling laws for neural language models
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei. 2020 · 2001
Earlier work this paper cites.
Benjamin Muller, Benoît Sagot, and Djamé Seddah. 2020 · 2005
Earlier work this paper cites.
Linguistically naïve != language independent: Why NLP needs linguistic typology
Emily M. Bender. 2009 · 2009
Earlier work this paper cites.
Automatic gap-fill question generation from text books
Manish Agarwal and Prashanth Mannem. 2011 · 2011
Earlier work this paper cites.
On achieving and evaluating language-independence in nlp
Emily M. Bender. 2011 · 2011
Earlier work this paper cites.
Choice of plausible alternatives: An evaluation of commonsense causal reasoning
Melissa Roemmele, Cosmin Bejan, and Andrew Gordon. 2011 · 2011
Earlier work this paper cites.
MCTest: A challenge dataset for the open-domain machine comprehension of text
Matthew Richardson, Christopher J.C. Burges, and Erin Renshaw. 2013 · 2013
Earlier work this paper cites.
SQuAD: 100,000+ questions for machine comprehension of text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016 · 2016
Earlier work this paper cites.
Towards ai-complete question answering: A set of prerequisite toy tasks
Jason Weston, Antoine Bordes, Sumit Chopra, and Tomás Mikolov. 2016 · 2016
Earlier work this paper cites.
Transfer learning for low-resource neural machine translation
Barret Zoph, Deniz Yuret, Jonathan May, and Kevin Knight. 2016 · 2016
Earlier work this paper cites.
Enriching word vectors with subword information
Piotr Bojanowski, Edouard Grave, Armand Joulin, and Tomas Mikolov. 2017 · 2017
Earlier work this paper cites.
TriviaQA: A large scale distantly supervised challenge dataset for reading comprehension
Mandar Joshi, Eunsol Choi, Daniel Weld, and Luke Zettlemoyer. 2017 · 2017
Earlier work this paper cites.
RACE: Large-scale ReAding comprehension dataset from examinations
Guokun Lai, Qizhe Xie, Hanxiao Liu, Yiming Yang, and Eduard Hovy. 2017 · 2017
Earlier work this paper cites.
Crowdsourcing multiple choice science questions
Johannes Welbl, Nelson F. Liu, and Matt Gardner. 2017 · 2017
Earlier work this paper cites.
XNLI: Evaluating cross-lingual sentence representations
Alexis Conneau, Ruty Rinott, Guillaume Lample, Adina Williams, Samuel Bowman, Holger Schwenk, and Veselin Stoyanov. 2018 · 2018
Earlier work this paper cites.
Neural learning for question answering in italian
Danilo Croce, Alexandra Zelenanska, and Roberto Basili. 2018 · 2018
Earlier work this paper cites.
MMQA: A multi-domain multi-lingual question-answering framework for English and Hindi
Deepak Gupta, Surabhi Kumari, Asif Ekbal, and Pushpak Bhattacharyya. 2018 · 2018
Earlier work this paper cites.
Looking beyond the surface: A challenge set for reading comprehension over multiple sentences
Daniel Khashabi, Snigdha Chaturvedi, Michael Roth, Shyam Upadhyay, and Dan Roth. 2018 · 2018
Earlier work this paper cites.
SentencePiece: A simple and language independent subword tokenizer and detokenizer for neural text processing
Taku Kudo and John Richardson. 2018 · 2018
Earlier work this paper cites.
Can a suit of armor conduct electricity? a new dataset for open book question answering
Todor Mihaylov, Peter Clark, Tushar Khot, and Ashish Sabharwal. 2018 · 2018
Cited alongside, same era.
A broad-coverage challenge corpus for sentence understanding through inference
Adina Williams, Nikita Nangia, and Samuel Bowman. 2018 · 2018
Cited alongside, same era.
SWAG: A large-scale adversarial dataset for grounded commonsense inference
Rowan Zellers, Yonatan Bisk, Roy Schwartz, and Yejin Choi. 2018 · 2018
Cited alongside, same era.
Hindirc: A dataset for reading comprehension in hindi
Kaveri Anuranjana, Vijjini Anvesh Rao, and Radhika Mamidi. 2019 · 2019
Cited alongside, same era.
Beyond English-only reading comprehension: Experiments in zero-shot multilingual transfer for Bulgarian
Momchil Hardalov, Ivan Koychev, and Preslav Nakov. 2019 · 2019
Cited alongside, same era.
XL-sum: Large-scale multilingual abstractive summarization for 44 languages
Tahmid Hasan, Abhik Bhattacharjee, Md. Saiful Islam, Kazi Mubasshir, Yuan-Fang Li, Yong-Bin Kang, M. Sohel Rahman, and Rifat Shahriyar. 2021 · 2021
Later among the works it cites.
GermanQuAD and GermanDPR: Improving non-English question answering and passage retrieval
Timo Möller, Julian Risch, and Malte Pietsch. 2021 · 2021
Later among the works it cites.
When being unseen from mBERT is just the beginning: Handling new languages with multilingual language models
Benjamin Muller, Antonios Anastasopoulos, Benoît Sagot, and Djamé Seddah. 2021a · 2021
Later among the works it cites.
What ingredients make for an effective crowdsourcing protocol for difficult NLU data collection tasks?
Nikita Nangia, Saku Sugawara, Harsh Trivedi, Alex Warstadt, Clara Vania, and Samuel R. Bowman. 2021 · 2021
Later among the works it cites.
UNKs everywhere: Adapting multilingual language models to new scripts
Jonas Pfeiffer, Ivan Vulić, Iryna Gurevych, and Sebastian Ruder. 2021 · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 2019
Cited alongside, same era.
Neural Arabic question answering
Hussein Mozannar, Elie Maamary, Karl El Hajal, and Hazem Hajj. 2019 · 2019
Cited alongside, same era.
Human vs. muppet: A conservative estimate of human performance on the GLUE benchmark
Nikita Nangia and Samuel R. Bowman. 2019 · 2019
Cited alongside, same era.
MCScript2.0: A machine comprehension corpus focused on script events and participants
Simon Ostermann, Michael Roth, and Manfred Pinkal. 2019 · 2019
Cited alongside, same era.
On the cross-lingual transferability of monolingual representations
Mikel Artetxe, Sebastian Ruder, and Dani Yogatama. 2020 · 2020
Cited alongside, same era.
New protocols and negative results for textual entailment data collection
Samuel R. Bowman, Jennimaria Palomaki, Livio Baldini Soares, and Emily Pitler. 2020 · 2020
Cited alongside, same era.
What question answering can learn from trivia nerds
Jordan Boyd-Graber and Benjamin Börschinger. 2020 · 2020
Cited alongside, same era.
Construction of high-quality Tibetan dataset for machine reading comprehension
Yuan Sun, Sisi Liu, Chaofan Chen, Zhengcuo Dan, and Xiaobing Zhao. 2021 · 2021
Later among the works it cites.
The Flores-101 evaluation benchmark for low-resource and multilingual machine translation
Naman Goyal, Cynthia Gao, Vishrav Chaudhary, Peng-Jen Chen, Guillaume Wenzek, Da Ju, Sanjana Krishnan, Marc’Aurelio Ranzato, Francisco Guzmán, and Angela Fan. 2022 · 2022
Later among the works it cites.
Challenges and strategies in cross-cultural NLP
Daniel Hershcovich, Stella Frank, Heather Lent, Miryam de Lhoneux, Mostafa Abdou, Stephanie Brandl, Emanuele Bugliarello, Laura Cabello Piqueras, Ilias Chalkidis, Ruixiang Cui, Constanza Fierro, Katerina Margatina, Phillip Rust, and Anders Søgaard. 2022 · 2022
Later among the works it cites.
Quality at a glance: An audit of web-crawled multilingual datasets
Julia Kreutzer, Isaac Caswell, Lisa Wang, Ahsan Wahab, Daan van Esch, Nasanbayar Ulzii-Orshikh, Allahsera Tapo, Nishant Subramani, Artem Sokolov, Claytone Sikasote, Monang Setyawan, Supheakmungkol Sarin, Sokhar Samb, Benoît Sagot, Clara Rivera, Annette Rios, Isabel Papadimitriou, Salomey Osei, Pedro Ortiz Suarez, Iroro Orife, Kelechi Ogueji, Andre Niyongabo Rubungo, Toan Q. Nguyen, Mathias Müller, André Müller, Shamsuddeen Hassan Muhammad, Nanda Muhammad, Ayanda Mnyakeni, Jamshidbek Mirzakhalov, Tapiwanashe Matangira, Colin Leong, Nze Lawson, Sneha Kudugunta, Yacine Jernite, Mathias Jenny, Orhan Firat, Bonaventure F. P. Dossou, Sakhile Dlamini, Nisansa de Silva, Sakine Çabuk Ballı, Stella Biderman, Alessia Battisti, Ahmed Baruwa, Ankur Bapna, Pallavi Baljekar, Israel Abebe Azime, Ayodele Awokoya, Duygu Ataman, Orevaoghene Ahia, Oghenefego Ahia, Sweta Agrawal, and Mofetoluwa Adeyemi. 2022 · 2022
Later among the works it cites.
Cascading biases: Investigating the effect of heuristic annotation strategies on data and models
Chaitanya Malaviya, Sudeep Bhatia, and Mark Yatskar. 2022 · 2022
Later among the works it cites.
No language left behind: Scaling human-centered machine translation
Team NLLB, Marta R. Costa-jussà, James Cross, Onur Çelebi, Maha Elbayad, Kenneth Heafield, Kevin Heffernan, Elahe Kalbassi, Janice Lam, Daniel Licht, Jean Maillard, Anna Sun, Skyler Wang, Guillaume Wenzek, Al Youngblood, Bapi Akula, Loic Barrault, Gabriel Mejia Gonzalez, Prangthip Hansanti, John Hoffman, Semarley Jarrett, Kaushik Ram Sadagopan, Dirk Rowe, Shannon Spruit, Chau Tran, Pierre Andrews, Necip Fazil Ayan, Shruti Bhosale, Sergey Edunov, Angela Fan, Cynthia Gao, Vedanuj Goswami, Francisco Guzmán, Philipp Koehn, Alexandre Mourachko, Christophe Ropers, Safiyyah Saleem, Holger Schwenk, and Jeff Wang. 2022 · 2022
Later among the works it cites.
Bloom: A 176b-parameter open-access multilingual language model
Teven Le Scao, Angela Fan, Christopher Akiki, Ellie Pavlick, Suzana Ilić, Daniel Hesslow, Roman Castagné, Alexandra Sasha Luccioni, François Yvon, Matthias Gallé, et al. 2022 · 2022
Later among the works it cites.
Mintaka: A complex, natural, and multilingual dataset for end-to-end question answering
Priyanka Sen, Alham Fikri Aji, and Amir Saffari. 2022 · 2022
Later among the works it cites.
In what languages are generative language models the most formal? analyzing formality distribution across languages
Asım Ersoy, Gerson Vizcarra, Tahsin Mayeesha, and Benjamin Muller. 2023 · 2023
Closest in time.
MASSIVE: A 1M-example multilingual natural language understanding dataset with 51 typologically-diverse languages
Jack FitzGerald, Christopher Hench, Charith Peris, Scott Mackie, Kay Rottmann, Ana Sanchez, Aaron Nash, Liam Urbach, Vishesh Kakarala, Richa Singh, Swetha Ranganath, Laurie Crist, Misha Britan, Wouter Leeuwis, Gokhan Tur, and Prem Natarajan. 2023 · 2023
Closest in time.
Large language models as superpositions of cultural perspectives
Grgur Kovač, Masataka Sawayama, Rémy Portelas, Cédric Colas, Peter Ford Dominey, and Pierre-Yves Oudeyer. 2023 · 2023
Closest in time.
XLM-V: Overcoming the vocabulary bottleneck in multilingual masked language models
Davis Liang, Hila Gonen, Yuning Mao, Rui Hou, Naman Goyal, Marjan Ghazvininejad, Luke Zettlemoyer, and Madian Khabsa. 2023 · 2023
Closest in time.
Aksharantar: Open Indic-language transliteration datasets and models for the next billion users
Yash Madhani, Sushane Parthan, Priyanka Bedekar, Gokul Nc, Ruchi Khapra, Anoop Kunchukuttan, Pratyush Kumar, and Mitesh Khapra. 2023 · 2023
Closest in time.
Crosslingual generalization through multitask finetuning
Niklas Muennighoff, Thomas Wang, Lintang Sutawika, Adam Roberts, Stella Biderman, Teven Le Scao, M Saiful Bari, Sheng Shen, Zheng Xin Yong, Hailey Schoelkopf, Xiangru Tang, Dragomir Radev, Alham Fikri Aji, Khalid Almubarak, Samuel Albanie, Zaid Alyafeai, Albert Webson, Edward Raff, and Colin Raffel. 2023 · 2023
Closest in time.
Languages you know influence those you learn: Impact of language characteristics on multi-lingual text-to-text transfer
Benjamin Muller, Deepanshu Gupta, Jean-Philippe Fauconnier, Siddharth Patwardhan, David Vandyke, and Sachin Agarwal. 2023 · 2023
Closest in time.
Guilherme Penedo, Quentin Malartic, Daniel Hesslow, Ruxandra Cojocaru, Alessandro Cappelli, Hamza Alobeidli, Baptiste Pannier, Ebtesam Almazrouei, and Julien Launay. 2023 · 2023
Closest in time.