Fetching the paper…
Reading the bibliography…
Datasets are foundational to many breakthroughs in modern artificial intelligence.
Explain yourself! leveraging language models for commonsense reasoning
Nazneen Fatema Rajani, Bryan McCann, Caiming Xiong, and Richard Socher · 1906
Earlier work this paper cites.
Quartz: An open-domain dataset of qualitative relationship questions
Oyvind Tafjord, Matt Gardner, Kevin Lin, and Peter Clark · 1909
Earlier work this paper cites.
Planning for a standard language in modern norway
Einar Haugen · 1959
Earlier work this paper cites.
Binary codes capable of correcting deletions, insertions, and reversals
Vladimir I Levenshtein et al · 1966
Earlier work this paper cites.
THE PRONOUNS OF POWER AND SOLIDARITY , pp. 252–275
Roger Brown and Albert Gilman · 1968
Earlier work this paper cites.
What makes a social class? on the theoretical and practical existence of groups
Pierre Bourdieu · 1987
Earlier work this paper cites.
Issues in dialect obsolescence: An introduction
Walt Wolfram · 1997
Earlier work this paper cites.
The Adinkra dictionary: A visual primer on the language of Adinkra
W Bruce Willis · 1998
Earlier work this paper cites.
Toward semantics-based answer pinpointing
Eduard Hovy, Laurie Gerber, Ulf Hermjakob, Chin-Yew Lin, and Deepak Ravichandran · 2001
Earlier work this paper cites.
Learning question classifiers
Xin Li and Dan Roth · 2002
Earlier work this paper cites.
English gigaword
David Graff, Junbo Kong, Ke Chen, and Kazuaki Maeda · 2003
Earlier work this paper cites.
Automatically constructing a corpus of sentential paraphrases
Bill Dolan and Chris Brockett · 2005
Earlier work this paper cites.
AG’s Corpus of News Articles
Antonio Gulli · 2005
Earlier work this paper cites.
Seeing stars: Exploiting class relationships for sentiment categorization with respect to rating scales
Bo Pang and Lillian Lee · 2005
Earlier work this paper cites.
Language and Social Relations
Asif Agha · 2006
Earlier work this paper cites.
Frontiers in linguistic annotation for lower-density languages
Mike Maxwell and Baden Hughes · 2006
Earlier work this paper cites.
An empirical analysis of open source software developers’ motivations and continuance intentions
Chorng-Guang Wu, James H Gerlach, and Clifford E Young · 2007
Earlier work this paper cites.
Data quality from crowdsourcing: a study of annotation selection criteria
Pei-Yun Hsueh, Prem Melville, and Vikas Sindhwani · 2009
Earlier work this paper cites.
A Dataset for Assessing Machine Translation Evaluation Metrics
Lucia Specia, Nicola Cancedda, and Marc Dymetman · 2010
Earlier work this paper cites.
Estimating machine translation post-editing effort with HTER
Lucia Specia and Atefeh Farzindar · 2010
Earlier work this paper cites.
Learning word vectors for sentiment analysis
Andrew L. Maas, Raymond E. Daly, Peter T. Pham, Dan Huang, Andrew Y. Ng, and Christopher Potts · 2011
Earlier work this paper cites.
Dialects, cultural identity, and economic exchange
Oliver Falck, Stephan Heblich, Alfred Lameli, and Jens Südekum · 2012
Earlier work this paper cites.
Quora question pairs, 2012
Shankar Iyer, Nikhil Dandekar, and Kornäl Csernai · 2012
Earlier work this paper cites.
African American, Creole, and other vernacular Englishes in education: A bibliographic resource
John R Rickford, Julie Sweetland, Angela E Rickford, and Thomas Grano · 2012
Earlier work this paper cites.
Crowdsourcing research opportunities: Lessons from natural language processing
Marta Sabou, Kalina Bontcheva, and Arno Scharl · 2012
Earlier work this paper cites.
Language diversity and social action: A third locus of linguistic relativity
Jack Sidnell and N. J. Enfield · 2012
Earlier work this paper cites.
Parallel data, tools and interfaces in opus
Jörg Tiedemann · 2012
Earlier work this paper cites.
Semantic parsing on Freebase from question-answer pairs
Jonathan Berant, Andrew Chou, Roy Frostig, and Percy Liang · 2013
Earlier work this paper cites.
Crowd science: The organization of scientific research in open collaborative projects
Chiara Franzoni and Henry Sauermann · 2013
Earlier work this paper cites.
Real-world semi-supervised learning of pos-taggers for low-resource languages
Dan Garrette, Jason Mielens, and Jason Baldridge · 2013
Earlier work this paper cites.
Francophonie
Cécile B. Vigouroux · 2013
Earlier work this paper cites.
Dbpedia - a large-scale, multilingual knowledge base extracted from wikipedia
Jens Lehmann, Robert Isele, Max Jakob, Anja Jentzsch, Dimitris Kontokostas, Pablo Mendes, Sebastian Hellmann, Mohamed Morsey, Patrick Van Kleef, Sören Auer, and Christian Bizer · 2014
Earlier work this paper cites.
Big data, big questions| metaphors of big data
Cornelius Puschmann and Jean Burgess · 2014
Earlier work this paper cites.
Big data metaphors we live by
Kailash Awati and Simon Buckingham Shum · 2015
Earlier work this paper cites.
Soylent: a word processor with a crowd inside
Michael S. Bernstein, Greg Little, Robert C. Miller, Björn Hartmann, Mark S. Ackerman, David R. Karger, David Crowell, and Katrina Panovich · 2015
Earlier work this paper cites.
Designing a gamification mechanism to encourage contributions in a crowdsourcing system
Flavio A. de Franga, Adriana S. Vivacqua, and Maria Luiza M. Campos · 2015
Earlier work this paper cites.
What is big data? a consensual definition and a review of key research topics
Andrea De Mauro, Marco Greco, and Michele Grimaldi · 2015
Earlier work this paper cites.
Teaching machines to read and comprehend
Karl Moritz Hermann, Tomas Kocisky, Edward Grefenstette, Lasse Espeholt, Will Kay, Mustafa Suleyman, and Phil Blunsom · 2015
Earlier work this paper cites.
Challenges of studying and processing dialects in social media
Anna Jørgensen, Dirk Hovy, and Anders Søgaard · 2015
Earlier work this paper cites.
chrF: character n-gram F-score for automatic MT evaluation
Maja Popović · 2015
Earlier work this paper cites.
A neural attention model for abstractive sentence summarization
Alexander M. Rush, Sumit Chopra, and Jason Weston · 2015
Earlier work this paper cites.
WikiQA: A challenge dataset for open-domain question answering
Yi Yang, Wen-tau Yih, and Christopher Meek · 2015
Earlier work this paper cites.
Character-level convolutional networks for text classification
Xiang Zhang, Junbo Zhao, and Yann LeCun · 2015
Earlier work this paper cites.
Breaking the unwritten language barrier: The bulb project
Gilles Adda, Sebastian Stüker, Martine Adda-Decker, Odette Ambouroue, Laurent Besacier, David Blachon, Hélène Bonneau-Maynard, Pierre Godard, Fatima Hamlaoui, Dmitry Idiatov, Guy-Noël Kouarata, Lori Lamel, Emmanuel-Moselly Makasso, Annie Rialland, Mark Van de Velde, François Yvon, and Sabine Zerbian · 2016
Earlier work this paper cites.
Demographic dialectal variation in social media: A case study of African-American English
Su Lin Blodgett, Lisa Green, and Brendan O’Connor · 2016
Earlier work this paper cites.
Generating text from structured data with application to the biography domain
Rémi Lebret, David Grangier, and Michael Auli · 2016
Earlier work this paper cites.
Abstractive text summarization using sequence-to-sequence rnns and beyond, 2016
Ramesh Nallapati, Bowen Zhou, Cicero Nogueira dos santos, Caglar Gulcehre, and Bing Xiang · 2016
Earlier work this paper cites.
The lambada dataset: Word prediction requiring a broad discourse context, 2016
Denis Paperno, Germán Kruszewski, Angeliki Lazaridou, Quan Ngoc Pham, Raffaella Bernardi, Sandro Pezzelle, Marco Baroni, Gemma Boleda, and Raquel Fernández · 2016
Earlier work this paper cites.
SQuAD: 100,000+ Questions for Machine Comprehension of Text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang · 2016
Earlier work this paper cites.
Neural machine translation of rare words with subword units
Rico Sennrich, Barry Haddow, and Alexandra Birch · 2016
Earlier work this paper cites.
LORELEI language packs: Data, tools, and resources for technology development in low resource languages
Stephanie Strassel and Jennifer Tracey · 2016
Earlier work this paper cites.
Transfer learning for low-resource neural machine translation
Barret Zoph, Deniz Yuret, Jonathan May, and Kevin Knight · 2016
Earlier work this paper cites.
Co-Operative Action
Charles Goodwin · 2017
Earlier work this paper cites.
Android apps and user feedback: a dataset for software evolution and quality improvement
Giovanni Grano, Andrea Di Sorbo, Francesco Mercaldo, Corrado A Visaggio, Gerardo Canfora, and Sebastiano Panichella · 2017
Earlier work this paper cites.
triviaqa: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehension
Mandar Joshi, Eunsol Choi, Daniel Weld, and Luke Zettlemoyer · 2017
Earlier work this paper cites.
RACE: Large-scale ReAding comprehension dataset from examinations
Guokun Lai, Qizhe Xie, Hanxiao Liu, Yiming Yang, and Eduard Hovy · 2017
Earlier work this paper cites.
Gamified crowdsourcing: Conceptualization, literature review, and future agenda
Benedikt Morschheuser, Juho Hamari, Jonna Koivisto, and Alexander Maedche · 2017
Earlier work this paper cites.
Code-switching
Carol Myers-Scotton · 2017
Earlier work this paper cites.
Get to the point: Summarization with pointer-generator networks
Abigail See, Peter J. Liu, and Christopher D. Manning · 2017
Earlier work this paper cites.
Crowdsourcing multiple choice science questions
Johannes Welbl, Nelson F Liu, and Matt Gardner · 2017
Earlier work this paper cites.
Findings of the VarDial evaluation campaign 2017
Marcos Zampieri, Shervin Malmasi, Nikola Ljubešić, Preslav Nakov, Ahmed Ali, Jörg Tiedemann, Yves Scherrer, and Noëmi Aepli · 2017
Earlier work this paper cites.
Research on mobile phone data in the global south
Seyram Avle, Emmanuel Quartey, and David Hutchful · 2018
Earlier work this paper cites.
Learning to split and rephrase from Wikipedia edit history
Jan A. Botha, Manaal Faruqui, John Alex, Jason Baldridge, and Dipanjan Das · 2018
Earlier work this paper cites.
Deep learning for classical japanese literature
Tarin Clanuwat, Mikel Bober-Irizar, Asanobu Kitamoto, Alex Lamb, Kazuaki Yamamoto, and David Ha · 2018
Earlier work this paper cites.
Think you have solved question answering? try arc, the ai2 reasoning challenge
Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord · 2018
Earlier work this paper cites.
Xnli: Evaluating cross-lingual sentence representations
Alexis Conneau, Guillaume Lample, Ruty Rinott, Adina Williams, Samuel R Bowman, Holger Schwenk, and Veselin Stoyanov · 2018
Earlier work this paper cites.
Pronoun translation in english-french machine translation: An analysis of error types, 2018
Christian Hardmeier and Liane Guillou · 2018
Earlier work this paper cites.
Looking beyond the surface:a challenge set for reading comprehension over multiple sentences
Daniel Khashabi, Snigdha Chaturvedi, Michael Roth, Shyam Upadhyay, and Dan Roth · 2018
Earlier work this paper cites.
The natural language decathlon: Multitask learning as question answering
Bryan McCann, Nitish Shirish Keskar, Caiming Xiong, and Richard Socher · 2018
Earlier work this paper cites.
Can a suit of armor conduct electricity? a new dataset for open book question answering
Todor Mihaylov, Peter Clark, Tushar Khot, and Ashish Sabharwal · 2018
Earlier work this paper cites.
Shashi Narayan, Shay B. Cohen, and Mirella Lapata · 2018
Earlier work this paper cites.
A call for clarity in reporting BLEU scores
Matt Post · 2018
Earlier work this paper cites.
DuoRC: Towards Complex Language Understanding with Paraphrased Reading Comprehension
Amrita Saha, Rahul Aralikatte, Mitesh M. Khapra, and Karthik Sankaranarayanan · 2018
Earlier work this paper cites.
Language policy in french colonies and after independence
Bernard Spolsky · 2018
Earlier work this paper cites.
Glue: A multi-task benchmark and analysis platform for natural language understanding
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R Bowman · 2018
Earlier work this paper cites.
Neural network acceptability judgments
Alex Warstadt, Amanpreet Singh, and Samuel R Bowman · 2018
Earlier work this paper cites.
HotpotQA: A dataset for diverse, explainable multi-hop question answering
Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William W. Cohen, Ruslan Salakhutdinov, and Christopher D. Manning · 2018
Earlier work this paper cites.
Massively multilingual neural machine translation, 2019
Roee Aharoni, Melvin Johnson, and Orhan Firat · 2019
Earlier work this paper cites.
On the cross-lingual transferability of monolingual representations
Mikel Artetxe, Sebastian Ruder, and Dani Yogatama · 2019
Earlier work this paper cites.
Boolq: Exploring the surprising difficulty of natural yes/no questions
Christopher Clark, Kenton Lee, Ming-Wei Chang, Tom Kwiatkowski, Michael Collins, and Kristina Toutanova · 2019
Earlier work this paper cites.
A span-extraction dataset for Chinese machine reading comprehension
Yiming Cui, Ting Liu, Wanxiang Che, Li Xiao, Zhipeng Chen, Wentao Ma, Shijin Wang, and Guoping Hu · 2019
Earlier work this paper cites.
Quoref: A reading comprehension dataset with questions requiring coreferential reasoning
Pradeep Dasigi, Nelson F. Liu, Ana Marasovic, Noah A. Smith, and Matt Gardner · 2019
Earlier work this paper cites.
Multi-news: a large-scale multi-document summarization dataset and abstractive hierarchical model, 2019
Alexander R. Fabbri, Irene Li, Tianwei She, Suyi Li, and Dragomir R. Radev · 2019
Earlier work this paper cites.
SAMSum corpus: A human-annotated dialogue dataset for abstractive summarization
Bogdan Gliwa, Iwona Mochol, Maciej Biesek, and Aleksander Wawer · 2019
Earlier work this paper cites.
Cosmos QA: Machine reading comprehension with contextual commonsense reasoning
Lifu Huang, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi · 2019
Earlier work this paper cites.
Unsung challenges of building and deploying language technologies for low resource language communities
Pratik Joshi, Christain Barnes, Sebastin Santy, Simran Khanuja, Sanket Shah, Anirudh Srinivasan, Satwik Bhattamishra, Sunayana Sitaram, Monojit Choudhury, and Kalika Bali · 2019
Earlier work this paper cites.
Natural questions: A benchmark for question answering research
Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, Kristina Toutanova, Llion Jones, Matthew Kelcey, Ming-Wei Chang, Andrew M. Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov · 2019
Earlier work this paper cites.
Cross-lingual language model pretraining
Guillaume Lample and Alexis Conneau · 2019
Earlier work this paper cites.
Reasoning over paragraph effects in situations
Kevin Lin, Oyvind Tafjord, Peter Clark, and Matt Gardner · 2019
Earlier work this paper cites.
All Data Are Local: Thinking Critically in a Data-Driven Society
Yanni Alexander Loukissas · 2019
Earlier work this paper cites.
Asynchronous pipelines for processing huge corpora on medium to low resource infrastructures
Pedro Javier Ortiz Su’arez, Benoit Sagot, and Laurent Romary · 2019
Earlier work this paper cites.
Wic: the word-in-context dataset for evaluating context-sensitive meaning representations, 2019
Mohammad Taher Pilehvar and Jose Camacho-Collados · 2019
Earlier work this paper cites.
CoQA: A conversational question answering challenge
Siva Reddy, Danqi Chen, and Christopher D. Manning · 2019
Earlier work this paper cites.
Sentence-bert: Sentence embeddings using siamese bert-networks
Nils Reimers and Iryna Gurevych · 2019
Earlier work this paper cites.
Socialiqa: Commonsense reasoning about social interactions, 2019
Maarten Sap, Hannah Rashkin, Derek Chen, Ronan LeBras, and Yejin Choi · 2019
Earlier work this paper cites.
Drcd: a chinese machine reading comprehension dataset, 2019
Chih Chieh Shao, Trois Liu, Yuting Lai, Yiying Tseng, and Sam Tsai · 2019
Earlier work this paper cites.
Two centuries of spreading language loss
Gary F. Simons · 2019
Earlier work this paper cites.
Data is the new what? popular metaphors & professional ethics in emerging data culture
Luke Stark and Anna Lauren Hoffmann · 2019
Earlier work this paper cites.
Wiqa: A dataset for "what if…" reasoning over procedural text
Niket Tandon, Bhavana Dalvi Mishra, Keisuke Sakaguchi, Antoine Bosselut, and Peter Clark · 2019
Earlier work this paper cites.
Multi-hop reading comprehension across multiple documents by reasoning over heterogeneous graphs, 2019
Ming Tu, Guangtao Wang, Jing Huang, Yun Tang, Xiaodong He, and Bowen Zhou · 2019
Earlier work this paper cites.
Superglue: A stickier benchmark for general-purpose language understanding systems
Alex Wang, Yada Pruksachatkun, Nikita Nangia, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R Bowman · 2019
Earlier work this paper cites.
Paraphrasing with large language models
Sam Witteveen and Martin Andrews · 2019
Earlier work this paper cites.
PAWS-X: A Cross-lingual Adversarial Dataset for Paraphrase Identification
Yinfei Yang, Yuan Zhang, Chris Tar, and Jason Baldridge · 2019
Earlier work this paper cites.
Paws: Paraphrase adversaries from word scrambling, 2019
Yuan Zhang, Jason Baldridge, and Luheng He · 2019
Earlier work this paper cites.
Addressing power dynamics in community-engaged research partnerships
Lauri Andress, Tristen Hall, Sheila Davis, Judith Levine, Kimberly Cripps, and Dominique Guinn · 2020
Earlier work this paper cites.
Beat the ai: Investigating adversarial human annotation for reading comprehension
Max Bartolo, Alastair Roberts, Johannes Welbl, Sebastian Riedel, and Pontus Stenetorp · 2020
Earlier work this paper cites.
Piqa: Reasoning about physical commonsense in natural language
Yonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao, and Yejin Choi · 2020
Earlier work this paper cites.
Still in need of norms: the state of the data in citizen science
Anne Bowser, Caren Cooper, Alex De Sherbinin, Andrea Wiggins, Peter Brenton, Tyng-Ruey Chuang, Elaine Faustman, Mordechai Haklay, and Metis Meloche · 2020
Earlier work this paper cites.
Language models are few-shot learners, 2020
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei · 2020
Earlier work this paper cites.
Language ID in the wild: Unexpected challenges on the path to a thousand-language web text corpus
Isaac Caswell, Theresa Breiner, Daan van Esch, and Ankur Bapna · 2020
Earlier work this paper cites.
Tydi qa: A benchmark for information-seeking question answering in typologically diverse languages
Jonathan H. Clark, Eunsol Choi, Michael Collins, Dan Garrette, Tom Kwiatkowski, Vitaly Nikolaev, and Jennimaria Palomaki · 2020
Earlier work this paper cites.
A survey of multilingual neural machine translation
Raj Dabre, Chenhui Chu, and Anoop Kunchukuttan · 2020
Earlier work this paper cites.
Participatory research for low-resourced machine translation: A case study in African languages
∀ \forall , Wilhelmina Nekoto, Vukosi Marivate, Tshinondiwa Matsila, Timi Fasubaa, Taiwo Fagbohungbe, Solomon Oluwole Akinola, Shamsuddeen Muhammad, Salomon Kabongo Kabenamualu, Salomey Osei, Freshia Sackey, Rubungo Andre Niyongabo, Ricky Macharm, Perez Ogayo, Orevaoghene Ahia, Musie Meressa Berhe, Mofetoluwa Adeyemi, Masabata Mokgesi-Selinga, Lawrence Okegbemi, Laura Martinus, Kolawole Tajudeen, Kevin Degila, Kelechi Ogueji, Kathleen Siminyu, Julia Kreutzer, Jason Webster, Jamiil Toure Ali, Jade Abbott, Iroro Orife, Ignatius Ezeani, Idris Abdulkadir Dangana, Herman Kamper, Hady Elsahar, Goodness Duru, Ghollah Kioko, Murhabazi Espoir, Elan van Biljon, Daniel Whitenack, Christopher Onyefuluchi, Chris Chinenye Emezue, Bonaventure F. P. Dossou, Blessing Sibanda, Blessing Bassey, Ayodele Olabiyi, Arshath Ramkilowan, Alp Öktem, Adewale Akinfaderin, and Abdallah Bashir · 2020
Earlier work this paper cites.
Crowdsourcing Latin American Spanish for low-resource text-to-speech
Adriana Guevara-Rukoz, Isin Demirsahin, Fei He, Shan-Hui Cathy Chu, Supheakmungkol Sarin, Knot Pipatsrisawat, Alexander Gutkin, Alena Butryna, and Oddur Kjartansson · 2020
Earlier work this paper cites.
An ethical highlighter for people-centric dataset creation
Margot Hanley, Apoorv Khandelwal, Hadar Averbuch-Elor, Noah Snavely, and Helen Nissenbaum · 2020
Cited alongside, same era.
The state and fate of linguistic diversity and inclusion in the NLP world
Pratik Joshi, Sebastin Santy, Amar Budhiraja, Kalika Bali, and Monojit Choudhury · 2020
Cited alongside, same era.
Unifiedqa: Crossing format boundaries with a single qa system
Daniel Khashabi, Sewon Min, Tushar Khot, Ashish Sabharwal, Oyvind Tafjord, Peter Clark, and Hannaneh Hajishirzi · 2020
Cited alongside, same era.
Qasc: A dataset for question answering via sentence composition
Tushar Khot, Peter Clark, Michal Guerquin, Peter Jansen, and Ashish Sabharwal · 2020
Cited alongside, same era.
Open science and its enemies: Challenges for a sustainable science–society social contract
Lij news instruct lij-ita
ConseggioLigure · 2023
Later among the works it cites.
Seed instruct eng-lij
ConseggioLigure · 2023
Later among the works it cites.
Seed instruct lij-eng
ConseggioLigure · 2023
Later among the works it cites.
Power and public participation in ai
Eric Corbett, Emily Denton, and Sheena Erete · 2023
Later among the works it cites.
Efficient and effective text encoding for chinese llama and alpaca, 2023
Yiming Cui, Ziqing Yang, and Xin Yao · 2023
Later among the works it cites.
Okapi: Instruction-tuned large language models in multiple languages with reinforcement learning from human feedback
Viet Dac Lai, Chien Van Nguyen, Nghia Trung Ngo, Thuat Nguyen, Franck Dernoncourt, Ryan A Rossi, and Thien Huu Nguyen · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Venni V. Krishna · 2020
Cited alongside, same era.
WikiLingua: A new benchmark dataset for cross-lingual abstractive summarization
Faisal Ladhak, Esin Durmus, Claire Cardie, and Kathleen McKeown · 2020
Cited alongside, same era.
Mlqa: Evaluating cross-lingual extractive question answering, 2020
Patrick Lewis, Barlas Oğuz, Ruty Rinott, Sebastian Riedel, and Holger Schwenk · 2020
Cited alongside, same era.
CommonGen: A constrained text generation challenge for generative commonsense reasoning
Bill Yuchen Lin, Wangchunshu Zhou, Ming Shen, Pei Zhou, Chandra Bhagavatula, Yejin Choi, and Xiang Ren · 2020
Cited alongside, same era.
Investigating an approach for low resource language dataset creation, curation and classification: Setswana and sepedi
Vukosi Marivate, Tshephisho Sefara, Vongani Chabalala, Keamogetswe Makhaya, Tumisho Mokgonyane, Rethabile Mokoena, and Abiodun Modupe · 2020
Cited alongside, same era.
4 linguistic linked open data and under-resourced languages: From collection to application
Steven Moran and Christian Chiarcos · 2020
Cited alongside, same era.
Caring for data: Value creation in a data-intensive research laboratory
McKevitt C Pinel C, Prainsack B · 2020
Cited alongside, same era.
Xcopa: A multilingual dataset for causal commonsense reasoning
Edoardo Maria Ponti, Goran Glavaš, Olga Majewska, Qianchu Liu, Ivan Vulić, and Anna Korhonen · 2020
Cited alongside, same era.
The participatory turn in ai design: Theoretical foundations and the current state of practice
Fernando Delgado, Stephen Yang, Michael Madaio, and Qian Yang · 2023
Later among the works it cites.
Multilingual jailbreak challenges in large language models
Yue Deng, Wenxuan Zhang, Sinno Jialin Pan, and Lidong Bing · 2023
Later among the works it cites.
Telugu Riddles
desik98 · 2023
Later among the works it cites.
Towards leaving no Indic language behind: Building monolingual corpora, benchmark and models for Indic languages
Sumanth Doddapaneni, Rahul Aralikatte, Gowtham Ramesh, Shreya Goyal, Mitesh M. Khapra, Anoop Kunchukuttan, and Pratyush Kumar · 2023
Later among the works it cites.
A Seza Doğruöz, Sunayana Sitaram, and Zheng-Xin Yong · 2023
Later among the works it cites.
Towards measuring the representation of subjective global opinions in language models
Esin Durmus, Karina Nyugen, Thomas I Liao, Nicholas Schiefer, Amanda Askell, Anton Bakhtin, Carol Chen, Zac Hatfield-Dodds, Danny Hernandez, Nicholas Joseph, et al · 2023
Later among the works it cites.
Aya Indic Sentiment
el2e10 · 2023
Later among the works it cites.
Aya Paraphrase
el2e10 · 2023
Later among the works it cites.
Should chatgpt be biased? challenges and risks of bias in large language models
Emilio Ferrara · 2023
Later among the works it cites.
Hindi Article Summarization
ganeshjcs · 2023
Later among the works it cites.
Hindi Headline Article Generation
ganeshjcs · 2023
Later among the works it cites.
Measures to sustain endangered languages: A bilingual competition model with sliding mode control
Ya Gao and WenQi Liu · 2023
Later among the works it cites.
Chatgpt perpetuates gender bias in machine translation and ignores non-gendered pronouns: Findings across bengali and five other low-resource languages, 2023
Sourojit Ghosh and Aylin Caliskan · 2023
Later among the works it cites.
Textbooks are all you need, 2023
Suriya Gunasekar, Yi Zhang, Jyoti Aneja, Caio César Teodoro Mendes, Allie Del Giorno, Sivakanth Gopi, Mojan Javaheripi, Piero Kauffmann, Gustavo de Rosa, Olli Saarikivi, Adil Salim, Shital Shah, Harkirat Singh Behl, Xin Wang, Sébastien Bubeck, Ronen Eldan, Adam Tauman Kalai, Yin Tat Lee, and Yuanzhi Li · 2023
Later among the works it cites.
Targen: Targeted data generation with large language models, 2023
Himanshu Gupta, Kevin Scaria, Ujjwala Anantheswaran, Shreyas Verma, Mihir Parmar, Saurabh Arjun Sawant, Chitta Baral, and Swaroop Mishra · 2023
Later among the works it cites.
Measuring sentiment bias in machine translation, 2023
Kai Hartung, Aaricia Herygers, Shubham Kurlekar, Khabbab Zakaria, Taylan Volkan, Sören Gröttrup, and Munir Georges · 2023
Later among the works it cites.
A material lens on coloniality in nlp, 2023
William Held, Camille Harris, Michael Best, and Diyi Yang · 2023
Later among the works it cites.
FarsTail-Instruct-LLM
hghader1 · 2023
Later among the works it cites.
Making instruction finetuning accessible to non-English languages: A case study on Swedish models
Oskar Holmström and Ehsan Doostmohammadi · 2023
Later among the works it cites.
Unnatural instructions: Tuning language models with (almost) no human labor
Or Honovich, Thomas Scialom, Omer Levy, and Timo Schick · 2023
Later among the works it cites.
Acegpt, localizing large language models in arabic
Huang Huang, Fei Yu, Jianqing Zhu, Xuening Sun, Hao Cheng, Dingjie Song, Zhihong Chen, Abdulmohsen Alharthi, Bang An, Ziche Liu, et al · 2023
Later among the works it cites.
Indonesian Instruct Stories
Iftitahu · 2023
Later among the works it cites.
Javanese Instruct Stories
Iftitahu · 2023
Later among the works it cites.
Sudanese Instruct Stories
Iftitahu · 2023
Later among the works it cites.
You have got to know your language to understand your culture
Amirova Gulruh Ilhomovna and S Yuldasheva · 2023
Later among the works it cites.
IMDB Dutch Instruct
jjzha · 2023
Later among the works it cites.
GuanacoDataset (Revision 8cf0d29) , 2023
Joseph Cheung · 2023
Later among the works it cites.
Khyati Khandelwal, Manuel Tonneau, Andrew M. Bean, Hannah Rose Kirk, and Scott A. Hale · 2023
Later among the works it cites.
Gptaraeval: A comprehensive evaluation of chatgpt on arabic nlp, 2023
Md Tawkat Islam Khondaker, Abdul Waheed, El Moatez Billah Nagoudi, and Muhammad Abdul-Mageed · 2023
Later among the works it cites.
Seungone Kim, Se June Joo, Doyoung Kim, Joel Jang, Seonghyeon Ye, Jamin Shin, and Minjoon Seo · 2023
Later among the works it cites.
Openassistant conversations–democratizing large language model alignment
Andreas Köpf, Yannic Kilcher, Dimitri von Rütte, Sotiris Anagnostidis, Zhi-Rui Tam, Keith Stevens, Abdullah Barhoum, Nguyen Minh Duc, Oliver Stanley, Richárd Nagyfi, et al · 2023
Later among the works it cites.
Gender bias and stereotypes in large language models
Hadas Kotek, Rikker Dockum, and David Q. Sun · 2023
Later among the works it cites.
Madlad-400: A multilingual and document-level large audited dataset
Sneha Kudugunta, Isaac Caswell, Biao Zhang, Xavier Garcia, Christopher A Choquette-Choo, Katherine Lee, Derrick Xin, Aditya Kusupati, Romi Stella, Ankur Bapna, et al · 2023
Later among the works it cites.
Openassistant conversations – democratizing large language model alignment, 2023
Andreas Köpf, Yannic Kilcher, Dimitri von Rütte, Sotiris Anagnostidis, Zhi-Rui Tam, Keith Stevens, Abdullah Barhoum, Nguyen Minh Duc, Oliver Stanley, Richárd Nagyfi, Shahul ES, Sameer Suri, David Glushkov, Arnav Dantuluri, Andrew Maguire, Christoph Schuhmann, Huu Nguyen, and Alexander Mattick · 2023
Later among the works it cites.
Improving diversity of demographic representation in large language models via collective-critiques and self-voting, 2023
Preethi Lahoti, Nicholas Blumm, Xiao Ma, Raghavendra Kotikalapudi, Sahitya Potluri, Qijun Tan, Hansa Srinivasan, Ben Packer, Ahmad Beirami, Alex Beutel, and Jilin Chen · 2023
Later among the works it cites.
Viet Dac Lai, Chien Van Nguyen, Nghia Trung Ngo, Thuat Nguyen, Franck Dernoncourt, Ryan A Rossi, and Thien Huu Nguyen · 2023
Later among the works it cites.
Hate speech classifiers are culturally insensitive
Nayeon Lee, Chani Jung, and Alice Oh · 2023
Later among the works it cites.
Understanding crowdsourcing in science
Regina Lenart-Gansiniec, Wojciech Czakon, Łukasz Sułkowski, and Jasna Pocek · 2023
Later among the works it cites.
Analyzing leakage of personally identifiable information in language models
Nils Lukas, A. Salem, Robert Sim, Shruti Tople, Lukas Wutschitz, and Santiago Zanella-B’eguelin · 2023
Later among the works it cites.
Wizardcoder: Empowering code large language models with evol-instruct, 2023
Ziyang Luo, Can Xu, Pu Zhao, Qingfeng Sun, Xiubo Geng, Wenxiang Hu, Chongyang Tao, Jing Ma, Qingwei Lin, and Daxin Jiang · 2023
Later among the works it cites.
Fingpt: Large generative models for a small language
Risto Luukkonen, Ville Komulainen, Jouni Luoma, Anni Eskelinen, Jenna Kanerva, Hanna-Mari Kupari, Filip Ginter, Veronika Laippala, Niklas Muennighoff, Aleksandra Piktus, et al · 2023
Later among the works it cites.
Small data, big impact: Leveraging minimal data for effective machine translation
Jean Maillard, Cynthia Gao, Elahe Kalbassi, Kaushik Ram Sadagopan, Vedanuj Goswami, Philipp Koehn, Angela Fan, and Francisco Guzmán · 2023
Later among the works it cites.
When less is more: Investigating data pruning for pretraining llms at scale, 2023
Max Marion, Ahmet Üstün, Luiza Pozzobon, Alex Wang, Marzieh Fadaee, and Sara Hooker · 2023
Later among the works it cites.
AdversarialQA D(BERT)
maxbartolo · 2023
Later among the works it cites.
AdversarialQA D(BiDAF)
maxbartolo · 2023
Later among the works it cites.
AdversarialQA D(RoBERTa)
maxbartolo · 2023
Later among the works it cites.
Universal NER: A Gold-Standard Multilingual Named Entity Recognition Benchmark
Stephen Mayhew, Terra Blevins, Shuheng Liu, Marek Šuppa, Hila Gonen, Joseph Marvin Imperial, Börje F. Karlsson, Peiqin Lin, Nikola Ljubešić, LJ Miranda, Barbara Plank, Arij Riabi, and Yuval Pinter · 2023
Later among the works it cites.
Crosslingual generalization through multitask finetuning
Niklas Muennighoff, Thomas Wang, Lintang Sutawika, Adam Roberts, Stella Biderman, Teven Le Scao, M Saiful Bari, Sheng Shen, Zheng Xin Yong, Hailey Schoelkopf, Xiangru Tang, Dragomir Radev, Alham Fikri Aji, Khalid Almubarak, Samuel Albanie, Zaid Alyafeai, Albert Webson, Edward Raff, and Colin Raffel · 2023
Later among the works it cites.
AfriSenti: A Twitter sentiment analysis benchmark for African languages
Shamsuddeen Muhammad, Idris Abdulmumin, Abinew Ayele, Nedjma Ousidhoum, David Adelani, Seid Yimam, Ibrahim Ahmad, Meriem Beloucif, Saif Mohammad, Sebastian Ruder, Oumaima Hourrane, Alipio Jorge, Pavel Brazdil, Felermino Ali, Davis David, Salomey Osei, Bello Shehu-Bello, Falalu Lawan, Tajuddeen Gwadabe, Samuel Rutunda, Tadesse Belay, Wendimu Messelle, Hailu Balcha, Sisay Chala, Hagos Gebremichael, Bernard Opoku, and Stephen Arthur · 2023
Later among the works it cites.
Diversity of thought improves reasoning abilities of large language models, 2023
Ranjita Naik, Varun Chandrasekaran, Mert Yuksekgonul, Hamid Palangi, and Besmira Nushi · 2023
Later among the works it cites.
Three pathways to better recognize the expertise of global south researchers
Gabriel Nakamura, Bruno Soares, Valério Pillar, José Diniz-Filho, and Leandro Duarte · 2023
Later among the works it cites.
Having beer after prayer? measuring cultural bias in large language models
Tarek Naous, Michael Joseph Ryan, and Wei Xu · 2023
Later among the works it cites.
Scalable extraction of training data from (production) language models, 2023
Milad Nasr, Nicholas Carlini, Jonathan Hayase, Matthew Jagielski, A. Feder Cooper, Daphne Ippolito, Christopher A. Choquette-Choo, Eric Wallace, Florian Tramèr, and Katherine Lee · 2023
Later among the works it cites.
The open instruction generalist (oig) dataset
Huu Nguyen, Sameer Suri, Ken Tsui, and Christoph Schuhmann · 2023
Later among the works it cites.
Instruction in the wild: A user-based instruction dataset
Jinjie Ni, Fuzhao Xue, Kabir Jain, Mahir Hitesh Shah, Zangwei Zheng, and Yang You · 2023
Later among the works it cites.
Lost in translation: Large language models in non-english content analysis, 2023
Gabriel Nicholas and Aliya Bhatia · 2023
Later among the works it cites.
Multilegalpile: A 689gb multilingual legal corpus
Joel Niklaus, Veton Matoshi, Matthias Stürmer, Ilias Chalkidis, and Daniel E Ho · 2023
Later among the works it cites.
Toxic bias: Perspective api misreads german as more toxic, 2023
Gianluca Nogara, Francesco Pierri, Stefano Cresci, Luca Luceri, Petter Törnberg, and Silvia Giordano · 2023
Later among the works it cites.
Afriqa: Cross-lingual open-retrieval question answering for african languages
Odunayo Ogundepo, Tajuddeen Gwadabe, Clara Rivera, Jonathan Clark, Sebastian Ruder, David Adelani, Bonaventure Dossou, Abdou Diop, Claytone Sikasote, Gilles Hacheme, Happy Buzaaba, Ignatius Ezeani, Rooweither Mabuya, Salomey Osei, Chris Emezue, Albert Kahira, Shamsuddeen Muhammad, Akintunde Oladipo, Abraham Owodunni, Atnafu Tonja, Iyanuoluwa Shode, Akari Asai, Anuoluwapo Aremu, Ayodele Awokoya, Bernard Opoku, Chiamaka Chukwuneke, Christine Mwase, Clemencia Siro, Stephen Arthur, Tunde Ajayi, Verrah Otiende, Andre Rubungo, Boyd Sinkala, Daniel Ajisafe, Emeka Onwuegbuzia, Falalu Lawan, Ibrahim Ahmad, Jesujoba Alabi, Chinedu Mbonu, Mofetoluwa Adeyemi, Mofya Phiri, Orevaoghene Ahia, Ruqayya Iro, and Sonia Adhiambo · 2023
Later among the works it cites.
How good are large language models on african languages?, 2023
Jessica Ojo, Kelechi Ogueji, Pontus Stenetorp, and David I. Adelani · 2023
Later among the works it cites.
UA-GEC instruction tuning
osyvokon · 2023
Later among the works it cites.
Papers and patents are becoming less disruptive over time
Michael Park, Erin Leahey, and Russell J. Funk · 2023
Later among the works it cites.
scb_mt_2020_en2th_prompt
PyThaiNLP · 2023
Later among the works it cites.
scb_mt_2020_th2en_prompt
PyThaiNLP · 2023
Later among the works it cites.
Thai-Pos-prompt
PyThaiNLP · 2023
Later among the works it cites.
thai_usembassy_en2th_prompt
PyThaiNLP · 2023
Later among the works it cites.
thai_usembassy_th2en_prompt
PyThaiNLP · 2023
Later among the works it cites.
thai-wiktionary-prompt
PyThaiNLP · 2023
Later among the works it cites.
Chatgpt mt: Competitive for high- (but not low-) resource languages
Nathaniel R. Robinson, Perez Ogayo, David R. Mortensen, and Graham Neubig · 2023
Later among the works it cites.
Pdftriage: Question answering over long, structured documents
Jon Saad-Falcon, Joe Barrow, Alexa Siu, Ani Nenkova, Ryan A Rossi, and Franck Dernoncourt · 2023
Later among the works it cites.
Aya Persian Instruction pn Summary
Shafagh · 2023
Later among the works it cites.
Aya Persian Instruction pn Summary Title
Shafagh · 2023
Later among the works it cites.
Dialect-robust evaluation of generated text
Jiao Sun, Thibault Sellam, Elizabeth Clark, Tu Vu, Timothy Dozat, Dan Garrette, Aditya Siddhant, Jacob Eisenstein, and Sebastian Gehrmann · 2023
Later among the works it cites.
Aya Telugu Food Recipes
SuryaKrishna02 · 2023
Later among the works it cites.
Aya Telugu Jokes
SuryaKrishna02 · 2023
Later among the works it cites.
Aya Telugu News Articles
SuryaKrishna02 · 2023
Later among the works it cites.
Aya Telugu Paraphrase
SuryaKrishna02 · 2023
Later among the works it cites.
Aya Telugu Poems
SuryaKrishna02 · 2023
Later among the works it cites.
From base to conversational: Japanese instruction dataset and tuning large language models
Masahiro Suzuki, Masanori Hirano, and Hiroki Sakaji · 2023
Later among the works it cites.
Arpa aya
syntaxshill · 2023
Later among the works it cites.
UA-GEC: Grammatical error correction and fluency corpus for the Ukrainian language
Oleksiy Syvokon, Olena Nahorna, Pavlo Kuchmiichuk, and Nastasiia Osidach · 2023
Later among the works it cites.
Annotated News Summary
TahmidH · 2023
Later among the works it cites.
Stanford alpaca: An instruction-following llama model
Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto · 2023
Later among the works it cites.
LLM Japanese Dataset Vanilla Aya Format
Tellarin.ai · 2023
Later among the works it cites.
NTX LLM Instructions
Tellarin.ai · 2023
Later among the works it cites.
Joke explaination
theblackcat102 · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models, 2023
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Dan Bikel, Lukas Blecher, Cristian Canton Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fernandes, Jeremy Fu, Wenyin Fu, Brian Fuller, Cynthia Gao, Vedanuj Goswami, Naman Goyal, Anthony Hartshorn, Saghar Hosseini, Rui Hou, Hakan Inan, Marcin Kardas, Viktor Kerkez, Madian Khabsa, Isabel Kloumann, Artem Korenev, Punit Singh Koura, Marie-Anne Lachaux, Thibaut Lavril, Jenya Lee, Diana Liskovich, Yinghai Lu, Yuning Mao, Xavier Martinet, Todor Mihaylov, Pushkar Mishra, Igor Molybog, Yixin Nie, Andrew Poulton, Jeremy Reizenstein, Rashi Rungta, Kalyan Saladi, Alan Schelten, Ruan Silva, Eric Michael Smith, Ranjan Subramanian, Xiaoqing Ellen Tan, Binh Tang, Ross Taylor, Adina Williams, Jian Xiang Kuan, Puxin Xu, Zheng Yan, Iliyan Zarov, Yuchen Zhang, Angela Fan, Melanie Kambadur, Sharan Narang, Aurelien Rodriguez, Robert Stojnic, Sergey Edunov, and Thomas Scialom · 2023
Later among the works it cites.
Turku Paraphrase Corpus
TurkuNLP · 2023
Later among the works it cites.
UNER LLM Instructions
Universal NER · 2023
Later among the works it cites.
Not enough data to pre-train your language model? MT to the rescue!
Gorka Urbizu, Iñaki San Vicente, Xabier Saralegi, and Ander Corral · 2023
Later among the works it cites.
On evaluating and mitigating gender biases in multilingual settings, 2023
Aniket Vashishtha, Kabir Ahuja, and Sunayana Sitaram · 2023
Later among the works it cites.
Poisoning language models during instruction tuning, 2023
Alexander Wan, Eric Wallace, Sheng Shen, and Dan Klein · 2023
Later among the works it cites.
Clinicalgpt: Large language models finetuned with diverse medical data and comprehensive evaluation, 2023
Guangyu Wang, Guoxing Yang, Zongxin Du, Longjun Fan, and Xiaohu Li · 2023
Later among the works it cites.
Polylm: An open source polyglot large language model
Xiangpeng Wei, Haoran Wei, Huan Lin, Tianhao Li, Pei Zhang, Xingzhang Ren, Mei Li, Yu Wan, Zhiwei Cao, Binbin Xie, et al · 2023
Later among the works it cites.
Llm-powered data augmentation for enhanced crosslingual performance
Chenxi Whitehouse, Monojit Choudhury, and Alham Fikri Aji · 2023
Later among the works it cites.
Voicing an algorithm: trials of strength in artificial intelligence research
Joseph Wilson · 2023
Later among the works it cites.
The decades progress on code-switching research in NLP: A systematic survey on trends and challenges
Genta Winata, Alham Fikri Aji, Zheng Xin Yong, and Thamar Solorio · 2023
Later among the works it cites.
NusaX: Multilingual parallel sentiment dataset for 10 Indonesian local languages
Genta Indra Winata, Alham Fikri Aji, Samuel Cahyawijaya, Rahmad Mahendra, Fajri Koto, Ade Romadhony, Kemal Kurniawan, David Moeljadi, Radityo Eko Prasojo, Pascale Fung, Timothy Baldwin, Jey Han Lau, Rico Sennrich, and Sebastian Ruder · 2023
Later among the works it cites.
Dataset pruning: Reducing training data by examining generalization influence, 2023
Shuo Yang, Zeke Xie, Hanyu Peng, Min Xu, Mingming Sun, and Ping Li · 2023
Later among the works it cites.
BLOOM+1: Adding language support to BLOOM for zero-shot prompting
Zheng Xin Yong, Hailey Schoelkopf, Niklas Muennighoff, Alham Fikri Aji, David Ifeoluwa Adelani, Khalid Almubarak, M Saiful Bari, Lintang Sutawika, Jungo Kasai, Ahmed Baruwa, Genta Winata, Stella Biderman, Edward Raff, Dragomir Radev, and Vassilina Nikoulina · 2023
Later among the works it cites.
Large language model as attributed training data generator: A tale of diversity and bias, 2023
Yue Yu, Yuchen Zhuang, Jieyu Zhang, Yu Meng, Alexander Ratner, Ranjay Krishna, Jiaming Shen, and Chao Zhang · 2023
Later among the works it cites.
Lima: Less is more for alignment
Chunting Zhou, Pengfei Liu, Puxin Xu, Srini Iyer, Jiao Sun, Yuning Mao, Xuezhe Ma, Avia Efrat, Ping Yu, Lili Yu, et al · 2023
Later among the works it cites.
OLMo: Accelerating the Science of Language Models
Dirk Groeneveld, Iz Beltagy, Pete Walsh, Akshita Bhagia, Rodney Kinney, Oyvind Tafjord, Ananya Harsh Jha, Hamish Ivison, Ian Magnusson, Yizhong Wang, Shane Arora, David Atkinson, Russell Authur, Khyathi Chandu, Arman Cohan, Jennifer Dumas, Yanai Elazar, Yuling Gu, Jack Hessel, Tushar Khot, William Merrill, Jacob Morrison, Niklas Muennighoff, Aakanksha Naik, Crystal Nam, Matthew E. Peters, Valentina Pyatkin, Abhilasha Ravichander, Dustin Schwenk, Saurabh Shah, Will Smith, Nishant Subramani, Mitchell Wortsman, Pradeep Dasigi, Nathan Lambert, Kyle Richardson, Jesse Dodge, Kyle Lo, Luca Soldaini, Noah A. Smith, and Hannaneh Hajishirzi · 2024
Closest in time.
Aboutme: Using self-descriptions in webpages to document the effects of english pretraining data filters, 2024
Li Lucy, Suchin Gururangan, Luca Soldaini, Emma Strubell, David Bamman, Lauren Klein, and Jesse Dodge · 2024
Closest in time.
Explainer: Why is Myanmar’s military holding an election?, 2023
Martin Petty · 2024
Closest in time.
Explainer: What is happening between Armenia and Azerbaijan over Nagorno-Karabakh?, 2023
Reuters · 2024
Closest in time.
Dolma: An Open Corpus of Three Trillion Tokens for Language Model Pretraining Research
Luca Soldaini, Rodney Kinney, Akshita Bhagia, Dustin Schwenk, David Atkinson, Russell Authur, Ben Bogin, Khyathi Chandu, Jennifer Dumas, Yanai Elazar, Valentin Hofmann, Ananya Harsh Jha, Sachin Kumar, Li Lucy, Xinxi Lyu, Ian Magnusson, Jacob Morrison, Niklas Muennighoff, Aakanksha Naik, Crystal Nam, Matthew E. Peters, Abhilasha Ravichander, Kyle Richardson, Zejiang Shen, Emma Strubell, Nishant Subramani, Oyvind Tafjord, Evan Pete Walsh, Hannaneh Hajishirzi, Noah A. Smith, Luke Zettlemoyer, Iz Beltagy, Dirk Groeneveld, Jesse Dodge, and Kyle Lo · 2024
Closest in time.
Astraios: Parameter-efficient instruction tuning code large language models
Terry Yue Zhuo, Armel Zebaze, Nitchakarn Suppattarachai, Leandro von Werra, Harm de Vries, Qian Liu, and Niklas Muennighoff · 2024
Closest in time.