Fetching the paper…
Reading the bibliography…
The HuggingFace Datasets Hub hosts thousands of datasets, offering exciting opportunities for language model training and evaluation.
Equate: A benchmark evaluation framework for quantitative reasoning in natural language inference
Abhilasha Ravichander, Aakanksha Naik, Carolyn Rose, and Eduard Hovy. 2019 · 1901
Earlier work this paper cites.
Right for the wrong reasons: Diagnosing syntactic heuristics in natural language inference
R. Thomas McCoy, Ellie Pavlick, and Tal Linzen. 2019 · 1902
Earlier work this paper cites.
Nuanced metrics for measuring unintended bias with real data for text classification
Daniel Borkan, Lucas Dixon, Jeffrey Sorensen, Nithum Thain, and Lucy Vasserman. 2019 · 1903
Earlier work this paper cites.
Nabil Hossain, John Krumm, and Michael Gamon. 2019 · 1906
Earlier work this paper cites.
Wiqa: A dataset for "what if…" reasoning over procedural text
Niket Tandon, Bhavana Dalvi Mishra, Keisuke Sakaguchi, Antoine Bosselut, and Peter Clark. 2019 · 1909
Earlier work this paper cites.
Qasc: A dataset for question answering via sentence composition
Tushar Khot, Peter Clark, Michal Guerquin, Peter Jansen, and Ashish Sabharwal. 2020 · 1910
Earlier work this paper cites.
Blimp: A benchmark of linguistic minimal pairs for english
Alex Warstadt, Alicia Parrish, Haokun Liu, Anhad Mohananey, Wei Peng, Sheng-Fu Wang, and Samuel R Bowman. 2019 · 1912
Earlier work this paper cites.
The argument reasoning comprehension task: Identification and reconstruction of implicit warrants
Ivan Habernal, Henning Wachsmuth, Iryna Gurevych, and Benno Stein. 2018 · 1940
Earlier work this paper cites.
CrowS-pairs: A challenge dataset for measuring social biases in masked language models
Nikita Nangia, Clara Vania, Rasika Bhalerao, and Samuel R. Bowman. 2020 · 1967
Earlier work this paper cites.
Switchboard: Telephone speech corpus for research and development
John J. Godfrey, Edward C. Holliman, and Jane McDaniel. 1992 · 1992
Earlier work this paper cites.
Multitask learning: A knowledge-based source of inductive bias
Rich Caruana. 1993 · 1993
Earlier work this paper cites.
The hcrc map task corpus: natural dialogue for speech recognition
Henry S Thompson, Anne H Anderson, Ellen Gurman Bard, Gwyneth Doherty-Sneddon, Alison Newlands, and Cathy Sotillo. 1993 · 1993
Earlier work this paper cites.
Learning question classifiers
Xin Li and Dan Roth. 2002 · 2002
Earlier work this paper cites.
Generic speech act annotation for task-oriented dialogues
Geoffrey Leech and Martin Weisser. 2003 · 2003
Earlier work this paper cites.
Introduction to the CoNLL-2003 shared task: Language-independent named entity recognition
Erik F. Tjong Kim Sang and Fien De Meulder. 2003 · 2003
Earlier work this paper cites.
Introduction to the bio-entity recognition task at jnlpba
Jin-Dong Kim, Tomoko Ohta, Yoshimasa Tsuruoka, Yuka Tateisi, and Nigel Collier. 2004 · 2004
Earlier work this paper cites.
The icsi meeting recorder dialog act (mrda) corpus
Elizabeth Shriberg, Raj Dhillon, Sonali Bhagat, Jeremy Ang, and Hannah Carvey. 2004 · 2004
Earlier work this paper cites.
Seeing stars: Exploiting class relationships for sentiment categorization with respect to rating scales
Bo Pang and Lillian Lee. 2005 · 2005
Earlier work this paper cites.
Effects of age and gender on blogging
Jonathan Schler, Moshe Koppel, Shlomo Argamon, and James W Pennebaker. 2006 · 2006
Earlier work this paper cites.
Logiqa: A challenge dataset for machine reading comprehension with logical reasoning
Jian Liu, Leyang Cui, Hanmeng Liu, Dandan Huang, Yile Wang, and Yue Zhang. 2020 · 2007
Earlier work this paper cites.
Iemocap: Interactive emotional dyadic motion capture database
Carlos Busso, Murtaza Bulut, Chi-Chun Lee, Abe Kazemzadeh, Emily Mower, Samuel Kim, Jeannette N Chang, Sungbok Lee, and Shrikanth S Narayanan. 2008 · 2008
Earlier work this paper cites.
Aligning ai with shared human values
Dan Hendrycks, Collin Burns, Steven Basart, Andrew Critch, Jerry Li, Dawn Song, and Jacob Steinhardt. 2020 · 2008
Earlier work this paper cites.
The Penn Discourse TreeBank 2.0
Rashmi Prasad, Nikhil Dinesh, Alan Lee, Eleni Miltsakaki, Livio Robaldo, Aravind Joshi, and Bonnie Webber. 2008 · 2008
Earlier work this paper cites.
SemEval-2010 task 8: Multi-way classification of semantic relations between pairs of nominals
Iris Hendrickx, Su Nam Kim, Zornitsa Kozareva, Preslav Nakov, Diarmuid O Seaghdha, Sebastian Pado, Marco Pennacchiotti, Lorenza Romano, and Stan Szpakowicz. 2010 · 2010
Earlier work this paper cites.
Contributions to the study of sms spam filtering: New collection and results
Tiago A. Almeida, Jose Maria Gomez Hidalgo, and Akebo Yamakami. 2011 · 2011
Earlier work this paper cites.
Do fine-tuned commonsense language models really generalize?
Mayank Kejriwal and Ke Shen. 2020 · 2011
Earlier work this paper cites.
Learning word vectors for sentiment analysis
Andrew L. Maas, Raymond E. Daly, Peter T. Pham, Dan Huang, Andrew Y. Ng, and Christopher Potts. 2011 · 2011
Earlier work this paper cites.
The semaine database: Annotated multimodal records of emotionally colored conversations between a person and a limited agent
Gary McKeown, Michel Valstar, Roddy Cowie, Maja Pantic, and Marc Schroder. 2011 · 2011
Earlier work this paper cites.
Choice of plausible alternatives: An evaluation of commonsense causal reasoning
Melissa Roemmele, Cosmin Adrian Bejan, and Andrew S Gordon. 2011 · 2011
Earlier work this paper cites.
Moral stories: Situated reasoning about norms, intents, actions and their consequences
Denis Emelin, Ronan Le Bras, Jena D. Hwang, Maxwell Forbes, and Yejin Choi. 2021 · 2012
Earlier work this paper cites.
Development of a benchmark corpus to support the automatic extraction of drug-related adverse effects from medical case reports
Harsha Gurulingappa, Abdul Mateen Rajput, Angus Roberts, Juliane Fluck, Martin Hofmann-Apitius, and Luca Toldo. 2012 · 2012
Earlier work this paper cites.
DynaSent: A dynamic benchmark for sentiment analysis
Christopher Potts, Zhengxuan Wu, Atticus Geiger, and Douwe Kiela. 2020 · 2012
Earlier work this paper cites.
Resolving complex cases of definite pronouns: the winograd schema challenge
Altaf Rahman and Vincent Ng. 2012 · 2012
Earlier work this paper cites.
Designing a scalable crowdsourcing platform
Chris Van Pelt and Alex Sorokin. 2012 · 2012
Earlier work this paper cites.
Hidden factors and hidden topics: understanding rating dimensions with review text
Julian McAuley and Jure Leskovec. 2013 · 2013
Earlier work this paper cites.
The species and organisms resources for fast and accurate identification of taxonomic names in text
Evangelos Pafilis, Sune P Frankild, Lucia Fanini, Sarah Faulwetter, Christina Pavloudi, Aikaterini Vasileiadou, Christos Arvanitidis, and Lars Juhl Jensen. 2013 · 2013
Earlier work this paper cites.
Ncbi disease corpus: a resource for disease name recognition and concept normalization
Rezarta Islamaj Dogan, Robert Leaman, and Zhiyong Lu. 2014 · 2014
Earlier work this paper cites.
Good debt or bad debt: Detecting semantic orientations in economic texts
P. Malo, A. Sinha, P. Korhonen, J. Wallenius, and P. Takala. 2014 · 2014
Earlier work this paper cites.
A SICK cure for the evaluation of compositional distributional semantic models
Marco Marelli, Stefano Menini, Marco Baroni, Luisa Bentivogli, Raffaella Bernardi, and Roberto Zamparelli. 2014 · 2014
Earlier work this paper cites.
Identifying appropriate support for propositions in online user comments
Joonsuk Park and Claire Cardie. 2014 · 2014
Earlier work this paper cites.
A large annotated corpus for learning natural language inference
Samuel R. Bowman, Gabor Angeli, Christopher Potts, and Christopher D. Manning. 2015 · 2015
Earlier work this paper cites.
SQUINKY! A Corpus of Sentence-level Formality, Informativeness, and Implicature
Shibamouli Lahiri. 2015 · 2015
Earlier work this paper cites.
Unsupervised prediction of acceptability judgements
Jey Han Lau, Alexander Clark, and Shalom Lappin. 2015 · 2015
Earlier work this paper cites.
Dbpedia–a large-scale, multilingual knowledge base extracted from wikipedia
Jens Lehmann, Robert Isele, Max Jakob, Anja Jentzsch, Dimitris Kontokostas, Pablo N Mendes, Sebastian Hellmann, Mohamed Morsey, Patrick Van Kleef, Soren Auer, et al. 2015 · 2015
Earlier work this paper cites.
Towards ai-complete question answering: A set of prerequisite toy tasks
Jason Weston, Antoine Bordes, Sumit Chopra, Alexander M Rush, Bart Van Merrienboer, Armand Joulin, and Tomas Mikolov. 2015 · 2015
Earlier work this paper cites.
Discourse structure and dialogue acts in multiparty dialogue: the STAC corpus
Nicholas Asher, Julie Hunter, Mathieu Morey, Benamara Farah, and Stergos Afantenos. 2016 · 2016
Earlier work this paper cites.
Emergent: a novel data-set for stance classification
William Ferreira and Andreas Vlachos. 2016 · 2016
Earlier work this paper cites.
Fasttext.zip: Compressing text classification models
Armand Joulin, Edouard Grave, Piotr Bojanowski, Matthijs Douze, Hérve Jégou, and Tomas Mikolov. 2016 · 2016
Earlier work this paper cites.
Creating and characterizing a diverse corpus of sarcasm in dialogue
Shereen Oraby, Vrindavan Harrison, Lena Reed, Ernesto Hernandez, Ellen Riloff, and Marilyn Walker. 2016 · 2016
Earlier work this paper cites.
Program induction by rationale generation: Learning to solve and explain algebraic word problems
2017 · 2017
Earlier work this paper cites.
Software applications user reviews
2017 · 2017
Earlier work this paper cites.
EmoBank: Studying the impact of annotation perspective and representation format on dimensional emotion analysis
Sven Buechel and Udo Hahn. 2017 · 2017
Earlier work this paper cites.
SemEval-2017 task 1: Semantic textual similarity multilingual and crosslingual focused evaluation
Daniel Cer, Mona Diab, Eneko Agirre, Inigo Lopez-Gazpio, and Lucia Specia. 2017 · 2017
Earlier work this paper cites.
Supervised learning of universal sentence representations from natural language inference data
Alexis Conneau, Douwe Kiela, Holger Schwenk, Loïc Barrault, and Antoine Bordes. 2017 · 2017
Earlier work this paper cites.
Automated hate speech detection and the problem of offensive language
Thomas Davidson, Dana Warmsley, Michael Macy, and Ingmar Weber. 2017 · 2017
Earlier work this paper cites.
Results of the WNUT2017 shared task on novel and emerging entity recognition
Leon Derczynski, Eric Nichols, Marieke van Erp, and Nut Limsopatham. 2017 · 2017
Earlier work this paper cites.
Crowdsourcing multiple choice science questions
Matt Gardner Johannes Welbl, Nelson F. Liu. 2017 · 2017
Earlier work this paper cites.
Race: Large-scale reading comprehension dataset from examinations
Guokun Lai, Qizhe Xie, Hanxiao Liu, Yiming Yang, and Eduard Hovy. 2017 · 2017
Earlier work this paper cites.
Dailydialog: A manually labelled multi-turn dialogue dataset
Yanran Li, Hui Su, Xiaoyu Shen, Wenjie Li, Ziqiang Cao, and Shuzi Niu. 2017 · 2017
Earlier work this paper cites.
“liar, liar pants on fire”: A new benchmark dataset for fake news detection
William Yang Wang. 2017 · 2017
Earlier work this paper cites.
The GUM corpus: Creating multilayer resources in the classroom
Amir Zeldes. 2017 · 2017
Earlier work this paper cites.
Sentimental analysis of tweets for detecting hate/racist speeches
2018 · 2018
Earlier work this paper cites.
Give me more feedback: Annotating argument persuasiveness and related attributes in student essays
Winston Carlile, Nishant Gurrapadi, Zixuan Ke, and Vincent Ng. 2018 · 2018
Cited alongside, same era.
Emotionlines: An emotion corpus of multi-party conversations
Sheng-Yeh Chen, Chao-Chun Hsu, Chuan-Chun Kuo, Lun-Wei Ku, et al. 2018 · 2018
Cited alongside, same era.
Think you have solved question answering? try arc, the ai2 reasoning challenge
Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord. 2018 · 2018
Cited alongside, same era.
What you can cram into a single $&!#* vector: Probing sentence embeddings for linguistic properties
Alexis Conneau, German Kruszewski, Guillaume Lample, Loic Barrault, and Marco Baroni. 2018 · 2018
Cited alongside, same era.
Ethos: an online hate speech detection dataset
Ioannis Mollas, Zoe Chrysopoulou, Stamatis Karlos, and Grigorios Tsoumakas. 2020 · 2020
Later among the works it cites.
Adversarial nli: A new benchmark for natural language understanding
Yixin Nie, Adina Williams, Emily Dinan, Mohit Bansal, Jason Weston, and Douwe Kiela. 2020 · 2020
Later among the works it cites.
jiant: A software toolkit for research on general-purpose text understanding models
Yada Pruksachatkun, Phil Yeres, Haokun Liu, Jason Phang, Phu Mon Htut, Alex Wang, Ian Tenney, and Samuel R. Bowman. 2020 · 2020
Later among the works it cites.
Getting closer to AI complete question answering: A set of prerequisite real tasks
Anna Rogers, Olga Kovaleva, Matthew Downey, and Anna Rumshisky. 2020 · 2020
Later among the works it cites.
Investigating societal biases in a poetry composition system
Emily Sheng and David Uthus. 2020 · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Alice Coucke, Alaa Saade, Adrien Ball, Theodore Bluche, Alexandre Caulier, David Leroy, Clement Doumouro, Thibault Gisselbrecht, Francesco Caltagirone, Thibaut Lavril, Mael Primet, and Joseph Dureau. 2018 · 2018
Cited alongside, same era.
Hate Speech Dataset from a White Supremacy Forum
Ona de Gibert, Naiara Perez, Aitor Garcia-Pablos, and Montse Cuadros. 2018 · 2018
Cited alongside, same era.
Identifying well-formed natural language questions
Manaal Faruqui and Dipanjan Das. 2018 · 2018
Cited alongside, same era.
Looking beyond the surface:a challenge set for reading comprehension over multiple sentences
Daniel Khashabi, Snigdha Chaturvedi, Michael Roth, Shyam Upadhyay, and Dan Roth. 2018 · 2018
Cited alongside, same era.
SciTail: A textual entailment dataset from science question answering
Tushar Khot, Ashish Sabharwal, and Peter Clark. 2018 · 2018
Cited alongside, same era.
Can a suit of armor conduct electricity? a new dataset for open book question answering
Todor Mihaylov, Peter Clark, Tushar Khot, and Ashish Sabharwal. 2018 · 2018
Cited alongside, same era.
Wic: 10, 000 example pairs for evaluating context-sensitive representations
Mohammad Taher Pilehvar and ose Camacho-Collados. 2018 · 2018
Cited alongside, same era.
Collecting diverse natural language inference problems for sentence representation evaluation
Adam Poliak, Aparajita Haldar, Rachel Rudinger, J. Edward Hu, Ellie Pavlick, Aaron Steven White, and Benjamin Van Durme. 2018 · 2018
Cited alongside, same era.
Later among the works it cites.
Supreme Court Database, Version 2020 Release 01
Harold J. Spaeth, Lee Epstein, Jeffrey A. Segal Andrew D. Martin, Theodore J. Ruger, and Sara C. Benesh. 2020 · 2020
Later among the works it cites.
Learning to summarize from human feedback
Nisan Stiennon, Long Ouyang, Jeff Wu, Daniel M. Ziegler, Ryan Lowe, Chelsea Voss, Alec Radford, Dario Amodei, and Paul Christiano. 2020 · 2020
Later among the works it cites.
LEDGAR: A large-scale multi-label corpus for text classification of legal provisions in contracts
Don Tuggener, Pius von Daniken, Thomas Peetz, and Mark Cieliebak. 2020 · 2020
Later among the works it cites.
What Does This Acronym Mean? Introducing a New Dataset for Acronym Identification and Disambiguation
Amir Pouran Ben Veyseh, Franck Dernoncourt, Quan Hung Tran, and Thien Huu Nguyen. 2020 · 2020
Later among the works it cites.
Datasets
Thomas Wolf, Quentin Lhoest, Patrick von Platen, Yacine Jernite, Mariama Drame, Julien Plu, Julien Chaumond, Clement Delangue, Clara Ma, Abhishek Thakur, Suraj Patil, Joe Davison, Teven Le Scao, Victor Sanh, Canwen Xu, Nicolas Patry, Angie McMillan-Major, Simon Brandeis, Sylvain Gugger, François Lagunas, Lysandre Debut, Morgan Funtowicz, Anthony Moi, Sasha Rush, Philipp Schmidd, Pierric Cistac, Victor Muštar, Jeff Boudier, and Anna Tordjmann. 2020 · 2020
Later among the works it cites.
Universal dependencies 2.7
Daniel Zeman, Joakim Nivre, Mitchell Abrams, Elia Ackermann, Noemi Aepli, Hamid Aghaei, veljko Agic, Amir Ahmadi, Lars Ahrenberg, Chika Kennedy Ajede, Gabriele Aleksandraviviute, Ika Alfina, Lene Antonsen, Katya Aplonova, Angelina Aquino, Carolina Aragon, Maria Jesus Aranzabe, torunn Arnardottir, Gashaw Arutie, Jessica Naraiswari Arwidarasti, Masayuki Asahara, Luma Ateyah, Furkan Atmaca, Mohammed Attia, Aitziber Atutxa, Liesbeth Augustinus, Elena Badmaeva, Keerthana Balasubramani, Miguel Ballesteros, Esha Banerjee, Sebastian Bank, Verginica Barbu Mititelu, Victoria Basmov, Colin Batchelor, John Bauer, Seyyit Talha Bedir, Kepa Bengoetxea, Gozde Berk, Yevgeni Berzak, Irshad Ahmad Bhat, Riyaz Ahmad Bhat, Erica Biagetti, Eckhard Bick, Agne Bielinskiene, Kristin Bjarnadottir, Rogier Blokland, Victoria Bobicev, Loic Boizou, Emanuel Borges Volker, Carl Borstell, Cristina Bosco, Gosse Bouma, Sam Bowman, Adriane Boyd, Kristina Brokaite, Aljoscha Burchardt, Marie Candito, Bernard Caron, Gauthier Caron, Tatiana Cavalcanti, Gulsen Cebiroglu Eryigit, Flavio Massimiliano Cecchini, Giuseppe G. A. Celano, Slavomir Ceplo, Savas Cetin, Ozlem Cetinoglu, Fabricio Chalub, Ethan Chi, Yongseok Cho, Jinho Choi, Jayeol Chun, Alessandra T. Cignarella, Silvie Cinkova, Aurelie Collomb, Cagri Coltekin, Miriam Connor, Marine Courtin, Elizabeth Davidson, Marie-Catherine de Marneffe, Valeria de Paiva, Mehmet Oguz Derin, Elvis de Souza, Arantza Diaz de Ilarraza, Carly Dickerson, Arawinda Dinakaramani, Bamba Dione, Peter Dirix, Kaja Dobrovoljc, Timothy Dozat, Kira Droganova, Puneet Dwivedi, Hanne Eckhoff, Marhaba Eli, Ali Elkahky, Binyam Ephrem, Olga Erina, Tomaz Erjavec, Aline Etienne, Wograine Evelyn, Sidney Facundes, Richard Farkas, Marilia Fernanda, Hector Fernandez Alcalde, Jennifer Foster, Claudia Freitas, Kazunori Fujita, Katarina Gajdosova, Daniel Galbraith, Marcos Garcia, Moa Gardenfors, Sebastian Garza, Fabricio Ferraz Gerardi, Kim Gerdes, Filip Ginter, Iakes Goenaga, Koldo Gojenola, Memduh Gokirmak, Yoav Goldberg, Xavier Gomez Guinovart, Berta Gonzalez Saavedra, Bernadeta Griciute, Matias Grioni, Loic Grobol, Normunds Gruzitis, Bruno Guillaume, Celine Guillot-Barbance, Tunga Gungor, Nizar Habash, Hinrik Hafsteinsson, Jan Hajiv, Jan Hajiv jr., Mika Hamalainen, Linh Ha My, Na-Rae Han, Muhammad Yudistira Hanifmuti, Sam Hardwick, Kim Harris, Dag Haug, Johannes Heinecke, Oliver Hellwig, Felix Hennig, Barbora Hladka, Jaroslava Hlavavova, Florinel Hociung, Petter Hohle, Eva Huber, Jena Hwang, Takumi Ikeda, Anton Karl Ingason, Radu Ion, Elena Irimia, dlajide Ishola, Tomav Jelinek, Anders Johannsen, Hildur Jonsdottir, Fredrik Jorgensen, Markus Juutinen, Sarveswaran K, Huner Kacikara, Andre Kaasen, Nadezhda Kabaeva, Sylvain Kahane, Hiroshi Kanayama, Jenna Kanerva, Boris Katz, Tolga Kayadelen, Jessica Kenney, Vaclava Kettnerova, Jesse Kirchner, Elena Klementieva, Arne Kohn, Abdullatif Koksal, Kamil Kopacewicz, Timo Korkiakangas, Natalia Kotsyba, Jolanta Kovalevskaite, Simon Krek, Parameswari Krishnamurthy, Sookyoung Kwak, Veronika Laippala, Lucia Lam, Lorenzo Lambertino, Tatiana Lando, Septina Dian Larasati, Alexei Lavrentiev, John Lee, Phng Le Hong, Alessandro Lenci, Saran Lertpradit, Herman Leung, Maria Levina, Cheuk Ying Li, Josie Li, Keying Li, Yuan Li, K Lim, Krister Linden, Nikola Ljubesic, Olga Loginova, Andry Luthfi, Mikko Luukko, Olga Lyashevskaya, Teresa Lynn, Vivien Macketanz, Aibek Makazhanov, Michael Mandl, Christopher Manning, Ruli Manurung, Catalina Maranduc, David Marcek, Katrin Marheinecke, Hector Martinez Alonso, Andre Martins, Jan Masek, Hiroshi Matsuda, Yuji Matsumoto, Ryan M, Sarah M, Gustavo Mendonca, Niko Miekka, Karina Mischenkova, Margarita Misirpashayeva, Anna Missila, Catalin Mititelu, Maria Mitrofan, Yusuke Miyao, A Mojiri Foroushani, Amirsaeid Moloodi, Simonetta Montemagni, Amir More, Laura Moreno Romero, Keiko Sophie Mori, Shinsuke Mori, Tomohiko Morioka, Shigeki Moro, Bjartur Mortensen, Bohdan Moskalevskyi, Kadri Muischnek, Robert Munro, Yugo Murawaki, Kaili Muurisep, Pinkey Nainwani, Mariam Nakhle, Juan Ignacio Navarro Horniacek, Anna Nedoluzhko, Gunta Nevpore-Berzkalne, Lng Nguyen Thd, Huyen Nguyen Thd Minh, Yoshihiro Nikaido, Vitaly Nikolaev, Rattima Nitisaroj, Alireza Nourian, Hanna Nurmi, Stina Ojala, Atul Kr. Ojha, Adedayd Oluokun, Mai Omura, Emeka Onwuegbuzia, Petya Osenova, Robert Ostling, Lilja Ovrelid, caziye Betul Ozatec, Arzucan Ozgur, Balkiz Ozturk Bacaran, Niko Partanen, Elena Pascual, Marco Passarotti, Agnieszka Patejuk, Guilherme Paulino-Passos, Angelika Peljak-Lapinska, Siyao Peng, Cenel-Augusto Perez, Natalia Perkova, Guy Perrier, Slav Petrov, Daria Petrova, Jason Phelan, Jussi Piitulainen, Tommi A Pirinen, Emily Pitler, Barbara Plank, Thierry Poibeau, Larisa Ponomareva, Martin Popel, Lauma Pretkalnina, Sophie Prevost, Prokopis Prokopidis, Adam Przepiorkowski, Tiina Puolakainen, Sampo Pyysalo, Peng Qi, Andriela Raabis, Alexandre Rademaker, Taraka Rama, Loganathan Ramasamy, Carlos Ramisch, Fam Rashel, Mohammad Sadegh Rasooli, Vinit Ravishankar, Livy Real, Petru Rebeja, Siva Reddy, Georg Rehm, Ivan Riabov, Michael Riesler, Erika Rimkute, Larissa Rinaldi, Laura Rituma, Luisa Rocha, Eirikur Rognvaldsson, Mykhailo Romanenko, Rudolf Rosa, Valentin Roca, Davide Rovati, Olga Rudina, Jack Rueter, Kristjan Runarsson, Shoval Sadde, Pegah Safari, Benoit Sagot, Aleksi Sahala, Shadi Saleh, Alessio Salomoni, Tanja Samardzic, Stephanie Samson, Manuela Sanguinetti, Dage Sarg, Baiba Saulite, Yanin Sawanakunanon, Kevin Scannell, Salvatore Scarlata, Nathan Schneider, Sebastian Schuster, Djame Seddah, Wolfgang Seeker, Mojgan Seraji, Mo Shen, Atsuko Shimada, Hiroyuki Shirasu, Muh Shohibussirri, Dmitry Sichinava, Einar Freyr Sigursson, Aline Silveira, Natalia Silveira, Maria Simi, Radu Simionescu, Katalin Simko, Maria vimkova, Kiril Simov, Maria Skachedubova, Aaron Smith, Isabela Soares-Bastos, Carolyn Spadine, Steintor Steingrimsson, Antonio Stella, Milan Straka, Emmett Strickland, Jana Strnadova, Alane Suhr, Yogi Lesmana Sulestio, Umut Sulubacak, Shingo Suzuki, Zsolt Szanto, Dima Taji, Yuta Takahashi, Fabio Tamburini, Mary Ann C. Tan, Takaaki Tanaka, Samson Tella, Isabelle Tellier, Guillaume Thomas, Liisi Torga, Marsida Toska, Trond Trosterud, Anna Trukhina, Reut Tsarfaty, Utku Turk, Francis Tyers, Sumire Uematsu, Roman Untilov, Zdenka Uresova, Larraitz Uria, Hans Uszkoreit, Andrius Utka, Sowmya Vajjala, Daniel van Niekerk, Gertjan van Noord, Viktor Varga, Eric Villemonte de la Clergerie, Veronika Vincze, Aya Wakasa, Joel C. Wallenberg, Lars Wallin, Abigail Walsh, Jing Xian Wang, Jonathan North Washington, Maximilan Wendt, Paul Widmer, Seyi Williams, Mats Wiren, Christian Wittern, Tsegay Woldemariam, Tak-sum Wong, Alina Wroblewska, Mary Yako, Kayo Yamashita, Naoki Yamazaki, Chunxiao Yan, Koichi Yasuoka, Marat M. Yavrumyan, Zhuoran Yu, Zdenek Zabokrtsky, Shorouq Zahra, Amir Zeldes, Hanzhi Zhu, and Anna Zhuravleva. 2020 · 2020
Later among the works it cites.
Muppet: Massive multi-task representations with pre-finetuning
Armen Aghajanyan, Anchit Gupta, Akshat Shrivastava, Xilun Chen, Luke Zettlemoyer, and Sonal Gupta. 2021 · 2021
Later among the works it cites.
PROST: Physical reasoning about objects through space and time
Stephane Aroca-Ouellette, Cory Paik, Alessandro Roncone, and Katharina Kann. 2021 · 2021
Later among the works it cites.
Controllable open-ended question generation with a new question type ontology
Shuyang Cao and Lu Wang. 2021 · 2021
Later among the works it cites.
Multieurlex – a multi-lingual and multi-label legal document classification dataset for zero-shot cross-lingual transfer
Ilias Chalkidis, Manos Fergadiotis, and Ion Androutsopoulos. 2021 · 2021
Later among the works it cites.
Latent hatred: A benchmark for understanding implicit hate speech
Mai ElSherief, Caleb Ziems, David Muchlinski, Vaishnavi Anupindi, Jordyn Seybolt, Munmun De Choudhury, and Diyi Yang. 2021 · 2021
Later among the works it cites.
Headlinecause: A dataset of news headlines for detecting casualties
Ilya Gusev and Alexey Tikhonov. 2021 · 2021
Later among the works it cites.
Pengcheng He, Jianfeng Gao, and Weizhu Chen. 2021 · 2021
Later among the works it cites.
Measuring massive multitask language understanding
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. 2021 · 2021
Later among the works it cites.
News headline grouping as a challenging nlu task
Philippe Laban and Lucas Bandarkar. 2021 · 2021
Later among the works it cites.
SPARTQA: A textual question answering benchmark for spatial reasoning
Roshanak Mirzaee, Hossein Rajaby Faghihi, Qiang Ning, and Parisa Kordjamshidi. 2021 · 2021
Later among the works it cites.
Webgpt: Browser-assisted question-answering with human feedback
Reiichiro Nakano, Jacob Hilton, Suchir Balaji, Jeff Wu, Long Ouyang, Christina Kim, Christopher Hesse, Shantanu Jain, Vineet Kosaraju, William Saunders, Xu Jiang, Karl Cobbe, Tyna Eloundou, Gretchen Krueger, Kevin Button, Matthew Knight, Benjamin Chess, and John Schulman. 2021 · 2021
Later among the works it cites.
James ONeill, Polina Rozenshtein, Ryuichi Kiryo, Motoko Kubota, and Danushka Bollegala. 2021 · 2021
Later among the works it cites.
Does putting a linguist in the loop improve NLU data collection?
Alicia Parrish, William Huang, Omar Agha, Soo-Hwan Lee, Nikita Nangia, Alexia Warstadt, Karmanya Aggarwal, Emily Allaway, Tal Linzen, and Samuel R. Bowman. 2021 · 2021
Later among the works it cites.
Get your vitamin C! robust fact verification with contrastive evidence
Tal Schuster, Adam Fisch, and Regina Barzilay. 2021 · 2021
Later among the works it cites.
A puzzle-based dataset for natural language inference
Roxana Szomiu and Adrian Groza. 2021 · 2021
Later among the works it cites.
Trusting roberta over bert: Insights from checklisting the natural language inference task
Ishan Tarunesh, Somak Aditya, and Monojit Choudhury. 2021 · 2021
Later among the works it cites.
Probing language models for understanding of temporal expressions
Shivin Thukral, Kunal Kukreja, and Christian Kavouras. 2021 · 2021
Later among the works it cites.
Learning from the worst: Dynamically generated datasets to improve online hate detection
Bertie Vidgen, Tristan Thrush, Zeerak Waseem, and Douwe Kiela. 2021 · 2021
Later among the works it cites.
Semi-Supervised Exaggeration Detection of Health Science Press Releases
Dustin Wright and Isabelle Augenstein. 2021 · 2021
Later among the works it cites.
Exploring transitivity in neural NLI models through veridicality
Hitomi Yanaka, Koji Mineshima, and Kentaro Inui. 2021 · 2021
Later among the works it cites.
CrossFit: A few-shot learning challenge for cross-task generalization in NLP
Qinyuan Ye, Bill Yuchen Lin, and Xiang Ren. 2021 · 2021
Later among the works it cites.
When does pretraining help? assessing self-supervised learning for law and the casehold dataset
Lucia Zheng, Neel Guha, Brandon R. Anderson, Peter Henderson, and Daniel E. Ho. 2021 · 2021
Later among the works it cites.
Avicenna: a challenge dataset for natural language generation toward commonsense syllogistic reasoning
Zeinab Aghahadi and Alireza Talebpour. 2022 · 2022
Later among the works it cites.
Ext5: Towards extreme multi-task scaling for transfer learning
Vamsi Aribandi, Yi Tay, Tal Schuster, Jinfeng Rao, Huaixiu Steven Zheng, Sanket Vaibhav Mehta, Honglei Zhuang, Vinh Q. Tran, Dara Bahri, Jianmo Ni, Jai Gupta, Kai Hui, Sebastian Ruder, and Donald Metzler. 2022 · 2022
Later among the works it cites.
Promptsource: An integrated development environment and repository for natural language prompts
Stephen H. Bach, Victor Sanh, Zheng-Xin Yong, Albert Webson, Colin Raffel, Nihal V. Nayak, Abheesht Sharma, Taewoon Kim, M Saiful Bari, Thibault Fevry, Zaid Alyafeai, Manan Dey, Andrea Santilli, Zhiqing Sun, Srulik Ben-David, Canwen Xu, Gunjan Chhablani, Han Wang, Jason Alan Fries, Maged S. Al-shaibani, Shanya Sharma, Urmish Thakker, Khalid Almubarak, Xiangru Tang, Xiangru Tang, Mike Tian-Jian Jiang, and Alexander M. Rush. 2022 · 2022
Later among the works it cites.
Multi-step deductive reasoning over natural language: An empirical study on out-of-distribution generalisation
Qiming Bao, Alex Yuxuan Peng, Tim Hartill, Neset Tan, Zhenyun Deng, Michael Witbrock, and Jiamou Liu. 2022 · 2022
Later among the works it cites.
LexGLUE: A benchmark dataset for legal language understanding in English
Ilias Chalkidis, Abhik Jana, Dirk Hartung, Michael Bommarito, Ion Androutsopoulos, Daniel Katz, and Nikolaos Aletras. 2022 · 2022
Later among the works it cites.
Where to start? analyzing the potential value of intermediate models
Leshem Choshen, Elad Venezian, Shachar Don-Yehia, Noam Slonim, and Yoav Katz. 2022 · 2022
Later among the works it cites.
Bigbio: A framework for data-centric biomedical natural language processing
Jason Alan Fries, Leon Weber, Natasha Seelam, Gabriel Altay, Debajyoti Datta, Samuele Garda, Myungsun Kang, Ruisi Su, Wojciech Kusa, Samuel Cahyawijaya, Fabio Barth, Simon Ott, Matthias Samwald, Stephen Bach, Stella Biderman, Mario Sänger, Bo Wang, Alison Callahan, Daniel León Periñán, Théo Gigant, Patrick Haller, Jenny Chim, Jose David Posada, John Michael Giorgi, Karthik Rangasai Sivaraman, Marc Pàmies, Marianna Nezhurina, Robert Martin, Michael Cullan, Moritz Freidank, Nathan Dahlberg, Shubhanshu Mishra, Shamik Bose, Nicholas Michio Broad, Yanis Labrak, Shlok S Deshmukh, Sid Kiblawi, Ayush Singh, Minh Chien Vu, Trishala Neeraj, Jonas Golde, Albert Villanova del Moral, and Benjamin Beilharz. 2022 · 2022
Later among the works it cites.
Cicero: A dataset for contextualized commonsense inference in dialogues
Deepanway Ghosal, Siqi Shen, Navonil Majumder, Rada Mihalcea, and Soujanya Poria. 2022 · 2022
Later among the works it cites.
Folio: Natural language reasoning with first-order logic
Simeng Han, Hailey Schoelkopf, Yilun Zhao, Zhenting Qi, Martin Riddell, Luke Benson, Lucy Sun, Ekaterina Zubova, Yujie Qiao, Matthew Burtell, David Peng, Jonathan Fan, Yixin Liu, Brian Wong, Malcolm Sailor, Ansong Ni, Linyong Nan, Jungo Kasai, Tao Yu, Rui Zhang, Shafiq Joty, Alexander R. Fabbri, Wojciech Kryscinski, Xi Victoria Lin, Caiming Xiong, and Dragomir Radev. 2022 · 2022
Later among the works it cites.
Less annotating, more classifying–addressing the data scarcity issue of supervised machine learning with deep transfer learning and bert-nli
Moritz Laurer, W v Atteveldt, Andreu Casas, and Kasper Welbers. 2022 · 2022
Later among the works it cites.
Wanli: Worker and ai collaboration for natural language inference dataset creation
Alisa Liu, Swabha Swayamdipta, Noah A. Smith, and Yejin Choi. 2022 · 2022
Later among the works it cites.
Cross-task generalization via natural language crowdsourcing instructions
Swaroop Mishra, Daniel Khashabi, Chitta Baral, and Hannaneh Hajishirzi. 2022 · 2022
Later among the works it cites.
Medmcqa: A large-scale multi-subject multi-choice dataset for medical domain question answering
Ankit Pal, Logesh Kumar Umapathi, and Malaikannan Sankarasubbu. 2022 · 2022
Later among the works it cites.
Discovering language model behaviors with model-written evaluations
Ethan Perez, Sam Ringer, Kamil Lukoit, Karina Nguyen, Edwin Chen, Scott Heiner, Craig Pettit, Catherine Olsson, Sandipan Kundu, Saurav Kadavath, Andy Jones, Anna Chen, Ben Mann, Brian Israel, Bryan Seethor, Cameron McKinnon, Christopher Olah, Da Yan, Daniela Amodei, Dario Amodei, Dawn Drain, Dustin Li, Eli Tran-Johnson, Guro Khundadze, Jackson Kernion, James Landis, Jamie Kerr, Jared Mueller, Jeeyoon Hyun, Joshua Landau, Kamal Ndousse, Landon Goldberg, Liane Lovitt, Martin Lucas, Michael Sellitto, Miranda Zhang, Neerav Kingsland, Nelson Elhage, Nicholas Joseph, Noem Mercado, Nova DasSarma, Oliver Rausch, Robin Larson, Sam McCandlish, Scott Johnston, Shauna Kravec, Sheer E, Tamera Lanham, Timothy Telleen-Lawton, Tom Brown, Tom Henighan, Tristan Hume, Yuntao Bai, Zac Hatfield-Dodds, Jack Clark, Samuel R. Bowman, Amanda Askell, Roger Grosse, Danny Hernandez, Deep Ganguli, Evan Hubinger, Nicholas Schiefer, and Jared Kaplan. 2022 · 2022
Later among the works it cites.
Pic: A phrase-in-context dataset for phrase understanding and semantic search
Thang M Pham, Seunghyun Yoon, Trung Bui, and Anh Nguyen. 2022 · 2022
Later among the works it cites.
Condaqa: A contrastive reading comprehension dataset for reasoning about negation
Abhilasha Ravichander, Matt Gardner, and Ana Marasovic. 2022 · 2022
Later among the works it cites.
SciNLI: A corpus for natural language inference on scientific text
Mobashir Sadat and Cornelia Caragea. 2022 · 2022
Later among the works it cites.
Scienceqa: A novel resource for question answering on scholarly articles
Tanik Saikh, Tirthankar Ghosal, Amish Mittal, Asif Ekbal, and Pushpak Bhattacharyya. 2022 · 2022
Later among the works it cites.
Can transformers reason in fragments of natural language?
Viktor Schlegel, Kamen V. Pavlov, and Ian Pratt-Hartmann. 2022 · 2022
Later among the works it cites.
Efficient few-shot learning without prompts
Lewis Tunstall, Nils Reimers, Unso Eun Seo Jo, Luke Bates, Daniel Korat, Moshe Wasserblat, and Oren Pereg. 2022 · 2022
Later among the works it cites.
Analyzing dynamic adversarial training data in the limit
Eric Wallace, Adina Williams, Robin Jia, and Douwe Kiela. 2022 · 2022
Later among the works it cites.
Super-naturalinstructions:generalization via declarative instructions on 1600+ tasks
Yizhong Wang, Swaroop Mishra, Pegah Alipoormolabashi, Yeganeh Kordi, Amirreza Mirzaei, Anjana Arunkumar, Arjun Ashok, Arut Selvan Dhanasekaran, Atharva Naik, David Stap, et al. 2022 · 2022
Later among the works it cites.
Distributed nli: Learning to predict human opinion distributions for language reasoning
Xiang Zhou, Yixin Nie, and Mohit Bansal. 2022 · 2022
Later among the works it cites.
Stanford human preferences dataset
Kawin Ethayarajh, Heidi Zhang, Yizhong Wang, and Dan Jurafsky. 2023 · 2023
Closest in time.
We’re afraid language models aren’t modeling ambiguity
Alisa Liu, Zhaofeng Wu, Julian Michael, Alane Suhr, Peter West, Alexander Koller, Swabha Swayamdipta, Noah A. Smith, and Yejin Choi. 2023 · 2023
Closest in time.
tasknet, multitask interface between Trainer and datasets
Damien Sileo. 2023 · 2023
Closest in time.
Mindgames: Targeting theory of mind in large language models with dynamic epistemic modal logic
Damien Sileo and Antoine Lernould. 2023 · 2023
Closest in time.
Generating multiple-choice questions for medical question answering with distractors and cue-masking
Damien Sileo, Kanimozhi Uma, and Marie-Francine Moens. 2023 · 2023
Closest in time.