Fetching the paper…
Reading the bibliography…
We introduce Dynabench, an open-source platform for dynamic dataset creation and model benchmarking.
RoBERTa: A robustly optimized BERT pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019b · 1907
Earlier work this paper cites.
CrowS-pairs: A challenge dataset for measuring social biases in masked language models
Nikita Nangia, Clara Vania, Rasika Bhalerao, and Samuel R. Bowman. 2020 · 1967
Earlier work this paper cites.
Studies on natural logic and categorial grammar, University of Amsterdam Ph. D
Victor Sánchez Valencia. 1991 · 1991
Earlier work this paper cites.
Building a large annotated corpus of English: The Penn Treebank
Mitchell P. Marcus, Beatrice Santorini, and Mary Ann Marcinkiewicz. 1993 · 1993
Earlier work this paper cites.
A model of textual affect sensing using real-world knowledge
Hugo Liu, Henry Lieberman, and Ted Selker. 2003 · 2003
Earlier work this paper cites.
Directions in Abusive Language Training Data: Garbage In, Garbage Out
Bertie Vidgen and Leon Derczynski. 2020 · 2004
Earlier work this paper cites.
Emotions from text: Machine learning for text-based emotion prediction
Cecilia Ovesdotter Alm, Dan Roth, and Richard Sproat. 2005 · 2005
Earlier work this paper cites.
Language models are few-shot learners
Tom B Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020 · 2005
Earlier work this paper cites.
Annotating expressions of opinions and emotions in language
Janyce Wiebe, Theresa Wilson, and Claire Cardie. 2005 · 2005
Earlier work this paper cites.
The PASCAL recognising textual entailment challenge
Ido Dagan, Oren Glickman, and Bernardo Magnini. 2006 · 2006
Earlier work this paper cites.
Towards ecologically valid research on language user interfaces
Harm de Vries, Dzmitry Bahdanau, and Christopher Manning. 2020 · 2007
Earlier work this paper cites.
Question and answer test-train overlap in open-domain question answering datasets
Patrick Lewis, Pontus Stenetorp, and Sebastian Riedel. 2020 · 2008
Earlier work this paper cites.
Opinion mining and sentiment analysis
Bo Pang and Lillian Lee. 2008 · 2008
Earlier work this paper cites.
Cheap and fast – but is it good? evaluating non-expert annotations for natural language tasks
Rion Snow, Brendan O’Connor, Daniel Jurafsky, and Andrew Ng. 2008 · 2008
Earlier work this paper cites.
A brief history of natural logic
Johan van Benthem. 2008 · 2008
Earlier work this paper cites.
Designing games with a purpose
Luis Von Ahn and Laura Dabbish. 2008 · 2008
Earlier work this paper cites.
William Huang, Haokun Liu, and Samuel R Bowman. 2020 · 2010
Earlier work this paper cites.
Crowdsourcing and language studies: the new generation of linguistic data
Robert Munro, Steven Bethard, Victor Kuperman, Vicky Tzuyin Lai, Robin Melnick, Christopher Potts, Tyler Schnoebelen, and Harry Tily. 2010 · 2010
Earlier work this paper cites.
Recognition of affect, judgment, and appreciation in text
Alena Neviarouskaya, Helmut Prendinger, and Mitsuru Ishizuka. 2010 · 2010
Earlier work this paper cites.
DynaSent: A dynamic benchmark for sentiment analysis
Christopher Potts, Zhengxuan Wu, Atticus Geiger, and Douwe Kiela. 2020 · 2012
Earlier work this paper cites.
CoNLL-2012 shared task: Modeling multilingual unrestricted coreference in OntoNotes
Sameer Pradhan, Alessandro Moschitti, Nianwen Xue, Olga Uryupina, and Yuchen Zhang. 2012 · 2012
Earlier work this paper cites.
Learning from the worst: Dynamically generated datasets to improve online hate detection
Bertie Vidgen, Tristan Thrush, Zeerak Waseem, and Douwe Kiela. 2020 · 2012
Earlier work this paper cites.
Lifelong machine learning systems: Beyond learning algorithms
Daniel L Silver, Qiang Yang, and Lianghao Li. 2013 · 2013
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D. Manning, Andrew Ng, and Christopher Potts. 2013 · 2013
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. 2014 · 2014
Earlier work this paper cites.
Sentiment expression conditioned by affective transitions and social forces
Moritz Sudhof, Andrés Gómez Emilsson, Andrew L. Maas, and Christopher Potts. 2014 · 2014
Earlier work this paper cites.
Beat the machine: Challenging humans to find a predictive model’s “unknown unknowns”
Joshua Attenberg, Panos Ipeirotis, and Foster Provost. 2015 · 2015
Earlier work this paper cites.
A large annotated corpus for learning natural language inference
Samuel R. Bowman, Gabor Angeli, Christopher Potts, and Christopher D. Manning. 2015 · 2015
Earlier work this paper cites.
Effective approaches to attention-based neural machine translation
Thang Luong, Hieu Pham, and Christopher D. Manning. 2015 · 2015
Earlier work this paper cites.
Dangerous Speech and Dangerous Ideology: An Integrated Model for Monitoring and Prevention
Jonathan Leader Maynard and Susan Benesch. 2016 · 2016
Earlier work this paper cites.
SQuAD: 100,000+ questions for machine comprehension of text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016 · 2016
Earlier work this paper cites.
Towards linguistically generalizable NLP systems: A workshop and shared task
Allyson Ettinger, Sudha Rao, Hal Daumé III, and Emily M. Bender. 2017 · 2017
Earlier work this paper cites.
Bidirectional attention flow for machine comprehension
Minjoon Seo, Aniruddha Kembhavi, Ali Farhadi, and Hannaneh Hajishirzi. 2017 · 2017
Earlier work this paper cites.
Axiomatic attribution for deep networks
Mukund Sundararajan, Ankur Taly, and Qiqi Yan. 2017 · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
Understanding Abuse: A Typology of Abusive Language Detection Subtasks
Zeerak Waseem, Thomas Davidson, Dana Warmsley, and Ingmar Weber. 2017 · 2017
Cited alongside, same era.
Mastering the dungeon: Grounded language learning by mechanical turker descent
Zhilin Yang, Saizheng Zhang, Jack Urbanek, Will Feng, Alexander H Miller, Arthur Szlam, Douwe Kiela, and Jason Weston. 2017 · 2017
Cited alongside, same era.
Synthetic and natural noise both break neural machine translation
Yonatan Belinkov and Yonatan Bisk. 2018 · 2018
Cited alongside, same era.
SentEval: An evaluation toolkit for universal sentence representations
Alexis Conneau and Douwe Kiela. 2018 · 2018
Cited alongside, same era.
Pytorch: An imperative style, high-performance deep learning library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Kopf, Edward Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala. 2019 · 2019
Later among the works it cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019 · 2019
Later among the works it cites.
Diversify your datasets: Analyzing generalization via controlled variance in adversarial datasets
Ohad Rozen, Vered Shwartz, Roee Aharoni, and Ido Dagan. 2019 · 2019
Later among the works it cites.
Do neural dialog systems use the conversation history effectively? an empirical study
Chinnadhurai Sankar, Sandeep Subramanian, Chris Pal, Sarath Chandar, and Yoshua Bengio. 2019 · 2019
Later among the works it cites.
Challenges and frontiers in abusive content detection
Bertie Vidgen, Alex Harris, Dong Nguyen, Rebekah Tromble, Scott Hale, and Helen Margetts. 2019 · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Atticus Geiger, Ignacio Cases, Lauri Karttunen, and Christopher Potts. 2018 · 2018
Cited alongside, same era.
Breaking NLI systems with sentences that require simple lexical inferences
Max Glockner, Vered Shwartz, and Yoav Goldberg. 2018 · 2018
Cited alongside, same era.
Annotation artifacts in natural language inference data
Suchin Gururangan, Swabha Swayamdipta, Omer Levy, Roy Schwartz, Samuel Bowman, and Noah A. Smith. 2018 · 2018
Cited alongside, same era.
Universal language model fine-tuning for text classification
Jeremy Howard and Sebastian Ruder. 2018 · 2018
Cited alongside, same era.
On the impact of various types of noise on neural machine translation
Huda Khayrallah and Philipp Koehn. 2018 · 2018
Cited alongside, same era.
The natural language decathlon: Multitask learning as question answering
Bryan McCann, Nitish Shirish Keskar, Caiming Xiong, and Richard Socher. 2018 · 2018
Cited alongside, same era.
Adversarially regularising neural NLI models to integrate logical background knowledge
Pasquale Minervini and Sebastian Riedel. 2018 · 2018
Cited alongside, same era.
Later among the works it cites.
Universal adversarial triggers for attacking and analyzing NLP
Eric Wallace, Shi Feng, Nikhil Kandpal, Matt Gardner, and Sameer Singh. 2019 · 2019
Later among the works it cites.
Superglue: A stickier benchmark for general-purpose language understanding systems
Alex Wang, Yada Pruksachatkun, Nikita Nangia, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel Bowman. 2019 · 2019
Later among the works it cites.
Investigating BERT’s knowledge of language: Five analysis methods with NPIs
Alex Warstadt, Yu Cao, Ioana Grosu, Wei Peng, Hagen Blix, Yining Nie, Anna Alsop, Shikha Bordia, Haokun Liu, Alicia Parrish, Sheng-Fu Wang, Jason Phang, Anhad Mohananey, Phu Mon Htut, Paloma Jeretic, and Samuel R. Bowman. 2019 · 2019
Later among the works it cites.
Beat the ai: Investigating adversarial human annotation for reading comprehension
Max Bartolo, Alastair Roberts, Johannes Welbl, Sebastian Riedel, and Pontus Stenetorp. 2020 · 2020
Later among the works it cites.
With little power comes great responsibility
Dallas Card, Peter Henderson, Urvashi Khandelwal, Robin Jia, Kyle Mahowald, and Dan Jurafsky. 2020 · 2020
Later among the works it cites.
Utility is in the eye of the user: A critique of NLP leaderboard design
Kawin Ethayarajh and Dan Jurafsky. 2020 · 2020
Later among the works it cites.
What BERT is not: Lessons from a new suite of psycholinguistic diagnostics for language models
Allyson Ettinger. 2020 · 2020
Later among the works it cites.
Evaluating models’ local decision boundaries via contrast sets
Matt Gardner, Yoav Artzi, Victoria Basmov, Jonathan Berant, Ben Bogin, Sihao Chen, Pradeep Dasigi, Dheeru Dua, Yanai Elazar, Ananth Gottumukkala, Nitish Gupta, Hannaneh Hajishirzi, Gabriel Ilharco, Daniel Khashabi, Kevin Lin, Jiangming Liu, Nelson F. Liu, Phoebe Mulcaire, Qiang Ning, Sameer Singh, Noah A. Smith, Sanjay Subramanian, Reut Tsarfaty, Eric Wallace, Ally Zhang, and Ben Zhou. 2020 · 2020
Later among the works it cites.
SyntaxGym: An online platform for targeted evaluation of language models
Jon Gauthier, Jennifer Hu, Ethan Wilcox, Peng Qian, and Roger Levy. 2020 · 2020
Later among the works it cites.
An analysis of natural language inference benchmarks through the lens of negation
Md Mosharaf Hossain, Venelin Kovatchev, Pranoy Dutta, Tiffany Kao, Elizabeth Wei, and Eduardo Blanco. 2020 · 2020
Later among the works it cites.
Are natural language inference models IMPPRESsive? Learning IMPlicature and PRESupposition
Paloma Jeretic, Alex Warstadt, Suvrat Bhooshan, and Adina Williams. 2020 · 2020
Later among the works it cites.
Learning the difference that makes a difference with counterfactually-augmented data
Divyansh Kaushik, Eduard Hovy, and Zachary C Lipton. 2020 · 2020
Later among the works it cites.
COGS: A compositional generalization challenge based on semantic interpretation
Najoung Kim and Tal Linzen. 2020 · 2020
Later among the works it cites.
How can we accelerate progress towards human-like linguistic generalization?
Tal Linzen. 2020 · 2020
Later among the works it cites.
Adversarial NLI: A new benchmark for natural language understanding
Yixin Nie, Adina Williams, Emily Dinan, Mohit Bansal, Jason Weston, and Douwe Kiela. 2020 · 2020
Later among the works it cites.
Resources and benchmark corpora for hate speech detection: a systematic review
Fabio Poletto, Valerio Basile, Manuela Sanguinetti, Cristina Bosco, and Viviana Patti. 2020 · 2020
Later among the works it cites.
Beyond accuracy: Behavioral testing of NLP models with CheckList
Marco Tulio Ribeiro, Tongshuang Wu, Carlos Guestrin, and Sameer Singh. 2020 · 2020
Later among the works it cites.
Hatecheck: Functional tests for hate speech detection models
Paul Röttger, Bertram Vidgen, Dong Nguyen, Zeerak Waseem, Helen Margetts, and Janet Pierrehumbert. 2020 · 2020
Later among the works it cites.
ConjNLI: Natural language inference over conjunctive sentences
Swarnadeep Saha, Yixin Nie, and Mohit Bansal. 2020 · 2020
Later among the works it cites.
Harnessing the linguistic signal to predict scalar inferences
Sebastian Schuster, Yuxing Chen, and Judith Degen. 2020 · 2020
Later among the works it cites.
The dialogue dodecathlon: Open-domain knowledge and image grounded conversational agents
Kurt Shuster, Da Ju, Stephen Roller, Emily Dinan, Y-Lan Boureau, and Jason Weston. 2020 · 2020
Later among the works it cites.
Learning to summarize with human feedback
Nisan Stiennon, Long Ouyang, Jeffrey Wu, Daniel Ziegler, Ryan Lowe, Chelsea Voss, Alec Radford, Dario Amodei, and Paul F Christiano. 2020 · 2020
Later among the works it cites.
Assessing the benchmarking capacity of machine reading comprehension datasets
Saku Sugawara, Pontus Stenetorp, Kentaro Inui, and Akiko Aizawa. 2020 · 2020
Later among the works it cites.
Dataset cartography: Mapping and diagnosing datasets with training dynamics
Swabha Swayamdipta, Roy Schwartz, Nicholas Lourie, Yizhong Wang, Hannaneh Hajishirzi, Noah A. Smith, and Yejin Choi. 2020 · 2020
Later among the works it cites.
BLiMP: A benchmark of linguistic minimal pairs for English
Alex Warstadt, Alicia Parrish, Haokun Liu, Anhad Mohananey, Wei Peng, Sheng-Fu Wang, and Samuel R. Bowman. 2020 · 2020
Later among the works it cites.
The universal decompositional semantics dataset and decomp toolkit
Aaron Steven White, Elias Stengel-Eskin, Siddharth Vashishtha, Venkata Subrahmanyan Govindarajan, Dee Ann Reisinger, Tim Vieira, Keisuke Sakaguchi, Sheng Zhang, Francis Ferraro, Rachel Rudinger, Kyle Rawlins, and Benjamin Van Durme. 2020 · 2020
Later among the works it cites.
Assessing phrasal representation and composition in transformers
Lang Yu and Allyson Ettinger. 2020 · 2020
Later among the works it cites.
The curse of performance instability in analysis datasets: Consequences, source, and suggestions
Xiang Zhou, Yixin Nie, Hao Tan, and Mohit Bansal. 2020 · 2020
Later among the works it cites.
Adversarial examples for evaluating reading comprehension systems
Robin Jia and Percy Liang. 2017 · 2031
Closest in time.