Fetching the paper…
Reading the bibliography…
In recent years, progress in NLU has been driven by benchmarks.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 1907
Earlier work this paper cites.
Dqi: Measuring data quality in nlp
Swaroop Mishra, Anjana Arunkumar, Bhavdeep Sachdeva, Chris Bryan, and Chitta Baral. 2020b · 2005
Earlier work this paper cites.
Our evaluation metric needs an update to encourage generalization
Swaroop Mishra, Anjana Arunkumar, Chris Bryan, and Chitta Baral. 2020a · 2007
Earlier work this paper cites.
Convai3: Generating clarifying questions for open-domain dialogue systems (clariq)
Mohammad Aliannejadi, Julia Kiseleva, Aleksandr Chuklin, Jeff Dalton, and Mikhail Burtsev. 2020 · 2009
Earlier work this paper cites.
Creating speech and language data with Amazon’s Mechanical Turk
Chris Callison-Burch and Mark Dredze. 2010 · 2010
Earlier work this paper cites.
The turking test: Can language models understand instructions?
Avia Efrat and Omer Levy. 2020 · 2010
Earlier work this paper cites.
Glance: Rapidly coding behavioral video with the crowd
Walter S Lasecki, Mitchell Gordon, Danai Koutra, Malte F Jung, Steven P Dow, and Jeffrey P Bigham. 2014 · 2014
Earlier work this paper cites.
Deep learning applications and challenges in big data analytics
Maryam M Najafabadi, Flavio Villanustre, Taghi M Khoshgoftaar, Naeem Seliya, Randall Wald, and Edin Muharemagic. 2015 · 2015
Earlier work this paper cites.
SQuAD: 100,000+ questions for machine comprehension of text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016 · 2016
Earlier work this paper cites.
Revolt: Collaborative crowdsourcing for labeling machine learning datasets
Joseph Chee Chang, Saleema Amershi, and Ece Kamar. 2017 · 2017
Earlier work this paper cites.
The effect of different writing tasks on linguistic style: A case study of the roc story cloze task
Roy Schwartz, Maarten Sap, Ioannis Konstas, Leila Zilles, Yejin Choi, and Noah A Smith. 2017 · 2017
Earlier work this paper cites.
Crowdsourcing multiple choice science questions
Johannes Welbl, Nelson F. Liu, and Matt Gardner. 2017 · 2017
Earlier work this paper cites.
Cognitive biases in crowdsourcing
Carsten Eickhoff. 2018 · 2018
Earlier work this paper cites.
Annotation artifacts in natural language inference data
Suchin Gururangan, Swabha Swayamdipta, Omer Levy, Roy Schwartz, Samuel Bowman, and Noah A. Smith. 2018 · 2018
Earlier work this paper cites.
Looking beyond the surface: A challenge set for reading comprehension over multiple sentences
Daniel Khashabi, Snigdha Chaturvedi, Michael Roth, Shyam Upadhyay, and Dan Roth. 2018 · 2018
Earlier work this paper cites.
Hypothesis only baselines in natural language inference
Adam Poliak, Jason Naradowsky, Aparajita Haldar, Rachel Rudinger, and Benjamin Van Durme. 2018 · 2018
Earlier work this paper cites.
Know what you don’t know: Unanswerable questions for SQuAD
Pranav Rajpurkar, Robin Jia, and Percy Liang. 2018 · 2018
Earlier work this paper cites.
DuoRC: Towards complex language understanding with paraphrased reading comprehension
Amrita Saha, Rahul Aralikatte, Mitesh M. Khapra, and Karthik Sankaranarayanan. 2018 · 2018
Earlier work this paper cites.
Performance impact caused by hidden bias of training data for recognizing textual entailment
Masatoshi Tsuchiya. 2018 · 2018
Cited alongside, same era.
Hotpotqa: A dataset for diverse, explainable multi-hop question answering
Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William Cohen, Ruslan Salakhutdinov, and Christopher D Manning. 2018 · 2018
Cited alongside, same era.
Crowdsourcing methods for data collection in geophysics: State of the art, issues, and future directions
Feifei Zheng, Ruoling Tao, Holger R Maier, Linda See, Dragan Savic, Tuqiao Zhang, Qiuwen Chen, Thaine H Assumpção, Pan Yang, Bardia Heidari, et al. 2018 · 2018
Cited alongside, same era.
Quoref: A reading comprehension dataset with questions requiring coreferential reasoning
Pradeep Dasigi, Nelson F. Liu, Ana Marasović, Noah A. Smith, and Matt Gardner. 2019 · 2019
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Cited alongside, same era.
QASC: A dataset for question answering via sentence composition
Tushar Khot, Peter Clark, Michal Guerquin, Peter Jansen, and Ashish Sabharwal. 2020 · 2020
Later among the works it cites.
Adversarial filters of dataset biases
Ronan Le Bras, Swabha Swayamdipta, Chandra Bhagavatula, Rowan Zellers, Matthew E Peters, Ashish Sabharwal, and Yejin Choi. 2020 · 2020
Later among the works it cites.
BART: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension
Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer. 2020 · 2020
Later among the works it cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020 · 2020
Later among the works it cites.
Winogrande: An adversarial winograd schema challenge at scale
Keisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi. 2020 · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
DROP: A reading comprehension benchmark requiring discrete reasoning over paragraphs
Dheeru Dua, Yizhong Wang, Pradeep Dasigi, Gabriel Stanovsky, Sameer Singh, and Matt Gardner. 2019 · 2019
Cited alongside, same era.
Are we modeling the task or the annotator? an investigation of annotator bias in natural language understanding datasets
Mor Geva, Yoav Goldberg, and Jonathan Berant. 2019 · 2019
Cited alongside, same era.
Cosmos QA: Machine reading comprehension with contextual commonsense reasoning
Lifu Huang, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi. 2019 · 2019
Cited alongside, same era.
Reasoning over paragraph effects in situations
Kevin Lin, Oyvind Tafjord, Peter Clark, and Matt Gardner. 2019 · 2019
Cited alongside, same era.
NumNet: Machine reading comprehension with numerical reasoning
Qiu Ran, Yankai Lin, Peng Li, Jie Zhou, and Zhiyuan Liu. 2019 · 2019
Cited alongside, same era.
“going on a vacation” takes longer than “going for a walk”: A study of temporal commonsense understanding
Ben Zhou, Daniel Khashabi, Qiang Ning, and Dan Roth. 2019 · 2019
Cited alongside, same era.
Real-time visual feedback for educative benchmark creation: A human-and-metric-in-the-loop workflow
Anjana Arunkumar, Swaroop Mishra, Bhavdeep Sachdeva, Chitta Baral, and Chris Bryan. 2020 · 2020
Cited alongside, same era.
Dataset cartography: Mapping and diagnosing datasets with training dynamics
Swabha Swayamdipta, Roy Schwartz, Nicholas Lourie, Yizhong Wang, Hannaneh Hajishirzi, Noah A Smith, and Yejin Choi. 2020 · 2020
Later among the works it cites.
Did aristotle use a laptop? a question answering benchmark with implicit reasoning strategies
Mor Geva, Daniel Khashabi, Elad Segal, Tushar Khot, Dan Roth, and Jonathan Berant. 2021 · 2021
Later among the works it cites.
Investigating and mitigating biases in crowdsourced data
Danula Hettiachchi, Mark Sanderson, Jorge Goncalves, Simo Hosio, Gabriella Kazai, Matthew Lease, Mike Schaekermann, and Emine Yilmaz. 2021 · 2021
Later among the works it cites.
Variational information bottleneck for effective low-resource fine-tuning
Rabeeh Karimi Mahabadi, Yonatan Belinkov, and James Henderson. 2021 · 2021
Later among the works it cites.
How robust are model rankings: A leaderboard customization approach for equitable evaluation
Swaroop Mishra and Anjana Arunkumar. 2021 · 2021
Later among the works it cites.
Cross-task generalization via natural language crowdsourcing instructions
Swaroop Mishra, Daniel Khashabi, Chitta Baral, and Hannaneh Hajishirzi. 2021 · 2021
Later among the works it cites.
What ingredients make for an effective crowdsourcing protocol for difficult NLU data collection tasks?
Nikita Nangia, Saku Sugawara, Harsh Trivedi, Alex Warstadt, Clara Vania, and Samuel R. Bowman. 2021 · 2021
Later among the works it cites.
Demographic biases of crowd workers in key opinion leaders finding
Hossein A Rahmani and Jie Yang. 2021 · 2021
Later among the works it cites.
Qa dataset explosion: A taxonomy of nlp resources for question answering and reading comprehension
Anna Rogers, Matt Gardner, and Isabelle Augenstein. 2021 · 2021
Later among the works it cites.
Crowdsourcing beyond annotation: Case studies in benchmark data collection
Alane Suhr, Clara Vania, Nikita Nangia, Maarten Sap, Mark Yatskar, Samuel R. Bowman, and Yoav Artzi. 2021 · 2021
Later among the works it cites.
Promptsource: An integrated development environment and repository for natural language prompts
Stephen H Bach, Victor Sanh, Zheng-Xin Yong, Albert Webson, Colin Raffel, Nihal V Nayak, Abheesht Sharma, Taewoon Kim, M Saiful Bari, Thibault Fevry, et al. 2022 · 2022
Closest in time.
Benchmarking generalization via in-context instructions on 1,600+ language tasks
Yizhong Wang, Swaroop Mishra, Pegah Alipoormolabashi, Yeganeh Kordi, Amirreza Mirzaei, Anjana Arunkumar, Arjun Ashok, Arut Selvan Dhanasekaran, Atharva Naik, David Stap, et al. 2022 · 2022
Closest in time.