Fetching the paper…
Reading the bibliography…
Crowdsourcing has been the prevalent paradigm for creating natural language understanding datasets in recent years.
Note on the sampling error of the difference between correlated proportions or percentages
Quinn McNemar. 1947 · 1947
Earlier work this paper cites.
Crowdsourcing translation: Professional quality from non-professionals
Omar F. Zaidan and Chris Callison-Burch. 2011 · 2011
Earlier work this paper cites.
An empirical investigation of statistical significance in nlp
Taylor Berg-Kirkpatrick, David Burkett, and Dan Klein. 2012 · 2012
Earlier work this paper cites.
MCTest: A challenge dataset for the open-domain machine comprehension of text
M. Richardson, C. J. Burges, and E. Renshaw. 2013 · 2013
Earlier work this paper cites.
Corpus annotation through crowdsourcing: Towards best practice guidelines
Marta Sabou, Kalina Bontcheva, Leon Derczynski, and Arno Scharl. 2014 · 2014
Earlier work this paper cites.
A large annotated corpus for learning natural language inference
S. Bowman, G. Angeli, C. Potts, and C. D. Manning. 2015 · 2015
Earlier work this paper cites.
Crowdsourcing for NLP
Chris Callison-Burch, Lyle Ungar, and Ellie Pavlick. 2015 · 2015
Earlier work this paper cites.
Do supervised distributional methods really learn lexical inference relations?
Omer Levy, Steffen Remus, Chris Biemann, and Ido Dagan. 2015 · 2015
Earlier work this paper cites.
SQuAD: 100,000+ questions for machine comprehension of text
P. Rajpurkar, J. Zhang, K. Lopyrev, and P. Liang. 2016 · 2016
Earlier work this paper cites.
Visual Genome: Connecting language and vision using crowdsourced dense image annotations
Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, Yannis Kalantidis, Li-Jia Li, David A Shamma, et al. 2017 · 2017
Earlier work this paper cites.
The effect of different writing tasks on linguistic style: A case study of the ROC story cloze task
R. Schwartz, M. Sap, Y. Konstas, L. Zilles, Y. Choi, and N. A. Smith. 2017 · 2017
Cited alongside, same era.
Conceptnet 5.5: An open multilingual graph of general knowledge
Robyn Speer, Joshua Chin, and Catherine Havasi. 2017 · 2017
Cited alongside, same era.
NewsQA: A machine comprehension dataset
A. Trischler, T. Wang, X. Yuan, J. Harris, A. Sordoni, P. Bachman, and K. Suleman. 2017 · 2017
Cited alongside, same era.
Split and rephrase: Better evaluation and stronger baselines
Roee Aharoni and Yoav Goldberg. 2018 · 2018
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018 · 2018
Cited alongside, same era.
Comparing bayesian models of annotation
Silviu Paun, Bob Carpenter, Jon Chamberlain, Dirk Hovy, Udo Kruschwitz, and Massimo Poesio. 2018 · 2018
Later among the works it cites.
Hypothesis only baselines in Natural Language Inference
A. Poliak, J. Naradowsky, A. Haldar, R. Rudinger, and B. V. Durme. 2018 · 2018
Later among the works it cites.
Know what you don’t know: Unanswerable questions for squad
Pranav Rajpurkar, Robin Jia, and Percy Liang. 2018 · 2018
Later among the works it cites.
A corpus for reasoning about natural language grounded in photographs
Alane Suhr, Stephanie Zhou, Iris Zhang, Huajun Bai, and Yoav Artzi. 2018 · 2018
Later among the works it cites.
Performance impact caused by hidden bias of training data for recognizing textual entailment
Masatoshi Tsuchiya. 2018 · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
The hitchhiker’s guide to testing statistical significance in natural language processing
Rotem Dror, Gili Baumer, Segev Shlomov, and Roi Reichart. 2018 · 2018
Cited alongside, same era.
Breaking NLI systems with sentences that require simple lexical inferences
Max Glockner, Vered Shwartz, and Yoav Goldberg. 2018 · 2018
Cited alongside, same era.
Annotation artifacts in natural language inference data
S. Gururangan, S. Swayamdipta, O. Levy, R. Schwartz, S. R. Bowman, and N. A. Smith. 2018 · 2018
Cited alongside, same era.
Can a suit of armor conduct electricity? A new dataset for open book question answering
Todor Mihaylov, Peter Clark, Tushar Khot, and Ashish Sabharwal. 2018 · 2018
Cited alongside, same era.
A broad-coverage challenge corpus for sentence understanding through inference
Adina Williams, Nikita Nangia, and Samuel Bowman. 2018 · 2018
Later among the works it cites.
DROP: A reading comprehension benchmark requiring discrete reasoning over paragraphs
Dheeru Dua, Yizhong Wang, Pradeep Dasigi, Gabriel Stanovsky, Sameer Singh, and Matt Gardner. 2019 · 2019
Closest in time.
CommonsenseQA: A question answering challenge targeting commonsense knowledge
A. Talmor, J. Herzig, N. Lourie, and J. Berant. 2019 · 2019
Closest in time.
From recognition to cognition: Visual commonsense reasoning
Rowan Zellers, Yonatan Bisk, Ali Farhadi, and Yejin Choi. 2019 · 2019
Closest in time.