Fetching the paper…
Reading the bibliography…
It has become a common pattern in our field: One group introduces a language task, exemplified by a dataset, which they argue is challenging enough to serve as a benchmark.
Language tasks and language games: On methodology in current natural language processing research
David Schlangen. 2019 · 1908
Earlier work this paper cites.
On the Measure of Intelligence
François Chollet. 2019 · 1911
Earlier work this paper cites.
Logik der Forschung
Karl Popper. 1934 · 1934
Earlier work this paper cites.
Measuring the Mind: Conceptual Issues in Contemporary Psychometrics
Denny Borsboom. 2005 · 2005
Earlier work this paper cites.
Inter-Coder Agreement for Computational Linguistics
Ron Artstein and Massimo Poesio. 2008 · 2008
Earlier work this paper cites.
Natural Language Annotation for Machine Learning
James Pustejovsky and Amber Stubbs. 2013 · 2013
Earlier work this paper cites.
Vqa: Visual question answering
Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C. Lawrence Zitnick, and Devi Parikh. 2015 · 2015
Cited alongside, same era.
A Dictionary of Computer Science , 7th edition
Andrew Butterfield, Gerard Ekembe Ngondi, and Anne Kerr, editors. 2016 · 2016
Cited alongside, same era.
Revisiting Visual Question Answering Baselines
Allan Jabri, Armand Joulin, and Laurens van der Maaten. 2016 · 2016
Cited alongside, same era.
Making the V in VQA Matter: Elevating the Role of Image Understanding in Visual Question Answering
Yash Goyal, Tejas Khot, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh. 2017 · 2017
Cited alongside, same era.
BERT: pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018 · 2018
Deep contextualized word representations
Matthew Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer. 2018 · 2018
Later among the works it cites.
A Broad-Coverage Challenge Corpus for Sentence Understanding through Inference
Adina Williams, Nikita Nangia, and Samuel R. Bowman. 2018 · 2018
Later among the works it cites.
UCI machine learning repository
Dheeru Dua and Casey Graff. 2019 · 2019
Later among the works it cites.
Pragmatic factors in (automatic) image description
Emiel van Miltenburg. 2019 · 2019
Later among the works it cites.
Inherent Disagreements in Human Textual Inferences
Ellie Pavlick and Tom Kwiatkowski. 2019 · 2019
Later among the works it cites.
GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Dual indicators to analyse ai benchmarks: Difficulty, discrimination, ability and generality
Fernando Martinez-Plumed and José Hernandez-Orallo. 2018 · 2018
Cited alongside, same era.
SuperGLUE: A Stickier Benchmark for General-Purpose Language Understanding Systems
Alex Wang, Yada Pruksachatkun, Nikita Nangia, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R. Bowman. 2019a
Cited in the paper.
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R. Bowman. 2019b · 2019
Later among the works it cites.