Fetching the paper…
Reading the bibliography…
Alongside huge volumes of research on deep learning models in NLP in the recent years, there has been also much work on benchmark datasets needed to track modeling progress.
Assessing BERT’s Syntactic Abilities
Yoav Goldberg. 2019 · 1901
Earlier work this paper cites.
Quizbowl: The Case for Incremental Question Answering
Pedro Rodriguez, Shi Feng, Mohit Iyyer, He He, and Jordan Boyd-Graber. 2021 · 1904
Earlier work this paper cites.
A Survey of Code-Switched Speech and Language Processing
Sunayana Sitaram, Khyathi Raghavi Chandu, Sai Krishna Rallabandi, and Alan W. Black. 2020 · 1904
Earlier work this paper cites.
ANTIQUE: A Non-Factoid Question Answering Benchmark
Helia Hashemi, Mohammad Aliannejadi, Hamed Zamani, and W. Bruce Croft. 2019 · 1905
Earlier work this paper cites.
HellaSwag: Can a Machine Really Finish Your Sentence?. In ACL 2019
Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi. 2019 · 1905
Earlier work this paper cites.
A Survey on Neural Machine Reading Comprehension
Boyu Qiu, Xu Chen, Jungang Xu, and Yingfei Sun. 2019 · 1906
Earlier work this paper cites.
HEAD-QA: A Healthcare Dataset for Complex Reasoning
David Vilares and Carlos Gómez-Rodríguez. 2019 · 1906
Earlier work this paper cites.
What BERT Is Not: Lessons from a New Suite of Psycholinguistic Diagnostics for Language Models
Allyson Ettinger. 2020 · 1907
Earlier work this paper cites.
RoBERTa: A Robustly Optimized BERT Pretraining Approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019c · 1907
Earlier work this paper cites.
WINOGRANDE: An Adversarial Winograd Schema Challenge at Scale
Keisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi. 2019 · 1907
Earlier work this paper cites.
Universal Adversarial Triggers for Attacking and Analyzing NLP
Eric Wallace, Shi Feng, Nikhil Kandpal, Matt Gardner, and Sameer Singh. 2019 · 1908
Earlier work this paper cites.
Question Answering Is a Format; When Is It Useful?
Matt Gardner, Jonathan Berant, Hannaneh Hajishirzi, Alon Talmor, and Sewon Min. 2019 · 1909
Earlier work this paper cites.
Divyansh Kaushik, Eduard Hovy, and Zachary Lipton. 2019 · 1909
Earlier work this paper cites.
KorQuAD1.0: Korean QA Dataset for Machine Reading Comprehension
Seungyoung Lim, Myungji Kim, and Jooyoul Lee. 2019 · 1909
Earlier work this paper cites.
Learning to Deceive with Attention-Based Explanations
Danish Pruthi, Mansi Gupta, Bhuwan Dhingra, Graham Neubig, and Zachary C. Lipton. 2019 · 1909
Earlier work this paper cites.
Mikel Artetxe, Sebastian Ruder, and Dani Yogatama. 2020 · 1910
Earlier work this paper cites.
What Question Answering Can Learn from Trivia Nerds
Jordan Boyd-Graber. 2019 · 1910
Earlier work this paper cites.
Tom McCoy, Junghyun Min, and Tal Linzen. 2019a · 1911
Earlier work this paper cites.
Assessing the Benchmarking Capacity of Machine Reading Comprehension Datasets. In Proc. of AAAI
Saku Sugawara, Pontus Stenetorp, Kentaro Inui, and Akiko Aizawa. 2020 · 1911
Earlier work this paper cites.
Automatic Spanish Translation of the SQuAD Dataset for Multilingual Question Answering
Casimiro Pio Carrino, Marta R. Costa-jussà, and José A. R. Fonollosa. 2019 · 1912
Earlier work this paper cites.
SberQuAD – Russian Reading Comprehension Dataset: Description and Analysis
Pavel Efimov, Andrey Chertok, Leonid Boytsov, and Pavel Braslavski. 2020 · 1912
Earlier work this paper cites.
Introducing MANtIS: A Novel Multi-Domain Information Seeking Dialogues Dataset
Gustavo Penha, Alexandru Balan, and Claudia Hauff. 2019 · 1912
Earlier work this paper cites.
VQD: Visual Query Detection In Natural Scenes. In Proc. of NAACL-HLT . ACL, Minneapolis, Minnesota, 1955–1961
Manoj Acharya, Karan Jariwala, and Christopher Kanan. 2019a · 1961
Earlier work this paper cites.
Some Philosophical Problems From the Standpoint of Artificial Intelligence
John McCarthy and Patrick Hayes. 1969 · 1969
Earlier work this paper cites.
Shifting the Baseline: Single Modality Performance on Visual Navigation & QA. In Proc. of NAACL . 1977–1983
Jesse Thomason, Daniel Gordon, and Yonatan Bisk. 2019 · 1983
Earlier work this paper cites.
Strategies of Discourse Comprehension
Teun A. van Dijk and Walter Kintsch. 1983 · 1983
Earlier work this paper cites.
The ATIS Spoken Language Systems Pilot Corpus. In Speech and Natural Language: Proceedings of a Workshop Held at Hidden Valley, Pennsylvania, June 24-27,1990
Charles T. Hemphill, John J. Godfrey, and George R. Doddington. 1990 · 1990
Earlier work this paper cites.
Dialogue Acts in Verbmobil 2
Jan Alexandersson, Bianka Buschbeck-Wolf, Tsutomu Fujinami, Michael Kipp, Stephan Koch, Elisabeth Maier, Norbert Reithinger, Birte Schmitz, and Melanie Siegel. 1998 · 1998
Earlier work this paper cites.
Life of Brian
Graham Chapman, John Cleese, Terry Gilliam, Eric Idle, Terry Jones, Michael Palin, John Goldstone, Spike Milligan, Monty Python (Comedy troupe), Handmade Films, and Criterion Collection (Firm). 1999 · 1999
Earlier work this paper cites.
Logical Models of Argument
Carlos Iván Chesnevar, Ana Gabriela Maguitman, and Ronald Prescott Loui. 2000 · 2000
Earlier work this paper cites.
Building a Question Answering Test Collection. In SIGIR (SIGIR ’00) . ACM, New York, NY, USA, 200–207
Ellen M. Voorhees and Dawn M. Tice. 2000 · 2000
Earlier work this paper cites.
Monty Python and the Holy Grail
Michael White, Graham Chapman, John Cleese, Eric Idle, Terry Gilliam, Terry Jones, Michael Palin, John Goldstone, Mark Forstater, Connie Booth, Carol Cleveland, Neil Innes, Bee Duffell, John Young, Rita Davies, Avril Stewart, Sally Kinghorn, Terry Bedford, Monty Python (Comedy troupe), Python (Monty) Pictures, and Columbia TriStar Home Entertainment (Firm). 2001 · 2001
Earlier work this paper cites.
FQuAD: French Question Answering Dataset
Martin d’Hoffschmidt, Maxime Vidal, Wacim Belblidia, and Tom Brendlé. 2020 · 2002
Earlier work this paper cites.
Temporal Order Relations in Language Comprehension
Elke van der Meer, Reinhard Beyer, Bertram Heinze, and Isolde Badel. 2002 · 2002
Earlier work this paper cites.
SemEval-2017 Task 3: Community Question Answering. In Proceedings of the 11th International Workshop on Semantic Evaluations (SemEval-2017) . Vancouver, Canada, August 3 - 4, 2017, 27–48
Preslav Nakov, Doris Hoogeveen, Lluís Màrquez, Alessandro Moschitti, Hamdy Mubarak, Timothy Baldwin, and Karin Verspoor. 2017 · 2003
Earlier work this paper cites.
Viktor Schlegel, Marco Valentino, André Freitas, Goran Nenadic, and Riza Batista-Navarro. 2020b · 2003
Earlier work this paper cites.
Event-QA: A Dataset for Event-Centric Question Answering over Knowledge Graphs
Tarcísio Souza Costa, Simon Gottschalk, and Elena Demidova. 2020 · 2004
Earlier work this paper cites.
Meanings and Configurations of Questions in English. In Proceedings of International Conference on Speech Prosody . Nara, Japan, 309–312
Nancy Hedberg, Juan M. Sosa, and Lorna Fadden. 2004 · 2004
Earlier work this paper cites.
Jiaqi Li, Ming Liu, Min-Yen Kan, Zihao Zheng, Zekun Wang, Wenqiang Lei, Ting Liu, and Bing Qin. 2020 · 2004
Earlier work this paper cites.
AmbigQA: Answering Ambiguous Open-Domain Questions
Sewon Min, Julian Michael, Hannaneh Hajishirzi, and Luke Zettlemoyer. 2020 · 2004
Earlier work this paper cites.
LAReQA: Language-Agnostic Answer Retrieval from a Multilingual Pool
Uma Roy, Noah Constant, Rami Al-Rfou, Aditya Barua, Aaron Phillips, and Yinfei Yang. 2020 · 2004
Earlier work this paper cites.
Language Models Are Few-Shot Learners
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020 · 2005
Earlier work this paper cites.
MultiReQA: A Cross-Domain Evaluation for Retrieval Question Answering Models
Mandy Guo, Yinfei Yang, Daniel Cer, Qinlan Shen, and Noah Constant. 2020 · 2005
Earlier work this paper cites.
UnifiedQA: Crossing Format Boundaries With a Single QA System
Daniel Khashabi, Tushar Khot, Ashish Sabharwal, Oyvind Tafjord, Peter Clark, and Hannaneh Hajishirzi. 2020 · 2005
Earlier work this paper cites.
RuBQ: A Russian Dataset for Question Answering over Wikidata
Vladislav Korablinov and Pavel Braslavski. 2020 · 2005
Earlier work this paper cites.
How Can We Accelerate Progress Towards Human-like Linguistic Generalization?
Tal Linzen. 2020 · 2005
Earlier work this paper cites.
Towards Question Format Independent Numerical Reasoning: A Set of Prerequisite Tasks
Swaroop Mishra, Arindam Mitra, Neeraj Varshney, Bhavdeep Sachdeva, and Chitta Baral. 2020 · 2005
Earlier work this paper cites.
TORQUE: A Reading Comprehension Dataset of Temporal Ordering Questions
Qiang Ning, Hao Wu, Rujun Han, Nanyun Peng, Matt Gardner, and Dan Roth. 2020 · 2005
Earlier work this paper cites.
Viktor Schlegel, Goran Nenadic, and Riza Batista-Navarro. 2020a · 2005
Earlier work this paper cites.
Evaluation of Text Generation: A Survey
Asli Celikyilmaz, Elizabeth Clark, and Jianfeng Gao. 2020 · 2006
Earlier work this paper cites.
TableQA: A Large-Scale Chinese Text-to-SQL Dataset for Table-Aware SQL Generation
Ningyuan Sun, Xuefeng Yang, and Yunfeng Liu. 2020 · 2006
Earlier work this paper cites.
Learning for Semantic Parsing with Statistical Machine Translation. In Proceedings of the Human Language Technology Conference of the NAACL, Main Conference . ACL, New York City, USA, 439–446
Yuk Wah Wong and Raymond Mooney. 2006 · 2006
Earlier work this paper cites.
MKQA: A Linguistically Diverse Benchmark for Multilingual Open Domain Question Answering
Shayne Longpre, Yi Lu, and Joachim Daiber. 2020 · 2007
Earlier work this paper cites.
A Dataset and Baselines for Visual Question Answering on Art
Noa Garcia, Chentao Ye, Zihua Liu, Qingtao Hu, Mayu Otani, Chenhui Chu, Yuta Nakashima, and Teruko Mitamura. 2020 · 2008
Earlier work this paper cites.
The Hitchhiker’s Guide to the Galaxy (del rey trade pbk. ed ed.)
Douglas Adams. 2009 · 2009
Earlier work this paper cites.
Chapter 9 Toward a Comprehensive Model of Comprehension
Danielle S. McNamara and Joe Magliano. 2009 · 2009
Earlier work this paper cites.
XOR QA: Cross-Lingual Open-Retrieval Question Answering
Akari Asai, Jungo Kasai, Jonathan H. Clark, Kenton Lee, Eunsol Choi, and Hannaneh Hajishirzi. 2020 · 2010
Earlier work this paper cites.
DaNetQA: A Yes/No Question Answering Dataset for the Russian Language
Taisia Glushkova, Alexey Machnev, Alena Fenogenova, Tatiana Shavrina, Ekaterina Artemova, and Dmitry I. Ignatov. 2020 · 2010
Earlier work this paper cites.
Explaining and Improving Model Behavior with k Nearest Neighbor Representations
Nazneen Fatema Rajani, Ben Krause, Wengpeng Yin, Tong Niu, Richard Socher, and Caiming Xiong. 2020 · 2010
Earlier work this paper cites.
Towards Data Distillation for End-to-End Spoken Conversational Question Answering
Chenyu You, Nuo Chen, Fenglin Liu, Dongchao Yang, and Yuexian Zou. 2020 · 2010
Earlier work this paper cites.
Beyond I.I.D.: Three Levels of Generalization for Question Answering on Knowledge Bases
Yu Gu, Sue Kase, Michelle Vanni, Brian Sadler, Percy Liang, Xifeng Yan, and Yu Su. 2021 · 2011
Earlier work this paper cites.
SRI’s Amex Travel Agent Data
SRI International. 2011 · 2011
Earlier work this paper cites.
Choice of Plausible Alternatives: An Evaluation of Commonsense Causal Reasoning. In AAAI Spring Symposium: Logical For- Malizations of Commonsense Reasoning . 6
Melissa Roemmele, Cosmin Adrian Bejan, and Andrew S Gordon. 2011 · 2011
Earlier work this paper cites.
When Do You Need Billions of Words of Pretraining Data?
Yian Zhang, Alex Warstadt, Haau-Sing Li, and Samuel R. Bowman. 2020 · 2011
Earlier work this paper cites.
SemEval-2012 Task 7: Choice of Plausible Alternatives: An Evaluation of Commonsense Causal Reasoning. In *SEM 2012: The First Joint Conference on Lexical and Computational Semantics . 394–398
Andrew Gordon, Zornitsa Kozareva, and Melissa Roemmele. 2012 · 2012
Earlier work this paper cites.
The Winograd Schema Challenge. In Proceedings of the Thirteenth International Conference on Principles of Knowledge Representation and Reasoning . 552–561
Hector J Levesque, Ernest Davis, and Leora Morgenstern. 2012 · 2012
Earlier work this paper cites.
Generating Natural Questions from Images for Multimodal Assistants
Alkesh Patel, Akanksha Bindal, Hadas Kotek, Christopher Klein, and Jason Williams. 2020 · 2012
Earlier work this paper cites.
Semantic Parsing on Freebase from Question-Answer Pairs. In Proc. of EMNLP . ACL, Seattle, Washington, USA, 1533–1544
Jonathan Berant, Andrew Chou, Roy Frostig, and Percy Liang. 2013 · 2013
Earlier work this paper cites.
Open Dataset for Development of Polish Question Answering Systems. In Proceedings of the 6th Language & Technology Conference: Human Language Technologies as a Challenge for Computer Science and Linguistics . Wydawnictwo Poznanskie, Fundacja Uniwersytetu im. Adama Mickiewicza
Michał Marcinczuk, Marcin Ptak, Adam Radziszewski, and Maciej Piasecki. 2013 · 2013
Earlier work this paper cites.
MCTest: A Challenge Dataset for the Open-Domain Machine Comprehension of Text. In EMNLP . Seattle, Washington, USA, 18-21 October 2013, 193–203
Matthew Richardson, Christopher J C Burges, and Erin Renshaw. 2013 · 2013
Earlier work this paper cites.
Modeling Biological Processes for Reading Comprehension. In Proc. of EMNLP . 1499–1510
Jonathan Berant, Vivek Srikumar, Pei-Chun Chen, Abby Vander Linden, Brittany Harding, Brad Huang, Peter Clark, and Christopher D. Manning. 2014 · 2014
Earlier work this paper cites.
Learning to Automatically Solve Algebra Word Problems. In ACL . ACL, Baltimore, Maryland, 271–281
Nate Kushman, Yoav Artzi, Luke Zettlemoyer, and Regina Barzilay. 2014 · 2014
Earlier work this paper cites.
Microsoft COCO: Common Objects in Context. In Computer Vision – ECCV 2014 , David Fleet, Tomas Pajdla, Bernt Schiele, and Tinne Tuytelaars (Eds.). Springer, Cham, 740–755
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C. Lawrence Zitnick. 2014 · 2014
Earlier work this paper cites.
Overview of CLEF Question Answering Track 2014. In Information Access Evaluation. Multilinguality, Multimodality, and Interaction . Springer, Cham, 300–306
Anselmo Peñas, Christina Unger, and Axel-Cyrille Ngonga Ngomo. 2014 · 2014
Earlier work this paper cites.
The NIPS Experiment
Eric Price. 2014 · 2014
Earlier work this paper cites.
Overview of the NTCIR-11 QA-Lab Task
Hideyuki Shibuki, Kotaro Sakamoto, Yoshionobu Kano, Teruko Mitamura, Madoka Ishioroshi, Kelly Y Itakura, Di Wang, Tatsunori Mori, and Noriko Kando. 2014 · 2014
Earlier work this paper cites.
VQA: Visual Question Answering. In 2015 IEEE International Conference on Computer Vision (ICCV) . Santiago, Chile, 2425–2433
Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C. Lawrence Zitnick, and Devi Parikh. 2015 · 2015
Earlier work this paper cites.
Large-Scale Simple Question Answering with Memory Networks
Antoine Bordes, Nicolas Usunier, Sumit Chopra, and Jason Weston. 2015 · 2015
Earlier work this paper cites.
Learning Hybrid Representations to Retrieve Semantically Equivalent Questions. In Proc. of ACL-IJCNLP . ACL, Beijing, China, 694–699
Cícero dos Santos, Luciano Barbosa, Dasha Bogdanova, and Bianca Zadrozny. 2015 · 2015
Earlier work this paper cites.
Teaching Machines to Read and Comprehend. In Proc. of NeurIPS . MIT Press, Cambridge, MA, USA, 1693–1701
Karl Moritz Hermann, Tomáš Kočiský, Edward Grefenstette, Lasse Espeholt, Will Kay, Mustafa Suleyman, and Phil Blunsom. 2015 · 2015
Earlier work this paper cites.
The Goldilocks Principle: Reading Children’s Books with Explicit Memory Representations
Felix Hill, Antoine Bordes, Sumit Chopra, and Jason Weston. 2015 · 2015
Earlier work this paper cites.
The Ubuntu Dialogue Corpus: A Large Dataset for Research in Unstructured Multi-Turn Dialogue Systems
Ryan Lowe, Nissan Pow, Iulian Serban, and Joelle Pineau. 2015 · 2015
Earlier work this paper cites.
SemEval-2015 Task 3: Answer Selection in Community Question Answering. 269–281
Preslav Nakov, Lluís Màrquez, Walid Magdy, Alessandro Moschitti, Jim Glass, and Bilal Randeree. 2015 · 2015
Earlier work this paper cites.
Compositional Semantic Parsing on Semi-Structured Tables. In ACL-IJCNLP . ACL, Beijing, China, 1470–1480
Panupong Pasupat and Percy Liang. 2015 · 2015
Earlier work this paper cites.
Overview of the CLEF Question Answering Track 2015. In Experimental IR Meets Multilinguality, Multimodality, and Interaction (Lecture Notes in Computer Science) . 539–544
Anselmo Peñas, Christina Unger, Georgios Paliouras, and Ioannis Kakadiaris. 2015 · 2015
Earlier work this paper cites.
"Answer Ka Type Kya He?": Learning to Classify Questions in Code-Mixed Language. In WWW (WWW ’15 Companion) . ACM, New York, NY, USA, 853–858
Khyathi Chandu Raghavi, Manoj Kumar Chinnakotla, and Manish Shrivastava. 2015 · 2015
Earlier work this paper cites.
Learning Answer-Entailing Structures for Machine Comprehension. In ACL-IJCNLP . ACL, Beijing, China, 239–249
Mrinmaya Sachan, Kumar Dubey, Eric Xing, and Matthew Richardson. 2015 · 2015
Earlier work this paper cites.
A Survey of Available Corpora for Building Data-Driven Dialogue Systems
Iulian Vlad Serban, Ryan Lowe, Peter Henderson, Laurent Charlin, and Joelle Pineau. 2015 · 2015
Earlier work this paper cites.
Automatically Solving Number Word Problems by Semantic Parsing and Reasoning. In EMNLP . ACL, Lisbon, Portugal, 1132–1142
Shuming Shi, Yuehui Wang, Chin-Yew Lin, Xiaojiang Liu, and Yong Rui. 2015 · 2015
Earlier work this paper cites.
An Overview of the BIOASQ Large-Scale Biomedical Semantic Indexing and Question Answering Competition
George Tsatsaronis, Georgios Balikas, Prodromos Malakasiotis, Ioannis Partalas, Matthias Zschunke, Michael R. Alvers, Dirk Weissenborn, Anastasia Krithara, Sergios Petridis, Dimitris Polychronopoulos, Yannis Almirantis, John Pavlopoulos, Nicolas Baskiotis, Patrick Gallinari, Thierry Artiéres, Axel-Cyrille Ngonga Ngomo, Norman Heino, Eric Gaussier, Liliana Barrio-Alvers, Michael Schroeder, Ion Androutsopoulos, and Georgios Paliouras. 2015 · 2015
Earlier work this paper cites.
Towards AI-Complete Question Answering: A Set of Prerequisite Toy Tasks
Jason Weston, Antoine Bordes, Sumit Chopra, Alexander M. Rush, Bart van Merriënboer, Armand Joulin, and Tomas Mikolov. 2015 · 2015
Earlier work this paper cites.
MS MARCO: A Human Generated MAchine Reading COmprehension Dataset
Payal Bajaj, Daniel Campos, Nick Craswell, Li Deng, Jianfeng Gao, Xiaodong Liu, Rangan Majumder, Andrew McNamara, Bhaskar Mitra, Tri Nguyen, Mir Rosenberg, Xia Song, Alina Stoica, Saurabh Tiwary, and Tong Wang. 2016 · 2016
Earlier work this paper cites.
The First Cross-Script Code-Mixed Question Answering Corpus. In Proc. of Workshop on Modeling, Learning and Mining for Cross/Multilinguality (MultiLingMine). Co-located with ECIR , Vol. 1589. CEUR-WS.org, Padua, Italy, 56–65
Somnath Banerjee, Sudip Kumar Naskar, and Paolo Rosso. 2016 · 2016
Earlier work this paper cites.
Consensus Attention-Based Neural Networks for Chinese Reading Comprehension. In Proc. of COLING . International Committee on Computational Linguistics, Osaka, Japan, 1777–1786
Yiming Cui, Ting Liu, Zhipeng Chen, Shijin Wang, and Guoping Hu. 2016 · 2016
Earlier work this paper cites.
WikiReading: A Novel Large-Scale Language Understanding Task over Wikipedia. In Proc. of ACL . ACL, Berlin, Germany, 1535–1545
Daniel Hewlett, Alexandre Lacoste, Llion Jones, Illia Polosukhin, Andrew Fandrianto, Jay Han, Matthew Kelcey, and David Berthelot. 2016 · 2016
Earlier work this paper cites.
Dialog State Tracking Challenge 5 Handbook v.3.1
Seokhwan Kim, Luis Ferdinando D’Haro, Rafael E. Banchs, Matthew Henderson, Jason Willisams, and Koichiro Yoshino. 2016 · 2016
Earlier work this paper cites.
Dataset and Neural Recurrent Sequence Labeling Model for Open-Domain Factoid Question Answering
Peng Li, Wei Li, Zhengyan He, Xuguang Wang, Ying Cao, Jie Zhou, and Wei Xu. 2016 · 2016
Earlier work this paper cites.
Addressing Complex and Subjective Product-Related Queries with Customer Reviews. In WWW (WWW ’16) . International WWW Conferences Steering Committee, Republic and Canton of Geneva, CHE, 625–635
Julian McAuley and Alex Yang. 2016 · 2016
Earlier work this paper cites.
InScript: Narrative texts annotated with script information. In Proceedings of the Tenth International Conference on Language Resources and Evaluation (LREC’16) . ELRA, Portorož, Slovenia, 3485–3493
Ashutosh Modi, Tatjana Anikina, Simon Ostermann, and Manfred Pinkal. 2016 · 2016
Earlier work this paper cites.
SemEval-2016 Task 3: Community Question Answering. 525–545
Preslav Nakov, Lluís Màrquez, Alessandro Moschitti, Walid Magdy, Hamdy Mubarak, abed Alhakim Freihat, Jim Glass, and Bilal Randeree. 2016 · 2016
Earlier work this paper cites.
Who Did What: A Large-Scale Person-Centered Cloze Dataset. In Proc. of EMNLP . ACL, Austin, Texas, 2230–2235
Takeshi Onishi, Hai Wang, Mohit Bansal, Kevin Gimpel, and David McAllester. 2016 · 2016
Earlier work this paper cites.
The LAMBADA Dataset: Word Prediction Requiring a Broad Discourse Context
Denis Paperno, Germán Kruszewski, Angeliki Lazaridou, Quan Ngoc Pham, Raffaella Bernardi, Sandro Pezzelle, Marco Baroni, Gemma Boleda, and Raquel Fernández. 2016 · 2016
Earlier work this paper cites.
SQuAD: 100,000+ Questions for Machine Comprehension of Text. In EMNLP . 2383–2392
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016 · 2016
Earlier work this paper cites.
"Why Should I Trust You?": Explaining the Predictions of Any Classifier. In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (KDD ’16) . ACM, San Francisco, California, USA, 1135–1144
Marco Tulio Ribeiro, Sameer Singh, and Carlos Guestrin. 2016 · 2016
Earlier work this paper cites.
An Analysis of Prerequisite Skills for Reading Comprehension. In Proceedings of the Workshop on Uphill Battles in Language Processing: Scaling Early Achievements to Robust Methods . ACL, Austin, TX, 1–5
Saku Sugawara and Akiko Aizawa. 2016 · 2016
Earlier work this paper cites.
MovieQA: Understanding Stories in Movies through Question-Answering. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR)
Makarand Tapaswi, Yukun Zhu, Rainer Stiefelhagen, Antonio Torralba, Raquel Urtasun, and Sanja Fidler. 2016 · 2016
Earlier work this paper cites.
Towards Machine Comprehension of Spoken Content: Initial TOEFL Listening Comprehension Test by Machine. In Interspeech 2016 . 2731–2735
Bo-Hsiang Tseng, Sheng-syun Shen, Hung-Yi Lee, and Lin-Shan Lee. 2016 · 2016
Cited alongside, same era.
Situation Models, Mental Simulations, and Abstract Concepts in Discourse Comprehension
Rolf A. Zwaan. 2016 · 2016
Cited alongside, same era.
Frames: A Corpus for Adding Memory to Goal-Oriented Dialogue Systems
Layla El Asri, Hannes Schulz, Shikhar Sharma, Jeremie Zumer, Justin Harris, Emery Fine, Rahul Mehrotra, and Kaheer Suleman. 2017 · 2017
Cited alongside, same era.
Embracing Data Abundance: BookTest Dataset for Reading Comprehension. In Proc. of ICLR
Ondrej Bajgar, Rudolf Kadlec, and Jan Kleindienst. 2017 · 2017
Cited alongside, same era.
Counting Everyday Objects in Everyday Scenes. In CVPR . 1135–1144
Prithvijit Chattopadhyay, Ramakrishna Vedantam, Ramprasaath R. Selvaraju, Dhruv Batra, and Devi Parikh. 2017 · 2017
How the Transformers Broke NLP Leaderboards
Anna Rogers. 2019 · 2019
Later among the works it cites.
Social IQa: Commonsense Reasoning about Social Interactions. In Proc. of EMNLP-IJCNLP . ACL, Hong Kong, China, 4453–4463
Maarten Sap, Hannah Rashkin, Derek Chen, Ronan Le Bras, and Yejin Choi. 2019 · 2019
Later among the works it cites.
DRCD: A Chinese Machine Reading Comprehension Dataset
Chih Chieh Shao, Trois Liu, Yuting Lai, Yiying Tseng, and Sam Tsai. 2019 · 2019
Later among the works it cites.
CLUTRR: A Diagnostic Benchmark for Inductive Reasoning from Text. In Proc. of EMNLP-IJCNLP . ACL, Hong Kong, China, 4496–4505
Koustuv Sinha, Shagun Sodhani, Jin Dong, Joelle Pineau, and William L. Hamilton. 2019 · 2019
Later among the works it cites.
A Corpus for Reasoning about Natural Language Grounded in Photographs. In Proc. of ACL . ACL, Florence, Italy, 6418–6428
Alane Suhr, Stephanie Zhou, Ally Zhang, Iris Zhang, Huajun Bai, and Yoav Artzi. 2019 · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Reading Wikipedia to Answer Open-Domain Questions. In Proc. of ACL . ACL, Vancouver, Canada, 1870–1879
Danqi Chen, Adam Fisch, Jason Weston, and Antoine Bordes. 2017 · 2017
Cited alongside, same era.
Abduction
Igor Douven. 2017 · 2017
Cited alongside, same era.
Overview of the NLPCC 2017 Shared Task: Open Domain Chinese Question Answering. In Natural Language Processing and Chinese Computing , Xuanjing Huang, Jing Jiang, Dongyan Zhao, Yansong Feng, and Yu Hong (Eds.). Springer, Cham, 954–961
Nan Duan and Duyu Tang. 2018 · 2017
Cited alongside, same era.
SearchQA: A New Q&A Dataset Augmented with Context from a Search Engine
Matthew Dunn, Levent Sagun, Mike Higgins, V. Ugur Guney, Volkan Cirik, and Kyunghyun Cho. 2017 · 2017
Cited alongside, same era.
Key-Value Retrieval Networks for Task-Oriented Dialogue
Mihail Eric and Christopher D. Manning. 2017 · 2017
Cited alongside, same era.
IJCNLP-2017 Task 5: Multi-Choice Question Answering in Examinations. In Proceedings of the IJCNLP 2017, Shared Tasks . Asian Federation of Natural Language Processing, Taipei, Taiwan, 34–40
Shangmin Guo, Kang Liu, Shizhu He, Cao Liu, Jun Zhao, and Zhuoyu Wei. 2017 · 2017
Cited alongside, same era.
Annotation Artifacts in Natural Language Inference Data. In Proc. of NAACL-HLT . ACL, New Orleans, Louisiana, 107–112
Suchin Gururangan, Swabha Swayamdipta, Omer Levy, Roy Schwartz, Samuel Bowman, and Noah A. Smith. 2018 · 2017
Cited alongside, same era.
Later among the works it cites.
DREAM: A Challenge Data Set and Models for Dialogue-Based Reading Comprehension
Kai Sun, Dian Yu, Jianshu Chen, Dong Yu, Yejin Choi, and Claire Cardie. 2019 · 2019
Later among the works it cites.
QuaRel: A Dataset and Models for Answering Questions about Qualitative Relationships. In AAAI 2019
Oyvind Tafjord, Peter Clark, Matt Gardner, Wen-tau Yih, and Ashish Sabharwal. 2019a · 2019
Later among the works it cites.
CommonsenseQA: A Question Answering Challenge Targeting Commonsense Knowledge. In NAACL . 4149–4158
Alon Talmor, Jonathan Herzig, Nicholas Lourie, and Jonathan Berant. 2019 · 2019
Later among the works it cites.
Best Practices for the Human Evaluation of Automatically Generated Text. In Proceedings of the 12th International Conference on Natural Language Generation . ACL, Tokyo, Japan, 355–368
Chris van der Lee, Albert Gatt, Emiel van Miltenburg, Sander Wubben, and Emiel Krahmer. 2019 · 2019
Later among the works it cites.
TWEETQA: A Social Media Focused Question Answering Dataset. In Proc. of ACL . ACL, Florence, Italy, 5020–5031
Wenhan Xiong, Jiawei Wu, Hong Wang, Vivek Kulkarni, Mo Yu, Shiyu Chang, Xiaoxiao Guo, and William Yang Wang. 2019 · 2019
Later among the works it cites.
FriendsQA: Open-Domain Question Answering on TV Show Transcripts. In Proceedings of the 20th Annual SIGdial Meeting on Discourse and Dialogue . ACL, Stockholm, Sweden, 188–197
Zhengzhe Yang and Jinho D. Choi. 2019 · 2019
Later among the works it cites.
A Qualitative Comparison of CoQA, SQuAD 2.0 and QuAC. In Proc. of NAACL-HLT . 2318–2323
Mark Yatskar. 2019 · 2019
Later among the works it cites.
ReClor: A Reading Comprehension Dataset Requiring Logical Reasoning. In Proc. of ICLR
Weihao Yu, Zihang Jiang, Yanfei Dong, and Jiashi Feng. 2019 · 2019
Later among the works it cites.
“Going on a Vacation” Takes Longer than “Going for a Walk”: A Study of Temporal Commonsense Understanding. In Proc. of EMNLP-IJCNLP . ACL, Hong Kong, China, 3361–3367
Ben Zhou, Daniel Khashabi, Qiang Ning, and Dan Roth. 2019 · 2019
Later among the works it cites.
Generating Fact Checking Explanations. In Proc. of ACL . ACL, Online, 7352–7364
Pepa Atanasova, Jakob Grue Simonsen, Christina Lioma, and Isabelle Augenstein. 2020 · 2020
Later among the works it cites.
STARC: Structured Annotations for Reading Comprehension. In Proc. of ACL . ACL, Online, 5726–5735
Yevgeni Berzak, Jonathan Malmaud, and Roger Levy. 2020 · 2020
Later among the works it cites.
Experience Grounds Language. In EMNLP . ACL, Online, 8718–8735
Yonatan Bisk, Ari Holtzman, Jesse Thomason, Jacob Andreas, Yoshua Bengio, Joyce Chai, Mirella Lapata, Angeliki Lazaridou, Jonathan May, Aleksandr Nisnevich, Nicolas Pinto, and Joseph Turian. 2020a · 2020
Later among the works it cites.
PIQA: Reasoning about Physical Commonsense in Natural Language
Yonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao, and Yejin Choi. 2020b · 2020
Later among the works it cites.
SubjQA: A Dataset for Subjectivity and Review Comprehension. In Proc. of EMNLP . ACL, Online, 5480–5494
Johannes Bjerva, Nikita Bhutani, Behzad Golshan, Wang-Chiew Tan, and Isabelle Augenstein. 2020 · 2020
Later among the works it cites.
Multiple Choice Questions: An Introductory Guide
Elisa Bone and Mike Prosser. 2020 · 2020
Later among the works it cites.
A Review of Public Datasets in Question Answering Research
B Barla Cambazoglu, Mark Sanderson, Falk Scholer, and Bruce Croft. 2020 · 2020
Later among the works it cites.
DoQA-Accessing Domain-Specific FAQs via Conversational QA. In Proc. of ACL . ACL, Online, 7302–7314
Jon Ander Campos, Arantxa Otegi, Aitor Soroa, Jan Deriu, Mark Cieliebak, and Eneko Agirre. 2020 · 2020
Later among the works it cites.
The TechQA Dataset. In Proc. of ACL . ACL, Online, 1269–1278
Vittorio Castelli, Rishav Chakravarti, Saswati Dana, Anthony Ferritto, Radu Florian, Martin Franz, Dinesh Garg, Dinesh Khandelwal, Scott McCarley, Michael McCawley, Mohamed Nasr, Lin Pan, Cezar Pendus, John Pitrelli, Saurabh Pujar, Salim Roukos, Andrzej Sakrajda, Avi Sil, Rosario Uceda-Sosa, Todd Ward, and Rong Zhang. 2020 · 2020
Later among the works it cites.
MOCHA: A Dataset for Training and Evaluating Generative Reading Comprehension Metrics. In Proc. of EMNLP . ACL, Online, 6521–6532
Anthony Chen, Gabriel Stanovsky, Sameer Singh, and Matt Gardner. 2020a · 2020
Later among the works it cites.
Open-Domain Question Answering. In Proc. of ACL: Tutorial Abstracts . ACL, Online, 34–37
Danqi Chen and Wen-tau Yih. 2020 · 2020
Later among the works it cites.
HybridQA: A Dataset of Multi-Hop Question Answering over Tabular and Textual Data. In Findings of EMNLP 2020 . ACL, Online, 1026–1036
Wenhu Chen, Hanwen Zha, Zhiyu Chen, Wenhan Xiong, Hong Wang, and William Yang Wang. 2020b · 2020
Later among the works it cites.
TyDi QA: A Benchmark for Information-Seeking Question Answering in Typologically Diverse Languages
Jonathan H. Clark, Eunsol Choi, Michael Collins, Dan Garrette, Tom Kwiatkowski, Vitaly Nikolaev, and Jennimaria Palomaki. 2020 · 2020
Later among the works it cites.
TutorialVQA: Question Answering Dataset for Tutorial Videos. In LREC . ELRA, Marseille, France, 5450–5455
Anthony Colas, Seokhwan Kim, Franck Dernoncourt, Siddhesh Gupte, Zhe Wang, and Doo Soon Kim. 2020 · 2020
Later among the works it cites.
A Sentence Cloze Dataset for Chinese Machine Reading Comprehension. In Proc. of COLING . ICCL, Barcelona, Spain (Online), 6717–6723
Yiming Cui, Ting Liu, Ziqing Yang, Zhipeng Chen, Wentao Ma, Wanxiang Che, Shijin Wang, and Guoping Hu. 2020 · 2020
Later among the works it cites.
To Test Machine Comprehension, Start by Defining Comprehension. In Proc. of ACL . ACL, Online, 7839–7859
Jesse Dunietz, Greg Burnham, Akash Bharadwaj, Owen Rambow, Jennifer Chu-Carroll, and Dave Ferrucci. 2020 · 2020
Later among the works it cites.
Temporal Reasoning via Audio Question Answering
H. M. Fayek and J. Johnson. 2020 · 2020
Later among the works it cites.
Read and Reason with MuSeRC and RuCoS: Datasets for Machine Reading Comprehension for Russian. In Proc. of COLING . ICCL, Barcelona, Spain (Online), 6481–6497
Alena Fenogenova, Vladislav Mikhailov, and Denis Shevelev. 2020 · 2020
Later among the works it cites.
IIRC: A Dataset of Incomplete Information Reading Comprehension Questions. In EMNLP . ACL, Online, 1137–1147
James Ferguson, Matt Gardner, Hannaneh Hajishirzi, Tushar Khot, and Pradeep Dasigi. 2020 · 2020
Later among the works it cites.
Evaluating Models’ Local Decision Boundaries via Contrast Sets. In Findings of EMNLP 2020 . ACL, Online, 1307–1323
Matt Gardner, Yoav Artzi, Victoria Basmov, Jonathan Berant, Ben Bogin, Sihao Chen, Pradeep Dasigi, Dheeru Dua, Yanai Elazar, Ananth Gottumukkala, Nitish Gupta, Hannaneh Hajishirzi, Gabriel Ilharco, Daniel Khashabi, Kevin Lin, Jiangming Liu, Nelson F. Liu, Phoebe Mulcaire, Qiang Ning, Sameer Singh, Noah A. Smith, Sanjay Subramanian, Reut Tsarfaty, Eric Wallace, Ally Zhang, and Ben Zhou. 2020 · 2020
Later among the works it cites.
TANDA: Transfer and Adapt Pre-Trained Transformer Models for Answer Sentence Selection. In Proc. of AAAI , Vol. 34:05. 7780–7788
Siddhant Garg, Thuy Vu, and Alessandro Moschitti. 2020 · 2020
Later among the works it cites.
Timnit Gebru, Jamie Morgenstern, Briana Vecchione, Jennifer Wortman Vaughan, Hanna Wallach, Hal Daumé III, and Kate Crawford. 2020 · 2020
Later among the works it cites.
Linguistic Appropriateness and Pedagogic Usefulness of Reading Comprehension Questions. In LREC . ELRA, Marseille, France, 1753–1762
Andrea Horbach, Itziar Aldabe, Marie Bexte, Oier Lopez de Lacalle, and Montse Maritxalar. 2020 · 2020
Later among the works it cites.
Project PIAF: Building a Native French Question-Answering Dataset. In Proc. of LREC . ELRA, Marseille, France, 5481–5490
Rachel Keraron, Guillaume Lancrenon, Mathilde Bras, Frédéric Allary, Gilles Moyse, Thomas Scialom, Edmundo-Pavel Soriano-Morales, and Jacopo Staiano. 2020 · 2020
Later among the works it cites.
SCDE: Sentence Cloze Dataset with High Quality Distractors From Examinations. In Proc. of ACL . ACL, Online, 5668–5683
Xiang Kong, Varun Gangal, and Eduard Hovy. 2020 · 2020
Later among the works it cites.
MLQA: Evaluating Cross-Lingual Extractive Question Answering. In Proc. of ACL . ACL, Online, 7315–7330
Patrick Lewis, Barlas Oguz, Ruty Rinott, Sebastian Riedel, and Holger Schwenk. 2020 · 2020
Later among the works it cites.
Birds Have Four Legs?! NumerSense: Probing Numerical Commonsense Knowledge of Pre-Trained Language Models. In Proc. of EMNLP . ACL, Online, 6862–6868
Bill Yuchen Lin, Seyeon Lee, Rahul Khanna, and Xiang Ren. 2020 · 2020
Later among the works it cites.
LogiQA: A Challenge Dataset for Machine Reading Comprehension with Logical Reasoning.. In IJCAI , Christian Bessiere (Ed.). ijcai.org, 3622–3628
Jian Liu, Leyang Cui, Hanmeng Liu, Dandan Huang, Yile Wang, and Yue Zhang. 2020 · 2020
Later among the works it cites.
A Diverse Corpus for Evaluating and Developing English Math Word Problem Solvers. In Proc. of ACL . 975–984
Shen-yun Miao, Chao-Chun Liang, and Keh-Yih Su. 2020 · 2020
Later among the works it cites.
COVID-QA: A Question Answering Dataset for COVID-19. In Proceedings of the 1st Workshop on NLP for COVID-19 at ACL 2020 . ACL, Online
Timo Möller, Anthony Reina, Raghavan Jayakumar, and Malte Pietsch. 2020 · 2020
Later among the works it cites.
A Vietnamese Dataset for Evaluating Machine Reading Comprehension. In ICLR . International Committee on Computational Linguistics, Barcelona, Spain (Online), 2595–2605
Kiet Nguyen, Vu Nguyen, Anh Nguyen, and Ngan Nguyen. 2020 · 2020
Later among the works it cites.
A Method for Building a Commonsense Inference Dataset Based on Basic Events. In Proc. of EMNLP . ACL, Online, 2450–2460
Kazumasa Omura, Daisuke Kawahara, and Sadao Kurohashi. 2020 · 2020
Later among the works it cites.
Neural Unsupervised Domain Adaptation in NLP—A Survey. In Proc. of COLING . ICCL, Barcelona, Spain (Online), 6838–6855
Alan Ramponi and Barbara Plank. 2020 · 2020
Later among the works it cites.
Beyond Accuracy: Behavioral Testing of NLP Models with CheckList. In Proc. of ACL . ACL, Online, 4902–4912
Marco Tulio Ribeiro, Tongshuang Wu, Carlos Guestrin, and Sameer Singh. 2020 · 2020
Later among the works it cites.
Information Seeking in the Spirit of Learning: A Dataset for Conversational Curiosity. In Proc. of EMNLP . ACL, Online, 8153–8172
Pedro Rodriguez, Paul Crook, Seungwhan Moon, and Zhiguang Wang. 2020 · 2020
Later among the works it cites.
What Can We Do to Improve Peer Review in NLP?. In Findings of EMNLP . ACL, Online, 1256–1262
Anna Rogers and Isabelle Augenstein. 2020 · 2020
Later among the works it cites.
Getting Closer to AI Complete Question Answering: A Set of Prerequisite Real Tasks. In AAAI . 8722–8731
Anna Rogers, Olga Kovaleva, Matthew Downey, and Anna Rumshisky. 2020 · 2020
Later among the works it cites.
Thinking Like a Skeptic: Defeasible Inference in Natural Language. In Findings of EMNLP 2020 . ACL, Online, 4661–4675
Rachel Rudinger, Vered Shwartz, Jena D. Hwang, Chandra Bhagavatula, Maxwell Forbes, Ronan Le Bras, Noah A. Smith, and Yejin Choi. 2020 · 2020
Later among the works it cites.
Text-to-SQL Generation for Question Answering on Electronic Medical Records. In Proceedings of The Web Conference 2020 (WWW ’20) . ACM, New York, NY, USA, 350–361
Ping Wang, Tian Shi, and Chandan K. Reddy. 2020a · 2020
Later among the works it cites.
Developing Dataset of Japanese Slot Filling Quizzes Designed for Evaluation of Machine Reading Comprehension. In LREC . ELRA, Marseille, France, 6895–6901
Takuto Watarai and Masatoshi Tsuchiya. 2020 · 2020
Later among the works it cites.
MATINF: A Jointly Labeled Large-Scale Dataset for Classification, Question Answering and Summarization. In Proc. of ACL . ACL, Online, 3586–3596
Canwen Xu, Jiaxin Pei, Hongtao Wu, Yiyu Liu, and Chenliang Li. 2020 · 2020
Later among the works it cites.
Question Answering with Long Multiple-Span Answers. In Findings of EMNLP 2020 . ACL, Online, 3840–3849
Ming Zhu, Aman Ahuja, Da-Cheng Juan, Wei Wei, and Chandan K. Reddy. 2020 · 2020
Later among the works it cites.
CrossVQA: Scalably Generating Benchmarks for Systematically Testing VQA Generalization. In Proc. of EMNLP . ACL, Online and Punta Cana, Dominican Republic, 2148–2166
Arjun Akula, Soravit Changpinyo, Boqing Gong, Piyush Sharma, Song-Chun Zhu, and Radu Soricut. 2021 · 2021
Closest in time.
Open-Domain Question Answering Goes Conversational via Question Rewriting. In NAACL . ACL, Online, 520–534
Raviteja Anantha, Svitlana Vakulenko, Zhucheng Tu, Shayne Longpre, Stephen Pulman, and Srinivas Chappidi. 2021 · 2021
Closest in time.
Challenges in Information-Seeking QA: Unanswerable Questions and Paragraph Retrieval. In Proc. of ACL-IJCNLP . ACL, Online, 1492–1504
Akari Asai and Eunsol Choi. 2021 · 2021
Closest in time.
Evaluating Entity Disambiguation and the Role of Popularity in Retrieval-Based NLP. In Proc. of ACL . ACL, Online, 4472–4485
Anthony Chen, Pallavi Gudipati, Shayne Longpre, Xiao Ling, and Sameer Singh. 2021b · 2021
Closest in time.
WebSRC: A Dataset for Web-Based Structural Reading Comprehension. In Proc. of EMNLP . ACL, Online and Punta Cana, Dominican Republic, 4173–4185
Xingyu Chen, Zihan Zhao, Lu Chen, JiaBao Ji, Danyang Zhang, Ao Luo, Yuxuan Xiong, and Kai Yu. 2021c · 2021
Closest in time.
FinQA: A Dataset of Numerical Reasoning over Financial Data. In Proc. of EMNLP . ACL, Online and Punta Cana, Dominican Republic, 3697–3711
Zhiyu Chen, Wenhu Chen, Charese Smiley, Sameena Shah, Iana Borova, Dylan Langdon, Reema Moussa, Matt Beane, Ting-Hao Huang, Bryan Routledge, and William Yang Wang. 2021a · 2021
Closest in time.
Decontextualization: Making Sentences Stand-Alone
Eunsol Choi, Jennimaria Palomaki, Matthew Lamm, Tom Kwiatkowski, Dipanjan Das, and Michael Collins. 2021 · 2021
Closest in time.
Perhaps PTLMs Should Go to School – A Task to Assess Open Book and Closed Book QA. In Proc. of EMNLP . ACL, Online and Punta Cana, Dominican Republic, 6104–6111
Manuel Ciosici, Joe Cecil, Dong-Ho Lee, Alex Hedges, Marjorie Freedman, and Ralph Weischedel. 2021 · 2021
Closest in time.
A Dataset of Information-Seeking Questions and Answers Anchored in Research Papers. In Proc. of NAACL-HLT . ACL, Online, 4599–4610
Pradeep Dasigi, Kyle Lo, Iz Beltagy, Arman Cohan, Noah A. Smith, and Matt Gardner. 2021 · 2021
Closest in time.
English Machine Reading Comprehension Datasets: A Survey. In Proc. of EMNLP . ACL, Online and Punta Cana, Dominican Republic, 8784–8804
Daria Dzendzik, Jennifer Foster, and Carl Vogel. 2021 · 2021
Closest in time.
Competency Problems: On Finding and Removing Artifacts in Language Data. In Proc. of EMNLP . ACL, Online and Punta Cana, Dominican Republic, 1801–1813
Matt Gardner, William Merrill, Jesse Dodge, Matthew Peters, Alexis Ross, Sameer Singh, and Noah A. Smith. 2021 · 2021
Closest in time.
On the Interaction of Belief Bias and Explanations. In Findings of ACL-IJCNLP 2021 . ACL, Online, 2930–2942
Ana Valeria González, Anna Rogers, and Anders Søgaard. 2021 · 2021
Closest in time.
Aditya Gupta, Jiacheng Xu, Shyam Upadhyay, Diyi Yang, and Manaal Faruqui. 2021 · 2021
Closest in time.
ESTER: A Machine Reading Comprehension Dataset for Reasoning about Event Semantic Relations. In Proc. of EMNLP . ACL, Online and Punta Cana, Dominican Republic, 7543–7559
Rujun Han, I-Hung Hsu, Jiao Sun, Julia Baylon, Qiang Ning, Dan Roth, and Nanyun Peng. 2021 · 2021
Closest in time.
Inductive Logic
James Hawthorne. 2021 · 2021
Closest in time.
Hurdles to Progress in Long-Form Question Answering. In NAACL-HLT . ACL, Online, 4940–4957
Kalpesh Krishna, Aurko Roy, and Mohit Iyyer. 2021 · 2021
Closest in time.
MLEC-QA: A Chinese Multi-Choice Biomedical Question Answering Dataset. In Proc. of EMNLP . ACL, Online and Punta Cana, Dominican Republic, 8862–8874
Jing Li, Shangping Zhong, and Kaizhi Chen. 2021 · 2021
Closest in time.
SPARTQA: A Textual Question Answering Benchmark for Spatial Reasoning. In Proc. of NAACL . ACL, Online, 4582–4598
Roshanak Mirzaee, Hossein Rajaby Faghihi, Qiang Ning, and Parisa Kordjamshidi. 2021 · 2021
Closest in time.
TIMEDIAL: Temporal Commonsense Reasoning in Dialog
Lianhui Qin, Aditya Gupta, Shyam Upadhyay, Luheng He, Yejin Choi, and Manaal Faruqui. 2021 · 2021
Closest in time.
Evaluation Paradigms in Question Answering. In Proc. of EMNLP . ACL, Online and Punta Cana, Dominican Republic, 9630–9642
Pedro Rodriguez and Jordan Boyd-Graber. 2021 · 2021
Closest in time.
Changing the World by Changing the Data. In ACL . ACL, Online, 2182–2194
Anna Rogers. 2021 · 2021
Closest in time.
Multi-Domain Multilingual Question Answering. In Proc. of EMNLP
Sebastian Ruder and Si Avirup. 2021 · 2021
Closest in time.
NLQuAD: A Non-Factoid Long Question Answering Data Set. In EACL . ACL, Online, 1245–1255
Amir Soleimani, Christof Monz, and Marcel Worring. 2021 · 2021
Closest in time.
MultimodalQA: Complex Question Answering Over Text, Tables and Images. In Proc. of ICLR . 12
Alon Talmor, Ori Yoran, Amnon Catav, Dan Lahav, Yizhong Wang, Akari Asai, Gabriel Ilharco, Hannaneh Hajishirzi, and Jonathan Berant. 2021 · 2021
Closest in time.
BEIR: A Heterogeneous Benchmark for Zero-shot Evaluation of Information Retrieval Models. In Thirty-Fifth Conference on Neural Information Processing Systems, Datasets and Benchmarks Track
Nandan Thakur, Nils Reimers, Andreas Rücklé, Abhishek Srivastava, and Iryna Gurevych. 2021 · 2021
Closest in time.
Improving Question Answering for Event-Focused Questions in Temporal Collections of News Articles
Jiexin Wang, Adam Jatowt, Michael Färber, and Masatoshi Yoshikawa. 2021 · 2021
Closest in time.
On the Faithfulness Measurements for Model Interpretations
Fan Yin, Zhouxing Shi, Cho-Jui Hsieh, and Kai-Wei Chang. 2021 · 2021
Closest in time.
SituatedQA: Incorporating Extra-Linguistic Contexts into QA. In Proc. of EMNLP . ACL, Online and Punta Cana, Dominican Republic, 7371–7387
Michael Zhang and Eunsol Choi. 2021 · 2021
Closest in time.
Calibrate Before Use: Improving Few-Shot Performance of Language Models. In Proc. of ICML
Tony Z. Zhao, Eric Wallace, Shi Feng, Dan Klein, and Sameer Singh. 2021 · 2021
Closest in time.
TAT-QA: A Question Answering Benchmark on a Hybrid of Tabular and Textual Content in Finance. In Proc. of ACL . ACL, Online, 3277–3287
Fengbin Zhu, Wenqiang Lei, Youcheng Huang, Chao Wang, Shuo Zhang, Jiancheng Lv, Fuli Feng, and Tat-Seng Chua. 2021a · 2021
Closest in time.
Retrieving and Reading: A Comprehensive Survey on Open-Domain Question Answering
Fengbin Zhu, Wenqiang Lei, Chao Wang, Jianming Zheng, Soujanya Poria, and Tat-Seng Chua. 2021b · 2021
Closest in time.
TopiOCQA: Open-domain Conversational Question Answering with Topic Switching
Vaibhav Adlakha, Shehzaad Dhuliawala, Kaheer Suleman, Harm de Vries, and Siva Reddy. 2022 · 2022
Closest in time.
KQA Pro: A Dataset with Explicit Compositional Programs for Complex Question Answering over Knowledge Base. In Proc. of ACL . ACL, Dublin, Ireland, 6101–6119
Shulin Cao, Jiaxin Shi, Liangming Pan, Lunyiu Nie, Yutong Xiang, Lei Hou, Juanzi Li, Bin He, and Hanwang Zhang. 2022 · 2022
Closest in time.
Machine Reading, Fast and Slow: When Do Models "Understand" Language?
Sagnik Ray Choudhury, Anna Rogers, and Isabelle Augenstein. 2022 · 2022
Closest in time.
CARETS: A Consistency And Robustness Evaluative Test Suite for VQA. In Proc. of ACL . ACL, Dublin, Ireland, 6392–6405
Carlos E. Jimenez, Olga Russakovsky, and Karthik Narasimhan. 2022 · 2022
Closest in time.
AIT-QA: Question Answering Dataset over Complex Tables in the Airline Industry. In Proc. of NAACL-HLT . ACL, Hybrid: Seattle, Washington + Online, 305–314
Yannis Katsis, Saneem Chemmengath, Vishwajeet Kumar, Samarth Bharadwaj, Mustafa Canim, Michael Glass, Alfio Gliozzo, Feifei Pan, Jaydeep Sen, Karthik Sankaranarayanan, and Soumen Chakrabarti. 2022 · 2022
Closest in time.
MultiSpanQA: A Dataset for Multi-Span Question Answering. In Proc. of NAACL . ACL, Seattle, United States, 1250–1260
Haonan Li, Martin Tomko, Maria Vasardani, and Timothy Baldwin. 2022b · 2022
Closest in time.
MMCoQA: Conversational Question Answering over Text, Tables, and Images. In Proc. of ACL . ACL, Dublin, Ireland, 4220–4231
Yongqi Li, Wenjie Li, and Liqiang Nie. 2022a · 2022
Closest in time.
TruthfulQA: Measuring How Models Mimic Human Falsehoods. In Proc. of ACL . ACL, Dublin, Ireland, 3214–3252
Stephanie Lin, Jacob Hilton, and Owain Evans. 2022 · 2022
Closest in time.
ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning. In Findings of ACL 2022 . ACL, Dublin, Ireland, 2263–2279
Ahmed Masry, Do Long, Jia Qing Tan, Shafiq Joty, and Enamul Hoque. 2022 · 2022
Closest in time.
QuALITY: Question Answering with Long Input Texts, Yes!. In Proc. of NAACL . ACL, Seattle, United States, 5336–5358
Richard Yuanzhe Pang, Alicia Parrish, Nitish Joshi, Nikita Nangia, Jason Phang, Angelica Chen, Vishakh Padmakumar, Johnny Ma, Jana Thompson, He He, and Samuel Bowman. 2022 · 2022
Closest in time.
xGQA: Cross-Lingual Visual Question Answering. In Findings of ACL . ACL, Dublin, Ireland, 2497–2511
Jonas Pfeiffer, Gregor Geigle, Aishwarya Kamath, Jan-Martin Steitz, Stefan Roth, Ivan Vulić, and Iryna Gurevych. 2022 · 2022
Closest in time.
ConditionalQA: A Complex Reading Comprehension Dataset with Conditional Answers. In Proc. of ACL . ACL, Dublin, Ireland, 3627–3637
Haitian Sun, William Cohen, and Ruslan Salakhutdinov. 2022 · 2022
Closest in time.
ArchivalQA: A Large-scale Benchmark Dataset for Open Domain Question Answering over Historical News Collections
Jiexin Wang, Adam Jatowt, and Masatoshi Yoshikawa. 2022 · 2022
Closest in time.
QAConv: Question Answering on Informative Conversations. In Proc. of ACL . ACL, Dublin, Ireland, 5389–5411
Chien-Sheng Wu, Andrea Madotto, Wenhao Liu, Pascale Fung, and Caiming Xiong. 2022 · 2022
Closest in time.
Fantastic Questions and Where to Find Them: FairytaleQA – An Authentic Dataset for Narrative Comprehension. In Proc. of ACL . ACL, Dublin, Ireland, 447–460
Ying Xu, Dakuo Wang, Mo Yu, Daniel Ritchie, Bingsheng Yao, Tongshuang Wu, Zheng Zhang, Toby Li, Nora Bradford, Branda Sun, Tran Hoang, Yisi Sang, Yufang Hou, Xiaojuan Ma, Diyi Yang, Nanyun Peng, Zhou Yu, and Mark Warschauer. 2022 · 2022
Closest in time.
Adversarial Examples for Evaluating Reading Comprehension Systems. In Proc. of EMNLP . ACL, 2021–2031
Robin Jia and Percy Liang. 2017 · 2031
Closest in time.
A Corpus of Natural Language for Visual Reasoning. In Proc. of ACL . 217–223
Alane Suhr, Mike Lewis, James Yeh, and Yoav Artzi. 2017 · 2034
Closest in time.
Challenging Reading Comprehension on Daily Conversation: Passage Completion on Multiparty Dialog. In Proc. of NAACL . ACL, New Orleans, Louisiana, 2039–2048
Kaixin Ma, Tomasz Jurczyk, and Jinho D. Choi. 2018 · 2048
Closest in time.
Interpretation of Natural Language Rules in Conversational Machine Reading. In Proc. of EMNLP . ACL, Brussels, Belgium, 2087–2097
Marzieh Saeidi, Max Bartolo, Patrick Lewis, Sameer Singh, Tim Rocktäschel, Mike Sheldon, Guillaume Bouchard, and Sebastian Riedel. 2018 · 2097
Closest in time.