Triviaqa: A large scale distantly supervised challenge dataset for reading comprehension
Original
Mandar Joshi, Eunsol Choi, Daniel S Weld, and Luke Zettlemoyer. 2017 · 2017
Later among the works it cites.
e-snli: Natural language inference with natural language explanations
Oana-Maria Camburu, Tim Rocktäschel, Thomas Lukasiewicz, and Phil Blunsom. 2018 · 2018
Later among the works it cites.
Learning to explain: An information-theoretic perspective on model interpretation
Original
Jianbo Chen, Le Song, Martin J Wainwright, and Michael I Jordan. 2018 · 2018
Later among the works it cites.
Multi-step retriever-reader interaction for scalable open-domain question answering
Rajarshi Das, Shehzaad Dhuliawala, Manzil Zaheer, and Andrew McCallum. 2018 · 2018
Later among the works it cites.
Comparing automatic and human evaluation of local explanations for text classification
Dong Nguyen. 2018 · 2018
Later among the works it cites.
Hotpotqa: A dataset for diverse, explainable multi-hop question answering
Original
Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William W Cohen, Ruslan Salakhutdinov, and Christopher D Manning. 2018 · 2018
Later among the works it cites.
Guidelines for human-ai interaction
Saleema Amershi, Dan Weld, Mihaela Vorvoreanu, Adam Fourney, Besmira Nushi, Penny Collisson, Jina Suh, Shamsi Iqbal, Paul N Bennett, Kori Inkpen, et al. 2019 · 2019
Later among the works it cites.
Beyond accuracy: The role of mental models in human-ai team performance
Gagan Bansal, Besmira Nushi, Ece Kamar, Walter S Lasecki, Daniel S Weld, and Eric Horvitz. 2019 · 2019
Later among the works it cites.
Natural questions: a benchmark for question answering research
Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, et al. 2019 · 2019
Later among the works it cites.
Towards explainable nlp: A generative explanation framework for text classification
Hui Liu, Qingyu Yin, and William Yang Wang. 2019 · 2019
Later among the works it cites.
Right for the wrong reasons: Diagnosing syntactic heuristics in natural language inference
R Thomas McCoy, Ellie Pavlick, and Tal Linzen. 2019 · 2019
Later among the works it cites.
Explanation in artificial intelligence: Insights from the social sciences
Tim Miller. 2019 · 2019
Later among the works it cites.
Explain yourself! leveraging language models for commonsense reasoning
Nazneen Fatema Rajani, Bryan McCann, Caiming Xiong, and Richard Socher. 2019 · 2019
Later among the works it cites.
Proxy tasks and subjective measures can be misleading in evaluating explainable ai systems
Zana Buçinca, Phoebe Lin, Krzysztof Z Gajos, and Elena L Glassman. 2020 · 2020
Closest in time.
Eraser: A benchmark to evaluate rationalized nlp models
Jay DeYoung, Sarthak Jain, Nazneen Fatema Rajani, Eric Lehman, Caiming Xiong, Richard Socher, and Byron C. Wallace. 2020 · 2020
Closest in time.