Fetching the paper…
Reading the bibliography…
Explainable NLP (ExNLP) has increasingly focused on collecting human-annotated textual explanations.
Automated rationale generation: A technique for explainable AI and its effects on human perceptions
Upol Ehsan, Pradyumna Tambwekar, Larry Chan, Brent Harrison, and Mark O. Riedl · 1901
Earlier work this paper cites.
Visual entailment: A novel task for fine-grained image understanding
Ning Xie, Farley Lai, Derek Doran, and Asim Kadav · 1901
Earlier work this paper cites.
Evaluating explanation without ground truth in interpretable machine learning
Fan Yang, Mengnan Du, and Xia Hu · 1907
Earlier work this paper cites.
Abductive commonsense reasoning
Chandra Bhagavatula, Ronan Le Bras, Chaitanya Malaviya, Keisuke Sakaguchi, Ari Holtzman, Hannah Rashkin, Doug Downey, Wen-tau Yih, and Yejin Choi · 1908
Earlier work this paper cites.
QASC: A dataset for question answering via sentence composition
Tushar Khot, Peter Clark, Michal Guerquin, Peter Jansen, and Ashish Sabharwal · 1910
Earlier work this paper cites.
Learning from explanations with neural execution tree
Ziqi Wang, Yujia Qin, Wenxuan Zhou, Jun Yan, Qinyuan Ye, Leonardo Neves, Zhiyuan Liu, and Xiang Ren · 1911
Earlier work this paper cites.
Telling more than we can know: Verbal reports on mental processes
R. Nisbett and T. Wilson · 1977
Earlier work this paper cites.
Conversational processes and causal explanation
Denis J Hilton · 1990
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu · 2002
Earlier work this paper cites.
e-SNLI-VE-2.0: Corrected Visual-Textual Entailment with Natural Language Explanations
Virginie Do, Oana-Maria Camburu, Zeynep Akata, and Thomas Lukasiewicz · 2004
Earlier work this paper cites.
WT5?! Training Text-to-Text Models to Explain their Predictions
Sharan Narang, Colin Raffel, Katherine Lee, Adam Roberts, Noah Fiedel, and Karishma Malkan · 2004
Earlier work this paper cites.
Seeing stars: Exploiting class relationships for sentiment categorization with respect to rating scales
Bo Pang and Lillian Lee · 2005
Earlier work this paper cites.
Aligning faithful interpretations with their social attribution
Alon Jacovi and Yoav Goldberg · 2006
Earlier work this paper cites.
Using “annotator rationales” to improve machine learning for text categorization
Omar Zaidan, Jason Eisner, and Christine Piatko · 2007
Earlier work this paper cites.
A survey on explainability in machine reading comprehension
Mokanarangan Thayaparan, Marco Valentino, and André Freitas · 2010
Earlier work this paper cites.
Counterfactual explanations for machine learning: A review
Sahil Verma, John P. Dickerson, and Keegan Hines · 2010
Earlier work this paper cites.
Measuring association between labels and free-text rationales
Sarah Wiegreffe, Ana Marasović, and Noah A. Smith · 2010
Earlier work this paper cites.
Learning to rationalize for nonmonotonic reasoning with distant supervision
Faeze Brahman, Vered Shwartz, Rachel Rudinger, and Yejin Choi · 2012
Earlier work this paper cites.
Hatexplain: A benchmark dataset for explainable hate speech detection
Binny Mathew, Punyajoy Saha, Seid Muhie Yimam, Chris Biemann, Pawan Goyal, and Animesh Mukherjee · 2012
Earlier work this paper cites.
Learning attitudes and attributes from multi-aspect reviews
Julian McAuley, Jure Leskovec, and Dan Jurafsky · 2012
Earlier work this paper cites.
Data and its (dis)contents: A survey of dataset development and use in machine learning research
Amandalynne Paullada, Inioluwa Deborah Raji, Emily M. Bender, Emily L. Denton, and A. Hanna · 2012
Earlier work this paper cites.
To what extent do human explanations of model behavior align with actual model behavior?
Grusha Prasad, Yixin Nie, Mohit Bansal, Robin Jia, Douwe Kiela, and Adina Williams · 2012
Earlier work this paper cites.
Crowd truth: Harnessing disagreement in crowdsourcing a relation extraction gold standard
Lora Aroyo and Chris Welty · 2013
Earlier work this paper cites.
MCTest: A challenge dataset for the open-domain machine comprehension of text
Matthew Richardson, Christopher J.C. Burges, and Erin Renshaw · 2013
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D. Manning, Andrew Ng, and Christopher Potts · 2013
Earlier work this paper cites.
From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions
Peter Young, Alice Lai, Micah Hodosh, and Julia Hockenmaier · 2014
Earlier work this paper cites.
Truth is a lie: Crowd truth and the seven myths of human annotation
Lora Aroyo and Chris Welty · 2015
Earlier work this paper cites.
A large annotated corpus for learning natural language inference
Samuel R. Bowman, Gabor Angeli, Christopher Potts, and Christopher D. Manning · 2015
Earlier work this paper cites.
An ontology design pattern to define explanations
Ilaria Tiddi, M. d’Aquin, and E. Motta · 2015
Earlier work this paper cites.
WikiQA: A challenge dataset for open-domain question answering
Yi Yang, Wen-tau Yih, and Christopher Meek · 2015
Earlier work this paper cites.
What’s in an explanation? characterizing knowledge and inference requirements for elementary science exams
Peter Jansen, Niranjan Balasubramanian, Mihai Surdeanu, and Peter Clark · 2016
Earlier work this paper cites.
Rationalizing neural predictions
Tao Lei, Regina Barzilay, and Tommi Jaakkola · 2016
Earlier work this paper cites.
SQuAD: 100,000+ questions for machine comprehension of text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang · 2016
Earlier work this paper cites.
Rationale-augmented convolutional neural networks for text classification
Ye Zhang, Iain Marshall, and Byron C. Wallace · 2016
Earlier work this paper cites.
Explanation and justification in machine learning: A survey
Or Biran and Courtenay Cotton · 2017
Earlier work this paper cites.
Towards a rigorous science of interpretable machine learning
Finale Doshi-Velez and Been Kim · 2017
Earlier work this paper cites.
Making the V in VQA Matter: Elevating the Role of Image Understanding in Visual Question Answering
Yash Goyal, Tejas Khot, Douglas Summers-Stay, Dhruv Batra, and D. Parikh · 2017
Earlier work this paper cites.
Annotation artifacts in natural language inference data
Suchin Gururangan, Swabha Swayamdipta, Omer Levy, Roy Schwartz, Samuel Bowman, and Noah A. Smith · 2017
Earlier work this paper cites.
RACE: Large-scale ReAding comprehension dataset from examinations
Guokun Lai, Qizhe Xie, Hanxiao Liu, Yiming Yang, and Eduard Hovy · 2017
Earlier work this paper cites.
Program induction by rationale generation: Learning to solve and explain algebraic word problems
Wang Ling, Dani Yogatama, Chris Dyer, and Phil Blunsom · 2017
Earlier work this paper cites.
Joint concept learning and semantic parsing from natural language explanations
Shashank Srivastava, Igor Labutov, and Tom Mitchell · 2017
Earlier work this paper cites.
Position-aware attention and supervised data improve slot filling
Yuhao Zhang, Victor Zhong, Danqi Chen, Gabor Angeli, and Christopher D. Manning · 2017
Earlier work this paper cites.
Peeking inside the black-box: A survey on explainable artificial intelligence (xai)
Amina Adadi and Mohammed Berrada · 2018
Earlier work this paper cites.
Where is your evidence: Improving fact-checking by justification modeling
Tariq Alhindi, Savvas Petridis, and Smaranda Muresan · 2018
Earlier work this paper cites.
Towards robust interpretability with self-explaining neural networks
David Alvarez-Melis and T. Jaakkola · 2018
Earlier work this paper cites.
Deriving machine attention from human rationales
Yujia Bao, Shiyu Chang, Mo Yu, and Regina Barzilay · 2018
Earlier work this paper cites.
Data statements for natural language processing: Toward mitigating system bias and enabling better science
Emily M. Bender and Batya Friedman · 2018
Earlier work this paper cites.
e-SNLI: Natural language inference with natural language explanations
Oana-Maria Camburu, Tim Rocktäschel, Thomas Lukasiewicz, and Phil Blunsom · 2018
Earlier work this paper cites.
Extractive adversarial networks: High-recall explanations for identifying personal attacks in social media posts
Samuel Carton, Qiaozhu Mei, and Paul Resnick · 2018
Earlier work this paper cites.
Explanations based on the missing: Towards contrastive explanations with pertinent negatives
Amit Dhurandhar, Pin-Yu Chen, Ronny Luss, Chun-Chen Tu, Paishun Ting, Karthikeyan Shanmugam, and Payel Das · 2018
Cited alongside, same era.
Datasheets for datasets
Timnit Gebru, Jamie Morgenstern, Briana Vecchione, Jennifer Wortman Vaughan, Hanna M. Wallach, Hal Daumé, and Kate Crawford · 2018
Cited alongside, same era.
Explaining explanations: An overview of interpretability of machine learning
Leilani H. Gilpin, David Bau, B. Yuan, A. Bajwa, M. Specter, and Lalana Kagal · 2018
Cited alongside, same era.
Training classifiers with natural language explanations
Braden Hancock, Paroma Varma, Stephanie Wang, Martin Bringmann, Percy Liang, and Christopher Ré · 2018
Cited alongside, same era.
Generating counterfactual explanations with natural language
Lisa Anne Hendricks, Ronghang Hu, Trevor Darrell, and Zeynep Akata · 2018
Cited alongside, same era.
An MTurk crisis? Shifts in data quality and the impact on study results
Michael Chmielewski and Sarah C Kucker · 2020
Later among the works it cites.
ERASER: A benchmark to evaluate rationalized NLP models
Jay DeYoung, Sarthak Jain, Nazneen Fatema Rajani, Eric Lehman, Caiming Xiong, Richard Socher, and Byron C. Wallace · 2020
Later among the works it cites.
Evidence inference 2.0: More data, better models
Jay DeYoung, Eric Lehman, Benjamin Nye, Iain Marshall, and Byron C. Wallace · 2020
Later among the works it cites.
The extraordinary failure of complement coercion crowdsourcing
Yanai Elazar, Victoria Basmov, Shauli Ravfogel, Yoav Goldberg, and Reut Tsarfaty · 2020
Later among the works it cites.
Social chemistry 101: Learning to reason about social and moral norms
Maxwell Forbes, Jena D. Hwang, Vered Shwartz, Maarten Sap, and Yejin Choi · 2020
Later among the works it cites.
Evaluating models’ local decision boundaries via contrast sets
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Robert R Hoffman, Shane T Mueller, Gary Klein, and Jordan Litman · 2018
Cited alongside, same era.
WorldTree: A corpus of explanation graphs for elementary science questions supporting multi-hop inference
Peter Jansen, Elizabeth Wainwright, Steven Marmorstein, and Clayton Morrison · 2018
Cited alongside, same era.
Looking beyond the surface: A challenge set for reading comprehension over multiple sentences
Daniel Khashabi, Snigdha Chaturvedi, Michael Roth, Shyam Upadhyay, and Dan Roth · 2018
Cited alongside, same era.
Textual Explanations for Self-Driving Vehicles
Jinkyu Kim, Anna Rohrbach, Trevor Darrell, John F. Canny, and Zeynep Akata · 2018
Cited alongside, same era.
VQA-E: Explaining, Elaborating, and Enhancing Your Answers for Visual Questions
Qing Li, Qingyi Tao, Shafiq R. Joty, Jianfei Cai, and Jiebo Luo · 2018
Cited alongside, same era.
The mythos of model interpretability
Zachary C Lipton · 2018
Cited alongside, same era.
Can a suit of armor conduct electricity? a new dataset for open book question answering
Todor Mihaylov, Peter Clark, Tushar Khot, and Ashish Sabharwal · 2018
Cited alongside, same era.
Matt Gardner, Yoav Artzi, Victoria Basmov, Jonathan Berant, Ben Bogin, Sihao Chen, Pradeep Dasigi, Dheeru Dua, Yanai Elazar, Ananth Gottumukkala, Nitish Gupta, Hannaneh Hajishirzi, Gabriel Ilharco, Daniel Khashabi, Kevin Lin, Jiangming Liu, Nelson F. Liu, Phoebe Mulcaire, Qiang Ning, Sameer Singh, Noah A. Smith, Sanjay Subramanian, Reut Tsarfaty, Eric Wallace, Ally Zhang, and Ben Zhou · 2020
Later among the works it cites.
R4C: A benchmark for evaluating RC systems to get the right answer for the right reason
Naoya Inoue, Pontus Stenetorp, and Kentaro Inui · 2020
Later among the works it cites.
Towards faithfully interpretable NLP systems: How should we define and evaluate faithfulness?
Alon Jacovi and Yoav Goldberg · 2020
Later among the works it cites.
Learning to explain: Datasets and models for identifying valid reasoning chains in multihop question-answering
Harsh Jhamtani and Peter Clark · 2020
Later among the works it cites.
Learning the difference that makes a difference with counterfactually-augmented data
Divyansh Kaushik, Eduard Hovy, and Zachary Lipton · 2020
Later among the works it cites.
Explainable AI for Cultural Minds
Hana Kopecká and Jose M Such · 2020
Later among the works it cites.
Explainable automated fact-checking for public health claims
Neema Kotonya and Francesca Toni · 2020
Later among the works it cites.
Explainable automated fact-checking: A survey
Neema Kotonya and Francesca Toni · 2020
Later among the works it cites.
NILE : Natural language inference with faithful natural language explanations
Sawan Kumar and Partha Talukdar · 2020
Later among the works it cites.
Annotator rationales for labeling tasks in crowdsourcing
Mucahid Kutlu, Tyler McDonnell, Matthew Lease, and Tamer Elsayed · 2020
Later among the works it cites.
What is more likely to happen next? video-and-language future event prediction
Jie Lei, Licheng Yu, Tamara Berg, and Mohit Bansal · 2020
Later among the works it cites.
Linguistically-informed transformations (LIT): A method for automatically generating contrast sets
Chuanrong Li, Lin Shengshuo, Zeyu Liu, Xinyi Wu, Xuhui Zhou, and Shane Steinert-Threlkeld · 2020
Later among the works it cites.
Natural language rationales with full-stack visual reasoning: From pixels to semantic frames to commonsense graphs
Ana Marasović, Chandra Bhagavatula, Jae sung Park, Ronan Le Bras, Noah A. Smith, and Yejin Choi · 2020
Later among the works it cites.
Details of data collection and crowd management for glucose (generalized and contextualized story explanations)
Lori Moon, Lauren Berkowitz, Jennifer Chu-Carroll, and Nasrin Mostafazadeh · 2020
Later among the works it cites.
GLUCOSE: GeneraLized and COntextualized story explanations
Nasrin Mostafazadeh, Aditya Kalyanpur, Lori Moon, David Buchanan, Lauren Berkowitz, Or Biran, and Jennifer Chu-Carroll · 2020
Later among the works it cites.
What can we learn from collective human opinions on natural language inference data?
Yixin Nie, Xiang Zhou, and Mohit Bansal · 2020
Later among the works it cites.
ToTTo: A controlled table-to-text generation dataset
Ankur Parikh, Xuezhi Wang, Sebastian Gehrmann, Manaal Faruqui, Bhuwan Dhingra, Diyi Yang, and Dipanjan Das · 2020
Later among the works it cites.
Achieving data excellence
Praveen Paritosh · 2020
Later among the works it cites.
QADiscourse - Discourse Relations as QA Pairs: Representation, Crowdsourcing and Baselines
Valentina Pyatkin, Ayal Klein, Reut Tsarfaty, and Ido Dagan · 2020
Later among the works it cites.
ESPRIT: Explaining solutions to physical reasoning tasks
Nazneen Fatema Rajani, Rui Zhang, Yi Chern Tan, Stephan Zheng, Jeremy Weiss, Aadit Vyas, Abhijit Gupta, Caiming Xiong, Richard Socher, and Dragomir Radev · 2020
Later among the works it cites.
Controlled crowdsourcing for high-quality QA-SRL annotation
Paul Roit, Ayal Klein, Daniela Stepanov, Jonathan Mamou, Julian Michael, Gabriel Stanovsky, Luke Zettlemoyer, and Ido Dagan · 2020
Later among the works it cites.
Human-data interaction in ai
Nithya Sambasivan · 2020
Later among the works it cites.
Social bias frames: Reasoning about social and power implications of language
Maarten Sap, Saadia Gabriel, Lianhui Qin, Dan Jurafsky, Noah A. Smith, and Yejin Choi · 2020
Later among the works it cites.
F1 is Not Enough! Models and Evaluation Towards User-Centered Explainable Question Answering
Hendrik Schuff, Heike Adel, and Ngoc Thang Vu · 2020
Later among the works it cites.
Explainability fact sheets: a framework for systematic assessment of explainable approaches
Kacper Sokol and Peter Flach · 2020
Later among the works it cites.
Fact or fiction: Verifying scientific claims
David Wadden, Shanchuan Lin, Kyle Lo, Lucy Lu Wang, Madeleine van Zuylen, Arman Cohan, and Hannaneh Hajishirzi · 2020
Later among the works it cites.
Investigating annotator bias with a graph-based approach
Maximilian Wich, Hala Al Kuwatly, and Georg Groh · 2020
Later among the works it cites.
WorldTree v2: A corpus of science-domain structured explanations and inference patterns supporting multi-hop inference
Zhengnan Xie, Sebastian Thiem, Jaycie Martin, Elizabeth Wainwright, Steven Marmorstein, and Peter Jansen · 2020
Later among the works it cites.
Generating plausible counterfactual explanations for deep transformers in financial text classification
Linyi Yang, Eoin Kenny, Tin Lok James Ng, Yi Yang, Barry Smyth, and Ruihai Dong · 2020
Later among the works it cites.
Teaching machine comprehension with compositional explanations
Qinyuan Ye, Xiao Huang, Elizabeth Boschee, and Xiang Ren · 2020
Later among the works it cites.
WinoWhy: A deep diagnosis of essential commonsense knowledge for answering Winograd schema challenge
Hongming Zhang, Xinran Zhao, and Yangqiu Song · 2020
Later among the works it cites.
Explanations for CommonsenseQA: New Dataset and Models
Shourya Aggarwal, Divyanshu Mandowara, Vishwajeet Agrawal, Dinesh Khandelwal, Parag Singla, and Dinesh Garg · 2021
Closest in time.
A survey on the explainability of supervised machine learning
Nadia Burkart and Marco F. Huber · 2021
Closest in time.
Paragraph-level rationale extraction through regularization: A case study on European court of human rights cases
Ilias Chalkidis, Manos Fergadiotis, Dimitrios Tsarapatsanis, Nikolaos Aletras, Ion Androutsopoulos, and Prodromos Malakasiotis · 2021
Closest in time.
Edited media understanding frames: Reasoning about the intent and implications of visual misinformation
Jeff Da, Maxwell Forbes, Rowan Zellers, Anthony Zheng, Jena D. Hwang, Antoine Bosselut, and Yejin Choi · 2021
Closest in time.
Did aristotle use a laptop? a question answering benchmark with implicit reasoning strategies
Mor Geva, Daniel Khashabi, Elad Segal, Tushar Khot, Dan Roth, and Jonathan Berant · 2021
Closest in time.
Does bert learn as humans perceive? understanding linguistic styles through lexica
Shirley Anugrah Hayati, Dongyeop Kang, and Lyle Ungar · 2021
Closest in time.
QED: A Framework and Dataset for Explanations in Question Answering
Matthew Lamm, Jennimaria Palomaki, Chris Alberti, Daniel Andor, Eunsol Choi, Livio Baldini Soares, and Michael Collins · 2021
Closest in time.
Explaining NLP models via minimal contrastive editing (MiCE)
Alexis Ross, Ana Marasović, and Matthew Peters · 2021
Closest in time.
" everyone wants to do the model work, not the data work": Data cascades in high-stakes ai
Nithya Sambasivan, Shivani Kapania, Hannah Highfill, Diana Akrong, Praveen Kumar Paritosh, and Lora Mois Aroyo · 2021
Closest in time.
Polyjuice: Generating counterfactuals for explaining, evaluating, and improving models
Tongshuang Wu, Marco Tulio Ribeiro, Jeffrey Heer, and Daniel Weld · 2021
Closest in time.
Do context-aware translation models pay the right attention?
Kayo Yin, Patrick Fernandes, Danish Pruthi, Aditi Chaudhary, André F. T. Martins, and Graham Neubig · 2021
Closest in time.
Hypothesis only baselines in natural language inference
Adam Poliak, Jason Naradowsky, Aparajita Haldar, Rachel Rudinger, and Benjamin Van Durme · 2023
Closest in time.