Fetching the paper…
Reading the bibliography…
When a model attribution technique highlights a particular part of the input, a user might understand this highlight as making a statement about counterfactuals (Miller, 2019): if that part of the input were to change, the model's prediction might change as well.
RoBERTa: A Robustly Optimized BERT Pretraining Approach
Y. Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, M. Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 1907
Earlier work this paper cites.
Deep inside convolutional networks: Visualising image classification models and saliency maps
Karen Simonyan, Andrea Vedaldi, and Andrew Zisserman. 2013 · 2013
Earlier work this paper cites.
Explaining prediction models and individual predictions with feature contributions
Erik Štrumbelj and Igor Kononenko. 2014 · 2014
Earlier work this paper cites.
Generating Visual Explanations
Lisa Anne Hendricks, Zeynep Akata, Marcus Rohrbach, Jeff Donahue, Bernt Schiele, and Trevor Darrell. 2016 · 2016
Earlier work this paper cites.
Rationalizing neural predictions
Tao Lei, Regina Barzilay, and Tommi Jaakkola. 2016 · 2016
Earlier work this paper cites.
SQuAD: 100,000+ questions for machine comprehension of text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016 · 2016
Earlier work this paper cites.
“Why should I trust you?” Explaining the predictions of any classifier
Marco Tulio Ribeiro, Sameer Singh, and Carlos Guestrin. 2016 · 2016
Earlier work this paper cites.
Network dissection: Quantifying interpretability of deep visual representations
David Bau, Bolei Zhou, Aditya Khosla, Aude Oliva, and Antonio Torralba. 2017 · 2017
Earlier work this paper cites.
Towards a rigorous science of interpretable machine learning
Finale Doshi-Velez and Been Kim. 2017 · 2017
Earlier work this paper cites.
Adversarial examples for evaluating reading comprehension systems
Robin Jia and Percy Liang. 2017 · 2017
Earlier work this paper cites.
A unified approach to interpreting model predictions
Scott Lundberg and Su-In Lee. 2017 · 2017
Earlier work this paper cites.
Axiomatic attribution for deep networks
Mukund Sundararajan, Ankur Taly, and Qiqi Yan. 2017 · 2017
Earlier work this paper cites.
Obfuscated gradients give a false sense of security: Circumventing defenses to adversarial examples
Anish Athalye, Nicholas Carlini, and David Wagner. 2018 · 2018
Earlier work this paper cites.
Verifiable reinforcement learning via policy extraction
Osbert Bastani, Yewen Pu, and Armando Solar-Lezama. 2018 · 2018
Earlier work this paper cites.
e-SNLI: Natural Language Inference with Natural Language Explanations
Oana-Maria Camburu, Tim Rocktäschel, Thomas Lukasiewicz, and Phil Blunsom. 2018 · 2018
Earlier work this paper cites.
Do explanations make VQA models more predictable to a human?
Arjun Chandrasekaran, Viraj Prabhu, Deshraj Yadav, Prithvijit Chattopadhyay, and Devi Parikh. 2018 · 2018
Earlier work this paper cites.
What you can cram into a single $&!#* vector: Probing sentence embeddings for linguistic properties
Alexis Conneau, German Kruszewski, Guillaume Lample, Loïc Barrault, and Marco Baroni. 2018 · 2018
Earlier work this paper cites.
The mythos of model interpretability: In machine learning, the concept of interpretability is both important and slippery
Zachary C Lipton. 2018 · 2018
Earlier work this paper cites.
Comparing automatic and human evaluation of local explanations for text classification
Dong Nguyen. 2018 · 2018
Earlier work this paper cites.
Anchors: High-precision model-agnostic explanations
Marco Tulio Ribeiro, Sameer Singh, and Carlos Guestrin. 2018 · 2018
Earlier work this paper cites.
Programmatically interpretable reinforcement learning
Abhinav Verma, Vijayaraghavan Murali, Rishabh Singh, Pushmeet Kohli, and Swarat Chaudhuri. 2018 · 2018
Cited alongside, same era.
HotpotQA: A dataset for diverse, explainable multi-hop question answering
Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William W. Cohen, Ruslan Salakhutdinov, and Christopher D. Manning. 2018 · 2018
Cited alongside, same era.
Understanding dataset design choices for multi-hop reasoning
Jifan Chen and Greg Durrett. 2019 · 2019
Cited alongside, same era.
BoolQ: Exploring the surprising difficulty of natural yes/no questions
Christopher Clark, Kenton Lee, Ming-Wei Chang, Tom Kwiatkowski, Michael Collins, and Kristina Toutanova. 2019 · 2019
Cited alongside, same era.
Towards a deep and unified understanding of deep neural models in NLP
Chaoyu Guan, Xiting Wang, Quanshi Zhang, Runjin Chen, Di He, and Xing Xie. 2019 · 2019
Cited alongside, same era.
Designing and interpreting probes with control tasks
How do decisions emerge across layers in neural models? interpretation with differentiable masking
Nicola De Cao, Michael Sejr Schlichtkrull, Wilker Aziz, and Ivan Titov. 2020 · 2020
Later among the works it cites.
Evaluating models’ local decision boundaries via contrast sets
Matt Gardner, Yoav Artzi, Victoria Basmov, Jonathan Berant, Ben Bogin, Sihao Chen, Pradeep Dasigi, Dheeru Dua, Yanai Elazar, Ananth Gottumukkala, Nitish Gupta, Hannaneh Hajishirzi, Gabriel Ilharco, Daniel Khashabi, Kevin Lin, Jiangming Liu, Nelson F. Liu, Phoebe Mulcaire, Qiang Ning, Sameer Singh, Noah A. Smith, Sanjay Subramanian, Reut Tsarfaty, Eric Wallace, Ally Zhang, and Ben Zhou. 2020 · 2020
Later among the works it cites.
A simple yet strong pipeline for HotpotQA
Dirk Groeneveld, Tushar Khot, Mausam, and Ashish Sabharwal. 2020 · 2020
Later among the works it cites.
Considering likelihood in NLP classification explanations with occlusion and language modeling
David Harbecke and Christoph Alt. 2020 · 2020
Later among the works it cites.
Evaluating explainable AI: Which algorithmic explanations help users predict model behavior?
Peter Hase and Mohit Bansal. 2020 · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
John Hewitt and Percy Liang. 2019 · 2019
Cited alongside, same era.
Avoiding reasoning shortcuts: Adversarial evaluation, training, and model development for multi-hop QA
Yichen Jiang and Mohit Bansal. 2019 · 2019
Cited alongside, same era.
Explanation in artificial intelligence: Insights from the social sciences
Tim Miller. 2019 · 2019
Cited alongside, same era.
Compositional questions do not necessitate multi-hop reasoning
Sewon Min, Eric Wallace, Sameer Singh, Matt Gardner, Hannaneh Hajishirzi, and Luke Zettlemoyer. 2019 · 2019
Cited alongside, same era.
Counterfactual story reasoning and generation
Lianhui Qin, Antoine Bosselut, Ari Holtzman, Chandra Bhagavatula, Elizabeth Clark, and Yejin Choi. 2019 · 2019
Cited alongside, same era.
Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead
Cynthia Rudin. 2019 · 2019
Cited alongside, same era.
Is attention interpretable?
Sofia Serrano and Noah A. Smith. 2019 · 2019
Cited alongside, same era.
Later among the works it cites.
Towards faithfully interpretable NLP systems: How should we define and evaluate faithfulness?
Alon Jacovi and Yoav Goldberg. 2020 · 2020
Later among the works it cites.
Towards hierarchical importance attribution: Explaining compositional semantics for neural sequence models
Xisen Jin, Zhongyu Wei, Junyi Du, Xiangyang Xue, and Xiang Ren. 2020 · 2020
Later among the works it cites.
Selective question answering under domain shift
Amita Kamath, Robin Jia, and Percy Liang. 2020 · 2020
Later among the works it cites.
Learning the difference that makes a difference with counterfactually-augmented data
Divyansh Kaushik, Eduard Hovy, and Zachary Lipton. 2020 · 2020
Later among the works it cites.
Compositional explanations of neurons
Jesse Mu and Jacob Andreas. 2020 · 2020
Later among the works it cites.
Beyond accuracy: Behavioral testing of NLP models with CheckList
Marco Tulio Ribeiro, Tongshuang Wu, Carlos Guestrin, and Sameer Singh. 2020 · 2020
Later among the works it cites.
Obtaining faithful interpretations from compositional neural networks
Sanjay Subramanian, Ben Bogin, Nitish Gupta, Tomer Wolfson, Sameer Singh, Jonathan Berant, and Matt Gardner. 2020 · 2020
Later among the works it cites.
How does this interaction affect me? interpretable attribution for feature interactions
Michael Tsang, Sirisha Rambhatla, and Yan Liu. 2020 · 2020
Later among the works it cites.
Select, answer and explain: Interpretable multi-hop reading comprehension over multiple documents
Ming Tu, Kevin Huang, Guangtao Wang, Jing Huang, Xiaodong He, and Bowen Zhou. 2020 · 2020
Later among the works it cites.
Information-theoretic probing with minimum description length
Elena Voita and Ivan Titov. 2020 · 2020
Later among the works it cites.
Transformers: State-of-the-art natural language processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander M. Rush. 2020 · 2020
Later among the works it cites.
Self-attention attribution: Interpreting information interactions inside transformer
Yaru Hao, Li Dong, Furu Wei, and Ke Xu. 2021 · 2021
Closest in time.
Aligning Faithful Interpretations with their Social Attribution
Alon Jacovi and Yoav Goldberg. 2021 · 2021
Closest in time.
Discretized Integrated Gradients for Explaining Language Models
Soumya Sanyal and Xiang Ren. 2021 · 2021
Closest in time.
Measuring association between labels and free-text rationales
Sarah Wiegreffe, Ana Marasović, and Noah A. Smith. 2021 · 2021
Closest in time.