Fetching the paper…
Reading the bibliography…
End-to-end neural Natural Language Processing (NLP) models are notoriously difficult to understand.
Visualizing Attention in Transformer-Based Language Representation Models
Vig, Jesse. 2019 · 1904
Earlier work this paper cites.
METGEN: A module-based entailment tree generation framework for answer explanation
Hong, Ruixin, Hongming Zhang, Xintong Yu, and Changshui Zhang. 2022 · 1905
Earlier work this paper cites.
RoBERTa: A Robustly Optimized BERT Pretraining Approach
Liu, Yinhan, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 1907
Earlier work this paper cites.
Attention Interpretability Across NLP Tasks
Vashishth, Shikhar, Shyam Upadhyay, Gaurav Singh Tomar, and Manaal Faruqui. 2019 · 1909
Earlier work this paper cites.
17. A Value for n-Person Games
Shapley, L. S. 1953 · 1953
Earlier work this paper cites.
Aspects of the theory of syntax
Chomsky, Noam. 1965 · 1965
Earlier work this paper cites.
Harvey Friedman’s Research on the Foundations of Mathematics
Harrington, L. A., M. D. Morley, A. Šcedrov, and S. G. Simpson. 1985 · 1985
Earlier work this paper cites.
Counterfactual thinking: A critical overview
Roese, Neal J. and James M. Olson. 1995 · 1995
Earlier work this paper cites.
CLEVR: A diagnostic dataset for compositional language and elementary visual reasoning
Johnson, Justin, Bharath Hariharan, Laurens van der Maaten, Li Fei-Fei, C. Lawrence Zitnick, and Ross B. Girshick. 2017 · 1997
Earlier work this paper cites.
Case-based explanation of non-case-based learning methods
Caruana, R., H. Kangarloo, J. D. Dionisio, U. Sinha, and D. Johnson. 1999 · 1999
Earlier work this paper cites.
The Estimation of Causal Effects from Observational Data
Winship, Christopher and Stephen L. Morgan. 1999 · 1999
Earlier work this paper cites.
WT5?! Training Text-to-Text Models to Explain their Predictions
Narang, Sharan, Colin Raffel, Katherine Lee, Adam Roberts, Noah Fiedel, and Karishma Malkan. 2020 · 2004
Earlier work this paper cites.
Causes and Explanations: A Structural-Model Approach. Part I: Causes
Halpern, Joseph Y. and Judea Pearl. 2005 · 2005
Earlier work this paper cites.
Using “Annotator Rationales” to Improve Machine Learning for Text Categorization
Zaidan, Omar, Jason Eisner, and Christine Piatko. 2007 · 2007
Earlier work this paper cites.
Captum: A unified and generic model interpretability library for PyTorch
Kokhlikyan, Narine, Vivek Miglani, Miguel Martin, Edward Wang, Bilal Alsallakh, Jonathan Reynolds, Alexander Melnikov, Natalia Kliushkina, Carlos Araya, Siqi Yan, and Orion Reblitz-Richardson. 2020 · 2009
Earlier work this paper cites.
How to Explain Individual Classification Decisions
Baehrens, David, Timon Schroeter, Stefan Harmeling, Motoaki Kawanabe, and Katja Hansen. 2010 · 2010
Earlier work this paper cites.
The Winograd Schema Challenge
Levesque, Hector, Ernest Davis, and Leora Morgenstern. 2012 · 2012
Earlier work this paper cites.
Jointly learning to parse and perceive: Connecting natural language to the physical world
Krishnamurthy, Jayant and Thomas Kollar. 2013 · 2013
Earlier work this paper cites.
Deep inside convolutional networks: Visualising image classification models and saliency maps
Simonyan, Karen, Andrea Vedaldi, and Andrew Zisserman. 2014 · 2014
Earlier work this paper cites.
Visualizing and Understanding Convolutional Networks
Zeiler, Matthew D. and Rob Fergus. 2014 · 2014
Earlier work this paper cites.
VQA: visual question answering
Antol, Stanislaw, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C. Lawrence Zitnick, and Devi Parikh. 2015 · 2015
Earlier work this paper cites.
On Pixel-Wise Explanations for Non-Linear Classifier Decisions by Layer-Wise Relevance Propagation
Bach, Sebastian, Alexander Binder, Grégoire Montavon, Frederick Klauschen, Klaus-Robert Müller, and Wojciech Samek. 2015 · 2015
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Bahdanau, Dzmitry, Kyunghyun Cho, and Yoshua Bengio. 2015 · 2015
Earlier work this paper cites.
A large annotated corpus for learning natural language inference
Bowman, Samuel R., Gabor Angeli, Christopher Potts, and Christopher D. Manning. 2015 · 2015
Earlier work this paper cites.
Extraction of Salient Sentences from Labelled Documents
Denil, Misha, Alban Demiraj, and Nando de Freitas. 2015 · 2015
Earlier work this paper cites.
Visualizing and Understanding Recurrent Networks
Karpathy, Andrej, Justin Johnson, and Li Fei-Fei. 2015 · 2015
Earlier work this paper cites.
Evaluating the visualization of what a Deep Neural Network has learned
Samek, Wojciech, Alexander Binder, Grégoire Montavon, Sebastian Bach, and Klaus-Robert Müller. 2015 · 2015
Earlier work this paper cites.
Striving for Simplicity: The All Convolutional Net
Springenberg, J., Alexey Dosovitskiy, Thomas Brox, and M. Riedmiller. 2015 · 2015
Earlier work this paper cites.
Learning to compose neural networks for question answering
Andreas, Jacob, Marcus Rohrbach, Trevor Darrell, and Dan Klein. 2016a · 2016
Earlier work this paper cites.
Neural module networks
Andreas, Jacob, Marcus Rohrbach, Trevor Darrell, and Dan Klein. 2016b · 2016
Earlier work this paper cites.
Explaining predictions of non-linear classifiers in NLP
Arras, Leila, Franziska Horn, Grégoire Montavon, Klaus-Robert Müller, and Wojciech Samek. 2016 · 2016
Earlier work this paper cites.
Generating Visual Explanations
Hendricks, Lisa Anne, Zeynep Akata, Marcus Rohrbach, Jeff Donahue, Bernt Schiele, and Trevor Darrell. 2016 · 2016
Earlier work this paper cites.
Rationalizing neural predictions
Lei, Tao, Regina Barzilay, and Tommi Jaakkola. 2016 · 2016
Earlier work this paper cites.
Visualizing and understanding neural models in NLP
Li, Jiwei, Xinlei Chen, Eduard Hovy, and Dan Jurafsky. 2016 · 2016
Earlier work this paper cites.
Understanding Neural Networks through Representation Erasure
Li, Jiwei, Will Monroe, and Dan Jurafsky. 2016 · 2016
Earlier work this paper cites.
The Mythos of Model Interpretability
Lipton, Zachary C. 2016 · 2016
Earlier work this paper cites.
From softmax to sparsemax: A sparse model of attention and multi-label classification
Martins, André F. T. and Ramón Fernandez Astudillo. 2016 · 2016
Earlier work this paper cites.
Analyzing linguistic knowledge in sequential model of sentence
Qian, Peng, Xipeng Qiu, and Xuanjing Huang. 2016 · 2016
Earlier work this paper cites.
"why should I trust you?": Explaining the predictions of any classifier
Ribeiro, Marco Túlio, Sameer Singh, and Carlos Guestrin. 2016 · 2016
Earlier work this paper cites.
Diagnostic classifiers: revealing how neural networks process hierarchical structure
Veldhoen, Sara, Dieuwke Hupkes, and Willem Zuidema. 2016 · 2016
Earlier work this paper cites.
Fine-grained analysis of sentence embeddings using auxiliary prediction tasks
Adi, Yossi, Einat Kermany, Yonatan Belinkov, Ofer Lavi, and Yoav Goldberg. 2017 · 2017
Earlier work this paper cites.
Explaining recurrent neural network predictions in sentiment analysis
Arras, Leila, Grégoire Montavon, Klaus-Robert Müller, and Wojciech Samek. 2017 · 2017
Earlier work this paper cites.
Towards A Rigorous Science of Interpretable Machine Learning
Doshi-Velez, Finale and Been Kim. 2017 · 2017
Earlier work this paper cites.
The Promise and Peril of Human Evaluation for Model Interpretability
Herman, Bernease. 2017 · 2017
Earlier work this paper cites.
Learning to reason: End-to-end module networks for visual question answering
Hu, Ronghang, Jacob Andreas, Marcus Rohrbach, Trevor Darrell, and Kate Saenko. 2017 · 2017
Earlier work this paper cites.
Representation of linguistic form and function in recurrent neural networks
Kádár, Ákos, Grzegorz Chrupała, and Afra Alishahi. 2017 · 2017
Earlier work this paper cites.
Understanding black-box predictions via influence functions
Koh, Pang Wei and Percy Liang. 2017 · 2017
Earlier work this paper cites.
Program induction by rationale generation: Learning to solve and explain algebraic word problems
Ling, Wang, Dani Yogatama, Chris Dyer, and Phil Blunsom. 2017 · 2017
Earlier work this paper cites.
A unified approach to interpreting model predictions
Lundberg, Scott M. and Su-In Lee. 2017 · 2017
Earlier work this paper cites.
Explanation in Artificial Intelligence: Insights from the Social Sciences
Miller, Tim. 2017 · 2017
Earlier work this paper cites.
Understanding Hidden Memories of Recurrent Neural Networks
Ming, Yao, Shaozu Cao, Ruixiang Zhang, Zhen Li, Yuanzhe Chen, Yangqiu Song, and Huamin Qu. 2017 · 2017
Earlier work this paper cites.
Explaining nonlinear classification decisions with deep Taylor decomposition
Montavon, Grégoire, Sebastian Lapuschkin, Alexander Binder, Wojciech Samek, and Klaus-Robert Müller. 2017 · 2017
Earlier work this paper cites.
Learning important features through propagating activation differences
Shrikumar, Avanti, Peyton Greenside, and Anshul Kundaje. 2017 · 2017
Earlier work this paper cites.
Not Just a Black Box: Learning Important Features Through Propagating Activation Differences
Shrikumar, Avanti, Peyton Greenside, Anna Shcherbina, and Anshul Kundaje. 2017 · 2017
Earlier work this paper cites.
SmoothGrad: removing noise by adding noise
Smilkov, Daniel, Nikhil Thorat, Been Kim, Fernanda Viégas, and Martin Wattenberg. 2017 · 2017
Earlier work this paper cites.
Axiomatic attribution for deep networks
Sundararajan, Mukund, Ankur Taly, and Qiqi Yan. 2017 · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, Ashish, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
An interpretable knowledge transfer model for knowledge base completion
Xie, Qizhe, Xuezhe Ma, Zihang Dai, and Eduard Hovy. 2017 · 2017
Earlier work this paper cites.
Sanity checks for saliency maps
Adebayo, Julius, Justin Gilmer, Michael Muelly, Ian J. Goodfellow, Moritz Hardt, and Been Kim. 2018 · 2018
Earlier work this paper cites.
On the Robustness of Interpretability Methods
Alvarez-Melis, David and Tommi S. Jaakkola. 2018 · 2018
Earlier work this paper cites.
Towards robust interpretability with self-explaining neural networks
Alvarez-Melis, David and Tommi S. Jaakkola. 2018 · 2018
Earlier work this paper cites.
Natural Language Multitasking: Analyzing and Improving Syntactic Saliency of Hidden Representations
Brunner, Gino, Yuyi Wang, Roger Wattenhofer, and Michael Weigelt. 2018 · 2018
Earlier work this paper cites.
e-snli: Natural language inference with natural language explanations
Camburu, Oana-Maria, Tim Rocktäschel, Thomas Lukasiewicz, and Phil Blunsom. 2018 · 2018
Earlier work this paper cites.
RNNbow: Visualizing Learning Via Backpropagation Gradients in RNNs
Cashman, Dylan, Geneviève Patterson, Abigail Mosca, Nathan Watts, Shannon Robinson, and Remco Chang. 2018 · 2018
Earlier work this paper cites.
Learning to explain: An information-theoretic perspective on model interpretation
Chen, Jianbo, Le Song, Martin J. Wainwright, and Michael I. Jordan. 2018 · 2018
Earlier work this paper cites.
Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
Clark, Peter, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord. 2018 · 2018
Earlier work this paper cites.
What you can cram into a single $&!#* vector: Probing sentence embeddings for linguistic properties
Conneau, Alexis, German Kruszewski, Guillaume Lample, Loïc Barrault, and Marco Baroni. 2018 · 2018
Earlier work this paper cites.
HotFlip: White-box adversarial examples for text classification
Ebrahimi, Javid, Anyi Rao, Daniel Lowd, and Dejing Dou. 2018 · 2018
Earlier work this paper cites.
Pathologies of neural models make interpretations difficult
Feng, Shi, Eric Wallace, Alvin Grissom II, Mohit Iyyer, Pedro Rodriguez, and Jordan Boyd-Graber. 2018 · 2018
Earlier work this paper cites.
Interpreting word-level hidden state behaviour of character-level LSTM language models
Hiebert, Avery, Cole Peterson, Alona Fyshe, and Nishant Mehta. 2018 · 2018
Earlier work this paper cites.
How much reading does reading comprehension require? a critical investigation of popular benchmarks
Kaushik, Divyansh and Zachary C. Lipton. 2018 · 2018
Earlier work this paper cites.
Interpretability beyond feature attribution: Quantitative testing with concept activation vectors (TCAV)
Kim, Been, Martin Wattenberg, Justin Gilmer, Carrie J. Cai, James Wexler, Fernanda B. Viégas, and Rory Sayres. 2018 · 2018
Earlier work this paper cites.
Learning how to explain neural networks: Patternnet and patternattribution
Kindermans, Pieter-Jan, Kristof T. Schütt, Maximilian Alber, Klaus-Robert Müller, Dumitru Erhan, Been Kim, and Sven Dähne. 2018 · 2018
Earlier work this paper cites.
Defining Locality for Surrogates in Post-hoc Interpretablity
Laugel, Thibault, Xavier Renard, Marie-Jeanne Lesot, Christophe Marsala, and Marcin Detyniecki. 2018 · 2018
Earlier work this paper cites.
Explainable prediction of medical codes from clinical text
Mullenbach, James, Sarah Wiegreffe, Jon Duke, Jimeng Sun, and Jacob Eisenstein. 2018 · 2018
Earlier work this paper cites.
A theoretical explanation for perplexing behaviors of backpropagation-based visualizations
Nie, Weili, Yang Zhang, and Ankit Patel. 2018 · 2018
Earlier work this paper cites.
Multimodal explanations: Justifying decisions and pointing to the evidence
Park, Dong Huk, Lisa Anne Hendricks, Zeynep Akata, Anna Rohrbach, Bernt Schiele, Trevor Darrell, and Marcus Rohrbach. 2018 · 2018
Earlier work this paper cites.
Interpretable textual neuron representations for NLP
Poerner, Nina, Benjamin Roth, and Hinrich Schütze. 2018 · 2018
Earlier work this paper cites.
Evaluating neural network explanation methods using hybrid documents and morphosyntactic agreement
Poerner, Nina, Hinrich Schütze, and Benjamin Roth. 2018 · 2018
Earlier work this paper cites.
Hypothesis only baselines in natural language inference
Poliak, Adam, Jason Naradowsky, Aparajita Haldar, Rachel Rudinger, and Benjamin Van Durme. 2018 · 2018
Earlier work this paper cites.
Anchors: High-precision model-agnostic explanations
Ribeiro, Marco Túlio, Sameer Singh, and Carlos Guestrin. 2018 · 2018
Cited alongside, same era.
Bridging CNNs, RNNs, and weighted finite-state machines
Schwartz, Roy, Sam Thomson, and Noah A. Smith. 2018 · 2018
Cited alongside, same era.
LSTMVis: A Tool for Visual Analysis of Hidden State Dynamics in Recurrent Neural Networks
Strobelt, Hendrik, Sebastian Gehrmann, Hanspeter Pfister, and Alexander M. Rush. 2018 · 2018
Cited alongside, same era.
Patient representation learning and interpretable evaluation using clinical notes
Sushil, Madhumita, Simon Šuster, Kim Luyckx, and Walter Daelemans. 2018 · 2018
Cited alongside, same era.
Interpreting neural networks with nearest neighbors
Wallace, Eric, Shi Feng, and Jordan Boyd-Graber. 2018 · 2018
Cited alongside, same era.
HotpotQA: A dataset for diverse, explainable multi-hop question answering
Gradient-based analysis of NLP models is manipulable
Wang, Junlin, Jens Tuyls, Eric Wallace, and Sameer Singh. 2020 · 2020
Later among the works it cites.
The What-If Tool: Interactive Probing of Machine Learning Models
Wexler, James, Mahima Pushkarna, Tolga Bolukbasi, Martin Wattenberg, Fernanda Viégas, and Jimbo Wilson. 2020 · 2020
Later among the works it cites.
On completeness-aware concept-based explanations in deep neural networks
Yeh, Chih-Kuan, Been Kim, Sercan Ömer Arik, Chun-Liang Li, Tomas Pfister, and Pradeep Ravikumar. 2020 · 2020
Later among the works it cites.
Does the Whole Exceed its Parts? The Effect of AI Explanations on Complementary Team Performance
Bansal, Gagan, Tongshuang Wu, Joyce Zhou, Raymond Fok, Besmira Nushi, Ece Kamar, Marco Tulio Ribeiro, and Daniel Weld. 2021 · 2021
Later among the works it cites.
Influence functions in deep learning are fragile
Basu, Samyadeep, Phillip Pope, and Soheil Feizi. 2021 · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yang, Zhilin, Peng Qi, Saizheng Zhang, Yoshua Bengio, William Cohen, Ruslan Salakhutdinov, and Christopher D. Manning. 2018 · 2018
Cited alongside, same era.
Neural-symbolic VQA: disentangling reasoning from vision and language understanding
Yi, Kexin, Jiajun Wu, Chuang Gan, Antonio Torralba, Pushmeet Kohli, and Josh Tenenbaum. 2018 · 2018
Cited alongside, same era.
Interpretable neural predictions with differentiable binary variables
Bastings, Jasmijn, Wilker Aziz, and Ivan Titov. 2019 · 2019
Cited alongside, same era.
Identifying and controlling important neurons in neural machine translation
Bau, Anthony, Yonatan Belinkov, Hassan Sajjad, Nadir Durrani, Fahim Dalvi, and James R. Glass. 2019 · 2019
Cited alongside, same era.
Analysis methods in neural language processing: A survey
Belinkov, Yonatan and James Glass. 2019 · 2019
Cited alongside, same era.
A general-purpose algorithm for constrained sequential inference
Deutsch, Daniel, Shyam Upadhyay, and Dan Roth. 2019 · 2019
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
Devlin, Jacob, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Cited alongside, same era.
Latent compositional representations improve systematic generalization in grounded question answering
Bogin, Ben, Sanjay Subramanian, Matt Gardner, and Jonathan Berant. 2021 · 2021
Later among the works it cites.
On the Opportunities and Risks of Foundation Models
Bommasani, Rishi, Drew A. Hudson, Ehsan Adeli, Russ Altman, Simran Arora, Sydney von Arx, Michael S. Bernstein, Jeannette Bohg, Antoine Bosselut, Emma Brunskill, Erik Brynjolfsson, Shyamal Buch, Dallas Card, Rodrigo Castellon, Niladri Chatterji, Annie Chen, Kathleen Creel, Jared Quincy Davis, Dora Demszky, Chris Donahue, Moussa Doumbouya, Esin Durmus, Stefano Ermon, John Etchemendy, Kawin Ethayarajh, Li Fei-Fei, Chelsea Finn, Trevor Gale, Lauren Gillespie, Karan Goel, Noah Goodman, Shelby Grossman, Neel Guha, Tatsunori Hashimoto, Peter Henderson, John Hewitt, Daniel E. Ho, Jenny Hong, Kyle Hsu, Jing Huang, Thomas Icard, Saahil Jain, Dan Jurafsky, Pratyusha Kalluri, Siddharth Karamcheti, Geoff Keeling, Fereshte Khani, Omar Khattab, Pang Wei Koh, Mark Krass, Ranjay Krishna, Rohith Kuditipudi, Ananya Kumar, Faisal Ladhak, Mina Lee, Tony Lee, Jure Leskovec, Isabelle Levent, Xiang Lisa Li, Xuechen Li, Tengyu Ma, Ali Malik, Christopher D. Manning, Suvir Mirchandani, Eric Mitchell, Zanele Munyikwa, Suraj Nair, Avanika Narayan, Deepak Narayanan, Ben Newman, Allen Nie, Juan Carlos Niebles, Hamed Nilforoshan, Julian Nyarko, Giray Ogut, Laurel Orr, Isabel Papadimitriou, Joon Sung Park, Chris Piech, Eva Portelance, Christopher Potts, Aditi Raghunathan, Rob Reich, Hongyu Ren, Frieda Rong, Yusuf Roohani, Camilo Ruiz, Jack Ryan, Christopher Ré, Dorsa Sadigh, Shiori Sagawa, Keshav Santhanam, Andy Shih, Krishnan Srinivasan, Alex Tamkin, Rohan Taori, Armin W. Thomas, Florian Tramèr, Rose E. Wang, William Wang, Bohan Wu, Jiajun Wu, Yuhuai Wu, Sang Michael Xie, Michihiro Yasunaga, Jiaxuan You, Matei Zaharia, Michael Zhang, Tianyi Zhang, Xikun Zhang, Yuhui Zhang, Lucia Zheng, Kaitlyn Zhou, and Percy Liang. 2021 · 2021
Later among the works it cites.
Transformer interpretability beyond attention visualization
Chefer, Hila, Shir Gur, and Lior Wolf. 2021 · 2021
Later among the works it cites.
Stepmothers are mean and academics are pretentious: What do pretrained language models learn about you?
Choenni, Rochelle, Ekaterina Shutova, and Robert van Rooij. 2021 · 2021
Later among the works it cites.
A study of automatic metrics for the evaluation of natural language explanations
Clinciu, Miruna-Adriana, Arash Eshghi, and Helen Hastie. 2021 · 2021
Later among the works it cites.
Training Verifiers to Solve Math Word Problems
Cobbe, Karl, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman. 2021 · 2021
Later among the works it cites.
Explaining answers with entailment trees
Dalvi, Bhavana, Peter Jansen, Oyvind Tafjord, Zhengnan Xie, Hannah Smith, Leighanna Pipatanangkura, and Peter Clark. 2021 · 2021
Later among the works it cites.
Evaluating saliency methods for neural language models
Ding, Shuoyang and Philipp Koehn. 2021 · 2021
Later among the works it cites.
Amnesic probing: Behavioral explanation with amnesic counterfactuals
Elazar, Yanai, Shauli Ravfogel, Alon Jacovi, and Yoav Goldberg. 2021 · 2021
Later among the works it cites.
Attention flows are shapley value explanations
Ethayarajh, Kawin and Dan Jurafsky. 2021 · 2021
Later among the works it cites.
CausaLM: Causal model explanation through counterfactual language models
Feder, Amir, Nadav Oved, Uri Shalit, and Roi Reichart. 2021 · 2021
Later among the works it cites.
Causal analysis of syntactic agreement mechanisms in neural language models
Finlayson, Matthew, Aaron Mueller, Sebastian Gehrmann, Stuart Shieber, Tal Linzen, and Yonatan Belinkov. 2021 · 2021
Later among the works it cites.
Self-attention attribution: Interpreting information interactions inside transformer
Hao, Yaru, Li Dong, Furu Wei, and Ke Xu. 2021 · 2021
Later among the works it cites.
Measuring massive multitask language understanding
Hendrycks, Dan, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. 2021 · 2021
Later among the works it cites.
Aligning faithful interpretations with their social attribution
Jacovi, Alon and Yoav Goldberg. 2021 · 2021
Later among the works it cites.
Contrastive explanations for model interpretability
Jacovi, Alon, Swabha Swayamdipta, Shauli Ravfogel, Yanai Elazar, Yejin Choi, and Yoav Goldberg. 2021 · 2021
Later among the works it cites.
Explaining explanations: Axiomatic feature interactions for deep networks
Janizek, Joseph D., Pascal Sturmfels, and Su-In Lee. 2021 · 2021
Later among the works it cites.
Putting words in BERT’s mouth: Navigating contextualized vector spaces with pseudowords
Karidi, Taelin, Yichu Zhou, Nathan Schneider, Omri Abend, and Vivek Srikumar. 2021 · 2021
Later among the works it cites.
BeliefBank: Adding memory to a pre-trained language model for a systematic notion of belief
Kassner, Nora, Oyvind Tafjord, Hinrich Schütze, and Peter Clark. 2021 · 2021
Later among the works it cites.
Influence patterns for explaining information flow in BERT
Lu, Kaiji, Zifan Wang, Piotr Mardziel, and Anupam Datta. 2021 · 2021
Later among the works it cites.
Improving Neural Model Performance through Natural Language Feedback on Their Explanations
Madaan, Aman, Niket Tandon, Dheeraj Rajagopal, Yiming Yang, Peter Clark, Keisuke Sakaguchi, and Ed Hovy. 2021 · 2021
Later among the works it cites.
Show Your Work: Scratchpads for Intermediate Computation with Language Models
Nye, Maxwell, Anders Johan Andreassen, Guy Gur-Ari, Henryk Michalewski, Jacob Austin, David Bieber, David Dohan, Aitor Lewkowycz, Maarten Bosma, David Luan, Charles Sutton, and Augustus Odena. 2021 · 2021
Later among the works it cites.
Telling BERT’s full story: from local attention to global aggregation
Pascual, Damian, Gino Brunner, and Roger Wattenhofer. 2021 · 2021
Later among the works it cites.
An empirical comparison of instance attribution methods for NLP
Pezeshkpour, Pouya, Sarthak Jain, Byron Wallace, and Sameer Singh. 2021 · 2021
Later among the works it cites.
SELFEXPLAIN: A self-explaining architecture for neural text classifiers
Rajagopal, Dheeraj, Vidhisha Balachandran, Eduard H Hovy, and Yulia Tsvetkov. 2021 · 2021
Later among the works it cites.
Counterfactual interventions reveal the causal effect of relative clause representations on agreement prediction
Ravfogel, Shauli, Grusha Prasad, Tal Linzen, and Yoav Goldberg. 2021 · 2021
Later among the works it cites.
Probing the probing paradigm: Does probing accuracy entail task relevance?
Ravichander, Abhilasha, Yonatan Belinkov, and Eduard Hovy. 2021 · 2021
Later among the works it cites.
ProofWriter: Generating implications, proofs, and abductive statements over natural language
Tafjord, Oyvind, Bhavana Dalvi, and Peter Clark. 2021 · 2021
Later among the works it cites.
What if this modified that? syntactic interventions with counterfactual embeddings
Tucker, Mycal, Peng Qian, and Roger Levy. 2021 · 2021
Later among the works it cites.
Measuring association between labels and free-text rationales
Wiegreffe, Sarah, Ana Marasović, and Noah A. Smith. 2021 · 2021
Later among the works it cites.
Polyjuice: Generating counterfactuals for explaining, evaluating, and improving models
Wu, Tongshuang, Marco Tulio Ribeiro, Jeffrey Heer, and Daniel Weld. 2021 · 2021
Later among the works it cites.
Connecting attributions and QA model behavior on realistic counterfactuals
Ye, Xi, Rohan Nair, and Greg Durrett. 2021 · 2021
Later among the works it cites.
CEBaB: Estimating the Causal Effects of Real-World Concepts on NLP Model Behavior
Abraham, Eldar David, Karel D’Oosterlinck, Amir Feder, Yair Ori Gat, Atticus Geiger, Christopher Potts, Roi Reichart, and Zhengxuan Wu. 2022 · 2022
Closest in time.
Naturalistic Causal Probing for Morpho-Syntax
Amini, Afra, Tiago Pimentel, Clara Meister, and Ryan Cotterell. 2022 · 2022
Closest in time.
“will you find these shortcuts?” a protocol for evaluating the faithfulness of input salience methods for text classification
Bastings, Jasmijn, Sebastian Ebert, Polina Zablotskaia, Anders Sandholm, and Katja Filippova. 2022 · 2022
Closest in time.
DoCoGen: Domain counterfactual generation for low resource domain adaptation
Calderon, Nitay, Eyal Ben-David, Amir Feder, and Roi Reichart. 2022 · 2022
Closest in time.
Brains and algorithms partially converge in natural language processing
Caucheteux, Charlotte and Jean-Rémi King. 2022 · 2022
Closest in time.
A comparative study of faithfulness metrics for model interpretability methods
Chan, Chun Sik, Huanqi Kong, and Liang Guanqing. 2022 · 2022
Closest in time.
Faithful Reasoning Using Large Language Models
Creswell, Antonia and Murray Shanahan. 2022 · 2022
Closest in time.
Discovering latent concepts learned in BERT
Dalvi, Fahim, Abdul Rafae Khan, Firoj Alam, Nadir Durrani, Jia Xu, and Hassan Sajjad. 2022 · 2022
Closest in time.
Sparse interventions in language models with differentiable masking
De Cao, Nicola, Leon Schmid, Dieuwke Hupkes, and Ivan Titov. 2022 · 2022
Closest in time.
Do transformer models show similar attention patterns to task-specific human gaze?
Eberle, Oliver, Stephanie Brandl, Jonas Pilot, and Anders Søgaard. 2022 · 2022
Closest in time.
Causal inference in natural language processing: Estimation, prediction, interpretation and beyond
Feder, Amir, Katherine A. Keith, Emaad Manzoor, Reid Pryzant, Dhanya Sridhar, Zach Wood-Doughty, Jacob Eisenstein, Justin Grimmer, Roi Reichart, Margaret E. Roberts, Brandon M. Stewart, Victor Veitch, and Diyi Yang. 2022 · 2022
Closest in time.
PAL: Program-aided Language Models
Gao, Luyu, Aman Madaan, Shuyan Zhou, Uri Alon, Pengfei Liu, Yiming Yang, Jamie Callan, and Graham Neubig. 2022 · 2022
Closest in time.
Better hit the nail on the head than beat around the bush: Removing protected attributes with a single projection
Haghighatkhah, Pantea, Antske Fokkens, Pia Sommerauer, Bettina Speckmann, and Kevin Verbeek. 2022 · 2022
Closest in time.
When can models learn from explanations? a formal framework for understanding the roles of explanation data
Hase, Peter and Mohit Bansal. 2022 · 2022
Closest in time.
Logic traps in evaluating attribution scores
Ju, Yiming, Yuanzhe Zhang, Zhao Yang, Zhongtao Jiang, Kang Liu, and Jun Zhao. 2022 · 2022
Closest in time.
Maieutic prompting: Logically consistent reasoning with recursive explanations
Jung, Jaehun, Lianhui Qin, Sean Welleck, Faeze Brahman, Chandra Bhagavatula, Ronan Le Bras, and Yejin Choi. 2022 · 2022
Closest in time.
Large Language Models are Zero-Shot Reasoners
Kojima, Takeshi, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa. 2022 · 2022
Closest in time.
Probing Classifiers are Unreliable for Concept Removal and Detection
Kumar, Abhinav, Chenhao Tan, and Amit Sharma. 2022 · 2022
Closest in time.
Can language models learn from explanations in context?
Lampinen, Andrew, Ishita Dasgupta, Stephanie Chan, Kory Mathewson, Mh Tessler, Antonia Creswell, James McClelland, Jane Wang, and Felix Hill. 2022 · 2022
Closest in time.
Solving Quantitative Reasoning Problems with Language Models
Lewkowycz, Aitor, Anders Andreassen, David Dohan, Ethan Dyer, Henryk Michalewski, Vinay Ramasesh, Ambrose Slone, Cem Anil, Imanol Schlag, Theo Gutman-Solo, Yuhuai Wu, Behnam Neyshabur, Guy Gur-Ari, and Vedant Misra. 2022 · 2022
Closest in time.
On the Advance of Making Language Models Better Reasoners
Li, Yifei, Zeqi Lin, Shizhuo Zhang, Qiang Fu, Bei Chen, Jian-Guang Lou, and Weizhu Chen. 2022 · 2022
Closest in time.
Rethinking attention-model explainability through faithfulness violation test
Liu, Yibing, Haoliang Li, Yangyang Guo, Chenqi Kong, Jing Li, and Shiqi Wang. 2022 · 2022
Closest in time.
Few-shot self-rationalization with natural language prompts
Marasovic, Ana, Iz Beltagy, Doug Downey, and Matthew Peters. 2022 · 2022
Closest in time.
SHAP-based explanation methods: A review for NLP interpretability
Mosca, Edoardo, Ferenc Szigeti, Stella Tragianni, Daniel Gallagher, and Georg Groh. 2022 · 2022
Closest in time.
Causal analysis of syntactic agreement neurons in multilingual language models
Mueller, Aaron, Yu Xia, and Tal Linzen. 2022 · 2022
Closest in time.
An Attention Matrix for Every Decision: Faithfulness-based Arbitration Among Multiple Attention-Based Interpretations of Transformers in Text Classification
Mylonas, Nikolaos, Ioannis Mollas, and Grigorios Tsoumakas. 2022 · 2022
Closest in time.
Evaluating explanations: How much do explanations from the teacher aid students?
Pruthi, Danish, Rachit Bansal, Bhuwan Dhingra, Livio Baldini Soares, Michael Collins, Zachary C. Lipton, Graham Neubig, and William W. Cohen. 2022 · 2022
Closest in time.
Limitations of Language Models in Arithmetic and Symbolic Induction
Qian, Jing, Hong Wang, Zekun Li, Shiyang Li, and Xifeng Yan. 2022 · 2022
Closest in time.
Linear Guardedness and its Implications
Ravfogel, Shauli, Yoav Goldberg, and Ryan Cotterell. 2022 · 2022
Closest in time.
Neuron-level interpretation of deep NLP models: A survey
Sajjad, Hassan, Nadir Durrani, and Fahim Dalvi. 2022 · 2022
Closest in time.
Logical Satisfiability of Counterfactuals for Faithful Explanations in NLI
Sia, Suzanna, Anton Belyy, Amjad Almahairi, Madian Khabsa, Luke Zettlemoyer, and Lambert Mathias. 2022 · 2022
Closest in time.
Chain of Thought Prompting Elicits Reasoning in Large Language Models
Wei, Jason, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed Chi, Quoc Le, and Denny Zhou. 2022 · 2022
Closest in time.
Reframing human-AI collaboration for generating free-text explanations
Wiegreffe, Sarah, Jack Hessel, Swabha Swayamdipta, Mark Riedl, and Yejin Choi. 2022 · 2022
Closest in time.
The Unreliability of Explanations in Few-shot Prompting for Textual Reasoning
Ye, Xi and Greg Durrett. 2022 · 2022
Closest in time.
Complementary Explanations for Effective In-Context Learning
Ye, Xi, Srinivasan Iyer, Asli Celikyilmaz, Ves Stoyanov, Greg Durrett, and Ramakanth Pasunuru. 2022 · 2022
Closest in time.
On the sensitivity and stability of model interpretations in NLP
Yin, Fan, Zhouxing Shi, Cho-Jui Hsieh, and Kai-Wei Chang. 2022 · 2022
Closest in time.
Interpreting language models with contrastive explanations
Yin, Kayo and Graham Neubig. 2022 · 2022
Closest in time.
The irrationality of neural rationale models
Zheng, Yiming, Serena Booth, Julie Shah, and Yilun Zhou. 2022 · 2022
Closest in time.
Do feature attribution methods correctly attribute features?
Zhou, Yilun, Serena Booth, Marco Túlio Ribeiro, and Julie Shah. 2022b · 2022
Closest in time.
ExSum: From local explanations to model understanding
Zhou, Yilun, Marco Tulio Ribeiro, and Julie Shah. 2022 · 2022
Closest in time.
Faithful Chain-of-Thought Reasoning
Lyu, Qing, Shreya Havaldar, Adam Stein, Li Zhang, Delip Rao, Eric Wong, Marianna Apidianaki, and Chris Callison-Burch. 2023 · 2023
Closest in time.
GPT-4 Technical Report
OpenAI. 2023 · 2023
Closest in time.
Explanation Selection Using Unlabeled Data for In-Context Learning
Ye, Xi and Greg Durrett. 2023 · 2023
Closest in time.