Fetching the paper…
Reading the bibliography…
Feature importance (FI) estimates are a popular form of explanation, and they are commonly created and evaluated by computing the change in model confidence caused by removing certain input features at test time.
Explaining a black-box using deep variational information bottleneck approach
Seo-Jin Bang, Pengtao Xie, Wei Wu, and Eric P. Xing · 1902
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov · 1907
Earlier work this paper cites.
The many shapley values for model explanation
Mukund Sundararajan and Amir Najmi · 1908
Earlier work this paper cites.
Attention is not not explanation
Sarah Wiegreffe and Yuval Pinter · 1908
Earlier work this paper cites.
Fine-grained sentiment analysis with faithful attention
Ruiqi Zhong, Steven Shao, and Kathleen McKeown · 1908
Earlier work this paper cites.
Weight of evidence as a basis for human-oriented explanations
David Alvarez-Melis, Hal Daumé III, Jennifer Wortman Vaughan, and Hanna Wallach · 1910
Earlier work this paper cites.
Feature relevance quantification in explainable ai: A causal problem
Dominik Janzing, Lenon Minorics, and Patrick Blöbaum · 1910
Earlier work this paper cites.
Eraser: A benchmark to evaluate rationalized nlp models
Jay DeYoung, Sarthak Jain, Nazneen Fatema Rajani, Eric Lehman, Caiming Xiong, Richard Socher, and Byron C. Wallace · 1911
Earlier work this paper cites.
An Introduction to the Bootstrap
Bradley Efron and Robert J Tibshirani · 1994
Earlier work this paper cites.
General local search methods
Marc Pirlot · 1996
Earlier work this paper cites.
Fine-Tuning Pretrained Language Models: Weight Initializations, Data Orders, and Early Stopping
Jesse Dodge, Gabriel Ilharco, Roy Schwartz, Ali Farhadi, Hannaneh Hajishirzi, and Noah Smith · 2002
Earlier work this paper cites.
Evaluating explainable ai: Which algorithmic explanations help users predict model behavior?
Peter Hase and Mohit Bansal · 2005
Earlier work this paper cites.
HotFlip: White-box adversarial examples for text classification
Javid Ebrahimi, Anyi Rao, Daniel Lowd, and Dejing Dou · 2006
Earlier work this paper cites.
Evaluations and methods for explanation through robustness analysis
Cheng-Yu Hsieh, Chih-Kuan Yeh, Xuanqing Liu, Pradeep Ravikumar, Seungyeon Kim, Sanjiv Kumar, and Cho-Jui Hsieh · 2006
Earlier work this paper cites.
Using “Annotator Rationales” to Improve Machine Learning for Text Categorization
Omar Zaidan, Jason Eisner, and Christine Piatko · 2007
Earlier work this paper cites.
Explaining classifications for individual instances
M. Robnik-Sikonja and I. Kononenko · 2008
Earlier work this paper cites.
Information-theoretic visual explanation for black-box classifiers
Jihun Yi, Eunji Kim, Siwon Kim, and Sungroh Yoon · 2009
Earlier work this paper cites.
Particle algorithms for optimization on binary spaces
Christian Schäfer · 2013
Earlier work this paper cites.
Deep inside convolutional networks: Visualising image classification models and saliency maps
Karen Simonyan, Andrea Vedaldi, and Andrew Zisserman · 2013
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D. Manning, Andrew Ng, and Christopher Potts · 2013
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D Manning, Andrew Y Ng, and Christopher Potts · 2013
Earlier work this paper cites.
A large annotated corpus for learning natural language inference
Samuel R. Bowman, Gabor Angeli, Christopher Potts, and Christopher D. Manning · 2015
Earlier work this paper cites.
Hyperparameter search in machine learning
Marc Claesen and Bart De Moor · 2015
Earlier work this paper cites.
Visualizing and understanding neural models in nlp
Jiwei Li, Xinlei Chen, Eduard Hovy, and Dan Jurafsky · 2015
Earlier work this paper cites.
Categorical reparameterization with gumbel-softmax
Eric Jang, Shixiang Gu, and Ben Poole · 2016
Earlier work this paper cites.
Understanding neural networks through representation erasure
Jiwei Li, Will Monroe, and Dan Jurafsky · 2016
Cited alongside, same era.
"Why Should I Trust You?": Explaining the predictions of any classifier
Marco Tulio Ribeiro, Sameer Singh, and Carlos Guestrin · 2016
Cited alongside, same era.
End-to-end learning for structured prediction energy networks
David Belanger, Bishan Yang, and Andrew McCallum · 2017
Cited alongside, same era.
Interpretable explanations of black boxes by meaningful perturbation
Ruth C Fong and Andrea Vedaldi · 2017
Cited alongside, same era.
Faithful and customizable explanations of black box models
Himabindu Lakkaraju, Ece Kamar, Rich Caruana, and Jure Leskovec · 2017
Cited alongside, same era.
A benchmark for interpretability methods in deep neural networks
Sara Hooker, Dumitru Erhan, Pieter-Jan Kindermans, and Been Kim · 2019
Later among the works it cites.
Attention is not explanation
Sarthak Jain and Byron C Wallace · 2019
Later among the works it cites.
Inferring which medical treatments work from reports of clinical trials
Eric Lehman, Jay DeYoung, Regina Barzilay, and Byron C. Wallace · 2019
Later among the works it cites.
A unifying view on dataset shift in classification
Jose G Moreno-Torres, Troy Raeder, Rocío Alaiz-Rodríguez, Nitesh V Chawla, and Francisco Herrera · 2019
Later among the works it cites.
Tunability: Importance of hyperparameters of machine learning algorithms
Philipp Probst, Anne-Laure Boulesteix, and Bernd Bischl · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Decoupled weight decay regularization, 2017
Ilya Loshchilov and Frank Hutter · 2017
Cited alongside, same era.
A Unified Approach to Interpreting Model Predictions
Scott M Lundberg and Su-In Lee · 2017
Cited alongside, same era.
The concrete distribution: A continuous relaxation of discrete random variables
Chris J. Maddison, Andriy Mnih, and Yee Whye Teh · 2017
Cited alongside, same era.
Learning important features through propagating activation differences
Avanti Shrikumar, Peyton Greenside, and Anshul Kundaje · 2017
Cited alongside, same era.
Smoothgrad: removing noise by adding noise
Daniel Smilkov, Nikhil Thorat, Been Kim, Fernanda Viégas, and Martin Wattenberg · 2017
Cited alongside, same era.
Axiomatic attribution for deep networks
Mukund Sundararajan, Ankur Taly, and Qiqi Yan · 2017
Cited alongside, same era.
Visualizing deep neural network decisions: Prediction difference analysis
Luisa M Zintgraf, Taco S Cohen, Tameem Adel, and Max Welling · 2017
Cited alongside, same era.
Huggingface’s transformers: State-of-the-art natural language processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, et al · 2019
Later among the works it cites.
Learning variational word masks to improve the interpretability of neural text classifiers
Hanjie Chen and Yangfeng Ji · 2020
Later among the works it cites.
How do decisions emerge across layers in neural models? interpretation with differentiable masking
Nicola De Cao, Michael Sejr Schlichtkrull, Wilker Aziz, and Ivan Titov · 2020
Later among the works it cites.
Interpretation of NLP models through input marginalization
Siwon Kim, Jihun Yi, Eunji Kim, and Sungroh Yoon · 2020
Later among the works it cites.
Explaining machine learning classifiers through diverse counterfactual explanations
Ramaravind Kommiya Mothilal, Amit Sharma, and Chenhao Tan · 2020
Later among the works it cites.
An information bottleneck approach for controlling conciseness in rationale extraction
Bhargavi Paranjape, Mandar Joshi, John Thickstun, Hannaneh Hajishirzi, and Luke Zettlemoyer · 2020
Later among the works it cites.
Visualizing the impact of feature attribution baselines
Pascal Sturmfels, Scott Lundberg, and Su-In Lee · 2020
Later among the works it cites.
Feature importance ranking for deep learning
Maksymilian Wojtas and Ke Chen · 2020
Later among the works it cites.
Transformers: State-of-the-art natural language processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander M. Rush · 2020
Later among the works it cites.
Model-agnostic local explanations with genetic algorithms for text classification
Qingfeng Du and Jincheng Xu · 2021
Closest in time.
The who in explainable ai: How ai background shapes perceptions of ai explanations
Upol Ehsan, Samir Passi, Q Vera Liao, Larry Chan, I Lee, Michael Muller, Mark O Riedl, et al · 2021
Closest in time.
On baselines for local feature attributions
Johannes Haug, Stefan Zürn, Peter El-Jiz, and Gjergji Kasneci · 2021
Closest in time.
Aligning faithful interpretations with their social attribution
Alon Jacovi and Yoav Goldberg · 2021
Closest in time.
Neil Jethani, Mukund Sudarshan, Yindalon Aphinyanaphongs, and Rajesh Ranganath · 2021
Closest in time.
Resisting out-of-distribution data problem in perturbation of xai
Luyu Qiu, Yi Yang, Caleb Chen Cao, Jing Liu, Yueyuan Zheng, Hilary Hei Ting Ngai, Janet Hsiao, and Lei Chen · 2021
Closest in time.
Discretized integrated gradients for explaining language models
Soumya Sanyal and Xiang Ren · 2021
Closest in time.
Rationales for sequential predictions
Keyon Vafa, Yuntian Deng, David M Blei, and Alexander M Rush · 2021
Closest in time.
On the faithfulness measurements for model interpretations, 2021
Fan Yin, Zhouxing Shi, Cho-Jui Hsieh, and Kai-Wei Chang · 2021
Closest in time.
Do feature attribution methods correctly attribute features?
Yilun Zhou, Serena Booth, Marco Tulio Ribeiro, and Julie Shah · 2021
Closest in time.