Fetching the paper…
Reading the bibliography…
With the ever-increasing complexity of neural language models, practitioners have turned to methods for understanding the predictions of these models.
RoBERTa: A Robustly Optimized BERT Pretraining Approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 1907
Earlier work this paper cites.
A Survey on Explainable Artificial Intelligence (XAI): Towards Medical XAI
Erico Tjoa and Cuntai Guan. 2020 · 1907
Earlier work this paper cites.
Benchmarking Attribution Methods with Relative Feature Importance
Mengjiao Yang and Been Kim. 2019 · 1907
Earlier work this paper cites.
An elementary mathematical theory of classification and prediction
T. T. Tanimoto. 1958 · 1958
Earlier work this paper cites.
Reading Tea Leaves: How Humans Interpret Topic Models
Jonathan Chang, Sean Gerrish, Chong Wang, Jordan Boyd-graber, and David Blei. 2009 · 2009
Earlier work this paper cites.
Captum: A Unified and Generic Model Interpretability Library for PyTorch
Narine Kokhlikyan, Vivek Miglani, Miguel Martin, Edward Wang, Bilal Alsallakh, Jonathan Reynolds, Alexander Melnikov, Natalia Kliushkina, Carlos Araya, Siqi Yan, and Orion Reblitz-Richardson. 2020 · 2009
Earlier work this paper cites.
Learning Word Vectors for Sentiment Analysis
Andrew L. Maas, Raymond E. Daly, Peter T. Pham, Dan Huang, Andrew Y. Ng, and Christopher Potts. 2011 · 2011
Earlier work this paper cites.
Recursive Deep Models for Semantic Compositionality Over a Sentiment Treebank
Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D. Manning, Andrew Ng, and Christopher Potts. 2013 · 2013
Earlier work this paper cites.
Good Debt or Bad Debt: Detecting Semantic Orientations in Economic Texts
Pekka Malo, Ankur Sinha, Pekka Korhonen, Jyrki Wallenius, and Pyry Takala. 2014 · 2014
Earlier work this paper cites.
Deep Inside Convolutional Networks: Visualising Image Classification Models and Saliency Maps
Karen Simonyan, Andrea Vedaldi, and Andrew Zisserman. 2014 · 2014
Earlier work this paper cites.
Delving Deep into Rectifiers: Surpassing Human-Level Performance on ImageNet Classification
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2015 · 2015
Earlier work this paper cites.
Explaining Predictions of Non-Linear Classifiers in NLP
Leila Arras, Franziska Horn, Grégoire Montavon, Klaus-Robert Müller, and Wojciech Samek. 2016 · 2016
Earlier work this paper cites.
Visualizing and Understanding Neural Models in NLP
Jiwei Li, Xinlei Chen, Eduard Hovy, and Dan Jurafsky. 2016 · 2016
Earlier work this paper cites.
”Why Should I Trust You?”: Explaining the Predictions of Any Classifier
Marco Tulio Ribeiro, Sameer Singh, and Carlos Guestrin. 2016 · 2016
Earlier work this paper cites.
Towards A Rigorous Science of Interpretable Machine Learning
Finale Doshi-Velez and Been Kim. 2017 · 2017
Earlier work this paper cites.
A Unified Approach to Interpreting Model Predictions
Scott M. Lundberg and Su-In Lee. 2017 · 2017
Earlier work this paper cites.
Evaluating the Visualization of What a Deep Neural Network Has Learned
W. Samek, A. Binder, G. Montavon, S. Lapuschkin, and K. Müller. 2017 · 2017
Earlier work this paper cites.
Learning Important Features Through Propagating Activation Differences
Avanti Shrikumar, Peyton Greenside, and Anshul Kundaje. 2017 · 2017
Earlier work this paper cites.
SmoothGrad: Removing Noise by Adding Noise
Daniel Smilkov, Nikhil Thorat, Been Kim, Fernanda Viégas, and Martin Wattenberg. 2017 · 2017
Earlier work this paper cites.
Axiomatic Attribution for Deep Networks
Mukund Sundararajan, Ankur Taly, and Qiqi Yan. 2017 · 2017
Cited alongside, same era.
Sanity Checks for Saliency Maps
Julius Adebayo, Justin Gilmer, Michael Muelly, Ian Goodfellow, Moritz Hardt, and Been Kim. 2018 · 2018
Cited alongside, same era.
Towards Robust Interpretability with Self-explaining Neural Networks
David Alvarez-Melis and Tommi S. Jaakkola. 2018 · 2018
Cited alongside, same era.
Pathologies of Neural Models Make Interpretations Difficult
Shi Feng, Eric Wallace, Alvin Grissom II, Mohit Iyyer, Pedro Rodriguez, and Jordan Boyd-Graber. 2018 · 2018
Cited alongside, same era.
Explaining Explanations: An Overview of Interpretability of Machine Learning
L. H. Gilpin, D. Bau, B. Z. Yuan, A. Bajwa, M. Specter, and L. Kagal. 2018 · 2018
Cited alongside, same era.
A Survey of Methods for Explaining Black Box Models
PyTorch: An Imperative Style, High-Performance Deep Learning Library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Kopf, Edward Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala. 2019 · 2019
Later among the works it cites.
Quantifying Interpretability and Trust in Machine Learning Systems
Philipp Schmidt and Felix Biessmann. 2019 · 2019
Later among the works it cites.
How to Fine-Tune BERT for Text Classification?
Chi Sun, Xipeng Qiu, Yige Xu, and Xuanjing Huang. 2019 · 2019
Later among the works it cites.
AllenNLP Interpret: A Framework for Explaining Predictions of NLP Models
Eric Wallace, Jens Tuyls, Junlin Wang, Sanjay Subramanian, Matt Gardner, and Sameer Singh. 2019 · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Riccardo Guidotti, Anna Monreale, Salvatore Ruggieri, Franco Turini, Fosca Giannotti, and Dino Pedreschi. 2018 · 2018
Cited alongside, same era.
Interpretability Beyond Feature Attribution: Quantitative Testing with Concept Activation Vectors (TCAV)
Been Kim, Martin Wattenberg, Justin Gilmer, Carrie Cai, James Wexler, Fernanda Viegas, and Rory Sayres. 2018 · 2018
Cited alongside, same era.
Automatic Shadow Detection in 2D Ultrasound Images
Qingjie Meng, Christian Baumgartner, Matthew Sinclair, James Housden, Martin Rajchl, Alberto Gomez, Benjamin Hou, Nicolas Toussaint, Veronika Zimmer, Jeremy Tan, Jacqueline Matthew, Daniel Rueckert, Julia Schnabel, and Bernhard Kainz. 2018 · 2018
Cited alongside, same era.
Comparing Automatic and Human Evaluation of Local Explanations for Text Classification
Dong Nguyen. 2018 · 2018
Cited alongside, same era.
Learning Global Additive Explanations for Neural Nets Using Model Distillation
Sarah Tan, Rich Caruana, Giles Hooker, Paul Koch, and Albert Gordo. 2018 · 2018
Cited alongside, same era.
Robust Attribution Regularization
Jiefeng Chen, Xi Wu, Vaibhav Rastogi, Yingyu Liang, and Somesh Jha. 2019 · 2019
Cited alongside, same era.
Bias in Bios: A Case Study of Semantic Representation Bias in a High-Stakes Setting
Maria De-Arteaga, Alexey Romanov, Hanna Wallach, Jennifer Chayes, Christian Borgs, Alexandra Chouldechova, Sahin Geyik, Krishnaram Kenthapadi, and Adam Tauman Kalai. 2019 · 2019
Cited alongside, same era.
Julius Adebayo, Michael Muelly, Ilaria Liccardi, and Been Kim. 2020 · 2020
Later among the works it cites.
Fairwashing Explanations with Off-manifold Detergent
Christopher Anders, Plamen Pasliev, Ann-Kathrin Dombrowski, Klaus-Robert Müller, and Pan Kessel. 2020 · 2020
Later among the works it cites.
A Diagnostic Study of Explainability Techniques for Text Classification
Pepa Atanasova, Jakob Grue Simonsen, Christina Lioma, and Isabelle Augenstein. 2020 · 2020
Later among the works it cites.
Explainable Machine Learning in Deployment
Umang Bhatt, Alice Xiang, Shubham Sharma, Adrian Weller, Ankur Taly, Yunhan Jia, Joydeep Ghosh, Ruchir Puri, José M. F. Moura, and Peter Eckersley. 2020 · 2020
Later among the works it cites.
Concise Explanations of Neural Networks using Adversarial Training
Prasad Chalasani, Jiefeng Chen, Amrita Roy Chowdhury, Xi Wu, and Somesh Jha. 2020 · 2020
Later among the works it cites.
ERASER: A Benchmark to Evaluate Rationalized NLP Models
Jay DeYoung, Sarthak Jain, Nazneen Fatema Rajani, Eric Lehman, Caiming Xiong, Richard Socher, and Byron C. Wallace. 2020 · 2020
Later among the works it cites.
Evaluating Explainable AI: Which Algorithmic Explanations Help Users Predict Model Behavior?
Peter Hase and Mohit Bansal. 2020 · 2020
Later among the works it cites.
Problems with Shapley-value-based explanations as feature importance measures
I. Elizabeth Kumar, Suresh Venkatasubramanian, Carlos Scheidegger, and Sorelle Friedler. 2020 · 2020
Later among the works it cites.
Robust and Stable Black Box Explanations
Himabindu Lakkaraju, Nino Arsov, and Osbert Bastani. 2020 · 2020
Later among the works it cites.
Fooling LIME and SHAP: Adversarial Attacks on Post hoc Explanation Methods
Dylan Slack, Sophie Hilgard, Emily Jia, Sameer Singh, and Himabindu Lakkaraju. 2020 · 2020
Later among the works it cites.
The Many Shapley Values for Model Explanation
Mukund Sundararajan and Amir Najmi. 2020 · 2020
Later among the works it cites.
Sanity Checks for Saliency Metrics
Richard Tomsett, Dan Harborne, Supriyo Chakraborty, Prudhvi Gurram, and Alun Preece. 2020 · 2020
Later among the works it cites.
Transformers: State-of-the-Art Natural Language Processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Remi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander Rush. 2020 · 2020
Later among the works it cites.
Manipulating and Measuring Model Interpretability
Forough Poursabzi-Sangdeh, Daniel G. Goldstein, Jake M. Hofman, Jennifer Wortman Vaughan, and Hanna Wallach. 2021 · 2021
Closest in time.