Fetching the paper…
Reading the bibliography…
The large size and complex decision mechanisms of state-of-the-art text classifiers make it difficult for humans to understand their predictions, leading to a potential lack of trust by the users.
A Survey on Explainable Artificial Intelligence (XAI): Towards Medical XAI
Erico Tjoa and Cuntai Guan · 1907
Earlier work this paper cites.
Random Forests
Leo Breiman · 2001
Earlier work this paper cites.
Topical N-Grams: Phrase and Topic Discovery, with an Application to Information Retrieval
X. Wang, A. McCallum, and X. Wei · 2007
Earlier work this paper cites.
Captum: A Unified and Generic Model Interpretability Library for PyTorch
Narine Kokhlikyan, Vivek Miglani, Miguel Martin, Edward Wang, Bilal Alsallakh, Jonathan Reynolds, Alexander Melnikov, Natalia Kliushkina, Carlos Araya, Siqi Yan, and Orion Reblitz-Richardson · 2009
Earlier work this paper cites.
Learning Word Vectors for Sentiment Analysis
Andrew L. Maas, Raymond E. Daly, Peter T. Pham, Dan Huang, Andrew Y. Ng, and Christopher Potts · 2011
Earlier work this paper cites.
Improving KernelSHAP: Practical Shapley Value Estimation via Linear Regression
Ian Covert and Su-In Lee · 2012
Earlier work this paper cites.
ImageNet Classification with Deep Convolutional Neural Networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E. Hinton · 2012
Earlier work this paper cites.
Sentence-Based Model Agnostic NLP Interpretability
Yves Rychener, Xavier Renard, Djamé Seddah, Pascal Frossard, and Marcin Detyniecki · 2012
Earlier work this paper cites.
Explaining Predictions of Non-Linear Classifiers in NLP
Leila Arras, Franziska Horn, Grégoire Montavon, Klaus-Robert Müller, and Wojciech Samek · 2016
Earlier work this paper cites.
Visualizing and Understanding Neural Models in NLP
Jiwei Li, Xinlei Chen, Eduard Hovy, and Dan Jurafsky · 2016
Earlier work this paper cites.
"Why Should I Trust You?": Explaining the Predictions of Any Classifier
Marco Tulio Ribeiro, Sameer Singh, and Carlos Guestrin · 2016
Earlier work this paper cites.
Towards A Rigorous Science of Interpretable Machine Learning
Finale Doshi-Velez and Been Kim · 2017
Earlier work this paper cites.
A Unified Approach to Interpreting Model Predictions
Scott M. Lundberg and Su-In Lee · 2017
Earlier work this paper cites.
Evaluating the Visualization of What a Deep Neural Network Has Learned
W. Samek, A. Binder, G. Montavon, S. Lapuschkin, and K. Müller · 2017
Earlier work this paper cites.
Learning Important Features Through Propagating Activation Differences
Avanti Shrikumar, Peyton Greenside, and Anshul Kundaje · 2017
Earlier work this paper cites.
Axiomatic Attribution for Deep Networks
Mukund Sundararajan, Ankur Taly, and Qiqi Yan · 2017
Earlier work this paper cites.
Sanity Checks for Saliency Maps
Julius Adebayo, Justin Gilmer, Michael Muelly, Ian Goodfellow, Moritz Hardt, and Been Kim · 2018
Earlier work this paper cites.
Towards Robust Interpretability with Self-explaining Neural Networks
David Alvarez-Melis and Tommi S. Jaakkola · 2018
Earlier work this paper cites.
Pathologies of Neural Models Make Interpretations Difficult
Shi Feng, Eric Wallace, Alvin Grissom II, Mohit Iyyer, Pedro Rodriguez, and Jordan Boyd-Graber · 2018
Earlier work this paper cites.
Explaining Explanations: An Overview of Interpretability of Machine Learning
L. H. Gilpin, D. Bau, B. Z. Yuan, A. Bajwa, M. Specter, and L. Kagal · 2018
Earlier work this paper cites.
A Survey of Methods for Explaining Black Box Models
Riccardo Guidotti, Anna Monreale, Salvatore Ruggieri, Franco Turini, Fosca Giannotti, and Dino Pedreschi · 2018
Earlier work this paper cites.
An Evaluation of the Human-Interpretability of Explanation
Isaac Lage, Emily Chen, Jeffrey He, Menaka Narayanan, Been Kim, Sam Gershman, and Finale Doshi-Velez · 2018
Earlier work this paper cites.
The Mythos of Model Interpretability
Zachary C. Lipton · 2018
Cited alongside, same era.
SHAP, 2018
Scott Lundberg · 2018
Cited alongside, same era.
Evaluating Neural Network Explanation Methods using Hybrid Documents and Morphosyntactic Agreement
Nina Poerner, Hinrich Schütze, and Benjamin Roth · 2018
Cited alongside, same era.
Anchors: High-Precision Model-Agnostic Explanations
Marco Tulio Ribeiro, Sameer Singh, and Carlos Guestrin · 2018
Cited alongside, same era.
Evaluating Recurrent Neural Network Explanations
Leila Arras, Ahmed Osman, Klaus-Robert Müller, and Wojciech Samek · 2019
Cited alongside, same era.
How to Fine-Tune BERT for Text Classification?
Chi Sun, Xipeng Qiu, Yige Xu, and Xuanjing Huang · 2019
Later among the works it cites.
AllenNLP Interpret: A Framework for Explaining Predictions of NLP Models
Eric Wallace, Jens Tuyls, Junlin Wang, Sanjay Subramanian, Matt Gardner, and Sameer Singh · 2019
Later among the works it cites.
Debugging Tests for Model Explanations
Julius Adebayo, Michael Muelly, Ilaria Liccardi, and Been Kim · 2020
Later among the works it cites.
A Diagnostic Study of Explainability Techniques for Text Classification
Pepa Atanasova, Jakob Grue Simonsen, Christina Lioma, and Isabelle Augenstein · 2020
Later among the works it cites.
Explainable Machine Learning in Deployment
Umang Bhatt, Alice Xiang, Shubham Sharma, Adrian Weller, Ankur Taly, Yunhan Jia, Joydeep Ghosh, Ruchir Puri, José M. F. Moura, and Peter Eckersley · 2020
Later among the works it cites.
Language Models are Few-Shot Learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Iz Beltagy, Kyle Lo, and Arman Cohan · 2019
Cited alongside, same era.
Robust Attribution Regularization
Jiefeng Chen, Xi Wu, Vaibhav Rastogi, Yingyu Liang, and Somesh Jha · 2019
Cited alongside, same era.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Cited alongside, same era.
Understanding Deep Networks via Extremal Perturbations and Smooth Masks
Ruth Fong, Mandela Patrick, and Andrea Vedaldi · 2019
Cited alongside, same era.
Causal structure based root cause analysis of outliers
Dominik Janzing, Kailash Budhathoki, Lenon Minorics, and Patrick Blöbaum · 2019
Cited alongside, same era.
XRAI: Better Attributions Through Regions
Andrei Kapishnikov, Tolga Bolukbasi, Fernanda Viegas, and Michael Terry · 2019
Cited alongside, same era.
Faithful and Customizable Explanations of Black Box Models
Himabindu Lakkaraju, Ece Kamar, Rich Caruana, and Jure Leskovec · 2019
Cited alongside, same era.
Later among the works it cites.
ERASER: A Benchmark to Evaluate Rationalized NLP Models
Jay DeYoung, Sarthak Jain, Nazneen Fatema Rajani, Eric Lehman, Caiming Xiong, Richard Socher, and Byron C. Wallace · 2020
Later among the works it cites.
Evaluating Explainable AI: Which Algorithmic Explanations Help Users Predict Model Behavior?
Peter Hase and Mohit Bansal · 2020
Later among the works it cites.
spaCy: Industrial-strength Natural Language Processing in Python, 2020
Matthew Honnibal, Ines Montani, Sofie Van Landeghem, and Adriane Boyd · 2020
Later among the works it cites.
Beyond Accuracy: Behavioral Testing of NLP Models with CheckList
Marco Tulio Ribeiro, Tongshuang Wu, Carlos Guestrin, and Sameer Singh · 2020
Later among the works it cites.
The Many Shapley Values for Model Explanation
Mukund Sundararajan and Amir Najmi · 2020
Later among the works it cites.
Ian Tenney, James Wexler, Jasmijn Bastings, Tolga Bolukbasi, Andy Coenen, Sebastian Gehrmann, Ellen Jiang, Mahima Pushkarna, Carey Radebaugh, Emily Reif, and Ann Yuan · 2020
Later among the works it cites.
Transformers: State-of-the-Art Natural Language Processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Remi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander Rush · 2020
Later among the works it cites.
FinBERT: A Pretrained Language Model for Financial Communications
Yi Yang, Mark Christopher Siy UY, and Allen Huang · 2020
Later among the works it cites.
Persistent Anti-Muslim Bias in Large Language Models
Abubakar Abid, Maheen Farooqi, and James Zou · 2021
Closest in time.
On the Dangers of Stochastic Parrots: Can Language Models Be Too Big? 🦜
Emily M. Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell · 2021
Closest in time.
BOLD: Dataset and Metrics for Measuring Biases in Open-Ended Language Generation
Jwala Dhamala, Tony Sun, Varun Kumar, Satyapriya Krishna, Yada Pruksachatkun, Kai-Wei Chang, and Rahul Gupta · 2021
Closest in time.
Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity
William Fedus, Barret Zoph, and Noam Shazeer · 2021
Closest in time.
Manipulating and Measuring Model Interpretability
Forough Poursabzi-Sangdeh, Daniel G. Goldstein, Jake M. Hofman, Jennifer Wortman Vaughan, and Hanna Wallach · 2021
Closest in time.
Reliable Post hoc Explanations: Modeling Uncertainty in Explainability
Dylan Slack, Sophie Hilgard, Sameer Singh, and Himabindu Lakkaraju · 2021
Closest in time.
On the Lack of Robust Interpretability of Neural Text Classifiers
Muhammad Bilal Zafar, Michele Donini, Dylan Slack, Cedric Archambeau, Sanjiv Das, and Krishnaram Kenthapadi · 2021
Closest in time.