Towards better understanding of gradient-based attribution methods for Deep Neural Networks
Marco Ancona, Enea Ceolini, Cengiz Öztireli, and Markus Gross · 2018
Later among the works it cites.
Explainable artificial intelligence: A survey
Filip Karlo Došilović, Mario Brčić, and Nikica Hlupić · 2018
Later among the works it cites.
Net2Vec: Quantifying and Explaining How Concepts are Encoded by Filters in Deep Neural Networks
Ruth Fong and Andrea Vedaldi · 2018
Later among the works it cites.
A Survey of Methods for Explaining Black Box Models
Riccardo Guidotti, Anna Monreale, Salvatore Ruggieri, Franco Turini, Fosca Giannotti, and Dino Pedreschi · 2018
Later among the works it cites.
Visual Analytics in Deep Learning: An Interrogative Survey for the Next Frontiers
Fred Hohman, Minsuk Kahng, Robert Pienta, and Duen Horng Chau · 2018
Later among the works it cites.
Transparency and Explanation in Deep Reinforcement Learning Neural Networks
Rahul Iyer, Yuezhang Li, Huao Li, Michael Lewis, Ramitha Sundar, and Katia Sycara · 2018
Later among the works it cites.
Interpretability Beyond Feature Attribution: Quantitative Testing with Concept Activation Vectors (TCAV)
Original
Been Kim, Martin Wattenberg, Justin Gilmer, Carrie Cai, James Wexler, Fernanda Viegas, and Rory Sayres · 2018
Later among the works it cites.
The mythos of model interpretability
Zachary C Lipton · 2018
Later among the works it cites.
Troubling Trends in Machine Learning Scholarship
Original
Zachary C. Lipton and Jacob Steinhardt · 2018
Later among the works it cites.
The Building Blocks of Interpretability
Chris Olah, Arvind Satyanarayan, Ian Johnson, Shan Carter, Ludwig Schubert, Katherine Ye, and Alexander Mordvintsev · 2018
Later among the works it cites.
High-dimensional probability: An introduction with applications in data science , volume 47
Roman Vershynin · 2018
Later among the works it cites.
TieNet: Text-Image Embedding Network for Common Thorax Disease Classification and Reporting in Chest X-rays
Original
Xiaosong Wang, Yifan Peng, Le Lu, Zhiyong Lu, and Ronald M. Summers · 2018
Later among the works it cites.
Convolutional neural networks: an overview and application in radiology
Rikiya Yamashita, Mizuho Nishio, Richard Kinh Gian Do, and Kaori Togashi · 2018
Later among the works it cites.
Interpreting Deep Visual Representations via Network Dissection
B. Zhou, D. Bau, A. Oliva, and A. Torralba · 2018
Later among the works it cites.
Revisiting the Importance of Individual Units in CNNs via Ablation
Original
Bolei Zhou, Yiyou Sun, David Bau, and Antonio Torralba · 2018
Later among the works it cites.
Exploratory Not Explanatory: Counterfactual Analysis of Saliency Maps for Deep Reinforcement Learning
Akanksha Atrey, Kaleigh Clary, and David Jensen · 2019
Later among the works it cites.
GAN Dissection: Visualizing and Understanding Generative Adversarial Networks
David Bau, Jun-Yan Zhu, Hendrik Strobelt, Bolei Zhou, Joshua B. Tenenbaum, William T. Freeman, and Antonio Torralba · 2019
Later among the works it cites.
Prototypical Examples in Deep Learning: Metrics, Characteristics, and Utility
Nicholas Carlini, Ulfar Erlingsson, and Nicolas Papernot · 2019
Later among the works it cites.
What Is One Grain of Sand in the Desert? Analyzing Individual Neurons in Deep NLP Models
Fahim Dalvi, Nadir Durrani, Hassan Sajjad, Yonatan Belinkov, Anthony Bau, and James Glass · 2019
Later among the works it cites.
How Important is a Neuron
Kedar Dhamdhere, Mukund Sundararajan, and Qiqi Yan · 2019
Later among the works it cites.
On Interpretability and Feature Representations: An Analysis of the Sentiment Neuron
Jonathan Donnelly and Adam Roegiest · 2019
Later among the works it cites.
A Benchmark for Interpretability Methods in Deep Neural Networks
Sara Hooker, Dumitru Erhan, Pieter-Jan Kindermans, and Been Kim · 2019
Later among the works it cites.
Discovery of Natural Language Concepts in Individual Units of CNNs
Seil Na, Yo Joong Choe, Dong-Hyun Lee, and Gunhee Kim · 2019
Later among the works it cites.
Explain Your Move: Understanding Agent Actions Using Specific and Relevant Feature Attribution
Nikaash Puri, Sukriti Verma, Piyush Gupta, Dhruv Kayastha, Shripad Deshmukh, Balaji Krishnamurthy, and Sameer Singh · 2019
Later among the works it cites.
Understanding trained CNNs by indexing neuron selectivity
Original
Ivet Rafegas, Maria Vanrell, Luis A. Alexandre, and Guillem Arias · 2019
Later among the works it cites.
Learning Reliable Visual Saliency for Model Explanations
Yulong Wang, Hang Su, Bo Zhang, and Xiaolin Hu · 2019
Later among the works it cites.
On the (In)fidelity and Sensitivity of Explanations
Chih-Kuan Yeh, Cheng-Yu Hsieh, Arun Suggala, David I Inouye, and Pradeep K Ravikumar · 2019
Later among the works it cites.
Foundations of Data Science
Avrim Blum, John Hopcroft, and Ravi Kannan · 2020
Closest in time.
Zoom In: An Introduction to Circuits
Chris Olah, Nick Cammarata, Ludwig Schubert, Gabriel Goh, Michael Petrov, and Shan Carter · 2020
Closest in time.
Explainable Machine Learning for Scientific Insights and Discoveries
Ribana Roscher, Bastian Bohn, Marco F. Duarte, and Jochen Garcke · 2020
Closest in time.
Attribution for graph neural networks
Benjamin Sanchez-Lengeling, Jennifer N Wei, Brian K Lee, Emily Reif, Peter Y Wang, Wesley Wei Qian, Kevin McCLoskey, Lucy Colwell, and Alexander B Wiltschko · 2020
Closest in time.
Attribution-based Salience Method towards Interpretable Reinforcement Learning
Yuyao Wang, Masayoshi Mase, and Masashi Egi · 2020
Closest in time.
Visual Perception and the Statistical Properties of Natural Scenes
Wilson S. Geisler · 2085
Closest in time.