Fetching the paper…
Reading the bibliography…
Interpreting the inner workings of deep learning models is crucial for establishing trust and ensuring model safety.
The effect of experimenter bias on the performance of the albino rat
Robert Rosenthal and Kermit L. Fode · 1963
Earlier work this paper cites.
The effect of race/ethnicity on sentencing: Examining sentence type, jail length, and prison length
Kareem L Jordan and Tina L Freiburger · 2015
Earlier work this paper cites.
Imagenet large scale visual recognition challenge
Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, et al · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Rethinking the inception architecture for computer vision
Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jon Shlens, and Zbigniew Wojna · 2016
Earlier work this paper cites.
Network dissection: Quantifying interpretability of deep visual representations
David Bau, Bolei Zhou, Aditya Khosla, Aude Oliva, and Antonio Torralba · 2017
Earlier work this paper cites.
Svcca: Singular vector canonical correlation analysis for deep learning dynamics and interpretability
Maithra Raghu, Justin Gilmer, Jason Yosinski, and Jascha Sohl-Dickstein · 2017
Earlier work this paper cites.
Sanity checks for saliency maps
Julius Adebayo, Justin Gilmer, Michael Muelly, Ian Goodfellow, Moritz Hardt, and Been Kim · 2018
Earlier work this paper cites.
Interpretability beyond feature attribution: Quantitative testing with concept activation vectors (tcav)
Been Kim, Martin Wattenberg, Justin Gilmer, Carrie Cai, James Wexler, Fernanda Viegas, et al · 2018
Cited alongside, same era.
The age of secrecy and unfairness in recidivism prediction, 11 2018
Cynthia Rudin, Caroline Wang, and Beau Coker · 2018
Cited alongside, same era.
Towards automatic concept-based explanations
Amirata Ghorbani, James Wexler, James Y Zou, and Been Kim · 2019
Cited alongside, same era.
Explaining classifiers with causal concept effect (cace)
Yash Goyal, Amir Feder, Uri Shalit, and Been Kim · 2019
Cited alongside, same era.
A benchmark for interpretability methods in deep neural networks
Sara Hooker, Dumitru Erhan, Pieter-Jan Kindermans, and Been Kim · 2019
Cited alongside, same era.
Concept bottleneck models
Pang Wei Koh, Thao Nguyen, Yew Siang Tang, Stephen Mussmann, Emma Pierson, Been Kim, and Percy Liang · 2020
Later among the works it cites.
Zoom in: An introduction to circuits
Chris Olah, Nick Cammarata, Ludwig Schubert, Gabriel Goh, Michael Petrov, and Shan Carter · 2020
Later among the works it cites.
Sharpening local interpretable model-agnostic explanations for histopathology: Improved understandability and reliability
Mara Graziani, Iam Palatnik de Sousa, Marley MBR Vellasco, Eduardo Costa da Silva, Henning Müller, and Vincent Andrearczyk · 2021
Later among the works it cites.
OpenXAI: Towards a transparent evaluation of model explanations
Chirag Agarwal, Satyapriya Krishna, Eshika Saxena, Martin Pawelczyk, Nari Johnson, Isha Puri, Marinka Zitnik, and Himabindu Lakkaraju · 2022
Later among the works it cites.
Nelson Elhage, Tristan Hume, Catherine Olsson, Nicholas Schiefer, Tom Henighan, Shauna Kravec, Zac Hatfield-Dodds, Robert Lasenby, Dawn Drain, Carol Chen, et al · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead
Cynthia Rudin · 2019
Cited alongside, same era.
Concept whitening for interpretable image recognition
Zhi Chen, Yijie Bei, and Cynthia Rudin · 2020
Cited alongside, same era.
Concept attribution: Explaining cnn decisions to physicians
Mara Graziani, Vincent Andrearczyk, Stephane Marchand-Maillet, and Henning Müller · 2020
Cited alongside, same era.
imagenette
Jeremy Howard
Cited in the paper.
Acquisition of chess knowledge in alphazero
Thomas McGrath, Andrei Kapishnikov, Nenad Tomašev, Adam Pearce, Martin Wattenberg, Demis Hassabis, Been Kim, Ulrich Paquet, and Vladimir Kramnik · 2022
Later among the works it cites.
Disentangling neuron representations with concept vectors
Laura O’Mahony, Vincent Andrearczyk, Henning Müller, and Mara Graziani · 2023
Closest in time.