Fetching the paper…
Reading the bibliography…
A precise understanding of why units in an artificial network respond to certain stimuli would constitute a big step towards explainable artificial intelligence.
A coefficient of agreement for nominal scales
Jacob Cohen · 1960
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Fei-Fei Li · 2009
Earlier work this paper cites.
Visualizing higher-layer features of a deep network
Dumitru Erhan, Yoshua Bengio, Aaron Courville, and Pascal Vincent · 2009
Earlier work this paper cites.
<i>why and why not</i> explanations improve the intelligibility of context-aware intelligent systems
Brian Y. Lim, Anind K. Dey, and Daniel Avrahami · 2009
Earlier work this paper cites.
Causality
Judea Pearl · 2009
Earlier work this paper cites.
Deep neural networks: a new framework for modeling biological vision and brain information processing
Nikolaus Kriegeskorte · 2015
Earlier work this paper cites.
Understanding deep image representations by inverting them
Aravindh Mahendran and Andrea Vedaldi · 2015
Earlier work this paper cites.
Inceptionism: Going deeper into neural networks, 2015
Alexander Mordvintsev, Christopher Olah, and Mike Tyka · 2015
Earlier work this paper cites.
Deep neural networks are easily fooled: High confidence predictions for unrecognizable images
Anh Mai Nguyen, Jason Yosinski, and Jeff Clune · 2015
Earlier work this paper cites.
ImageNet Large Scale Visual Recognition Challenge
Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, Alexander C. Berg, and Li Fei-Fei · 2015
Earlier work this paper cites.
Going deeper with convolutions
Christian Szegedy, Wei Liu, Yangqing Jia, Pierre Sermanet, Scott E. Reed, Dragomir Anguelov, Dumitru Erhan, Vincent Vanhoucke, and Andrew Rabinovich · 2015
Earlier work this paper cites.
Psychophysics in a web browser? comparing response times collected with javascript and psychophysics toolbox in a visual search task
Joshua R de Leeuw and Benjamin A Motz · 2016
Earlier work this paper cites.
Simr: an r package for power analysis of generalized linear mixed models by simulation
Peter Green and Catriona J MacLeod · 2016
Earlier work this paper cites.
Synthesizing the preferred inputs for neurons in neural networks via deep generator networks
Anh Mai Nguyen, Alexey Dosovitskiy, Jason Yosinski, Thomas Brox, and Jeff Clune · 2016
Earlier work this paper cites.
Towards a rigorous science of interpretable machine learning
Finale Doshi-Velez and Been Kim · 2017
Earlier work this paper cites.
Plug & play generative networks: Conditional iterative generation of images in latent space
Anh Nguyen, Jeff Clune, Yoshua Bengio, Alexey Dosovitskiy, and Jason Yosinski · 2017
Earlier work this paper cites.
Feature visualization
Chris Olah, Alexander Mordvintsev, and Ludwig Schubert · 2017
Earlier work this paper cites.
Diverse feature visualizations reveal invariances in early layers of deep neural networks
Santiago A Cadena, Marissa A Weis, Leon A Gatys, Matthias Bethge, and Alexander S Ecker · 2018
Earlier work this paper cites.
Manipulating and measuring model interpretability
Forough Poursabzi-Sangdeh, Daniel G Goldstein, Jake M Hofman, Jennifer Wortman Vaughan, and Hanna Wallach · 2018
Earlier work this paper cites.
Counterfactual explanations without opening the black box: Automated decisions and the gdpr, 2018
Sandra Wachter, Brent Mittelstadt, and Chris Russell · 2018
Earlier work this paper cites.
A psychophysics approach for quantitative comparison of interpretable computer vision models
Felix Biessmann and Dionysius Irza Refiano · 2019
Earlier work this paper cites.
The effects of example-based explanations in a machine learning interface
Carrie J Cai, Jonas Jongejan, and Jess Holbrook · 2019
Cited alongside, same era.
Adversarial robustness as a prior for learned representations
Logan Engstrom, Andrew Ilyas, Shibani Santurkar, Dimitris Tsipras, Brandon Tran, and Aleksander Madry · 2019
Cited alongside, same era.
Understanding deep networks via extremal perturbations and smooth masks
Ruth Fong, Mandela Patrick, and Andrea Vedaldi · 2019
Cited alongside, same era.
Counterfactual visual explanations
Yash Goyal, Ziyan Wu, Jan Ernst, Dhruv Batra, Devi Parikh, and Stefan Lee · 2019
Cited alongside, same era.
Understanding neural networks via feature visualization: A survey
Anh Nguyen, Jason Yosinski, and Jeff Clune · 2019
Cited alongside, same era.
Preserving causal constraints in counterfactual explanations for machine learning classifiers, 2020
Divyat Mahajan, Chenhao Tan, and Amit Sharma · 2020
Later among the works it cites.
An overview of early vision in inceptionv1
Chris Olah, Nick Cammarata, Ludwig Schubert, Gabriel Goh, Michael Petrov, and Shan Carter · 2020
Later among the works it cites.
Zoom in: An introduction to circuits
Chris Olah, Nick Cammarata, Ludwig Schubert, Gabriel Goh, Michael Petrov, and Shan Carter · 2020
Later among the works it cites.
OpenAI Microscope
OpenAI · 2020
Later among the works it cites.
To what extent do human explanations of model behavior align with actual model behavior?
Grusha Prasad, Yixin Nie, Mohit Bansal, Robin Jia, Douwe Kiela, and Adina Williams · 2020
Later among the works it cites.
From imagenet to image classification: Contextualizing progress on benchmarks
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Shubham Sharma, Jette Henderson, and Joydeep Ghosh · 2019
Cited alongside, same era.
Robustness may be at odds with accuracy
Dimitris Tsipras, Shibani Santurkar, Logan Engstrom, Alexander Turner, and Aleksander Madry · 2019
Cited alongside, same era.
Actionable recourse in linear classification
Berk Ustun, Alexander Spangher, and Yang Liu · 2019
Cited alongside, same era.
Interpretable counterfactual explanations guided by prototypes
Arnaud Van Looveren and Janis Klaise · 2019
Cited alongside, same era.
Understanding the effect of accuracy on trust in machine learning models
Ming Yin, Jennifer Wortman Vaughan, and Hanna M. Wallach · 2019
Cited alongside, same era.
COGAM: measuring and moderating cognitive load in machine learning model explanations
Ashraf M. Abdul, Christian von der Weth, Mohan S. Kankanhalli, and Brian Y. Lim · 2020
Cited alongside, same era.
Lucas Beyer, Olivier J Hénaff, Alexander Kolesnikov, Xiaohua Zhai, and Aäron van den Oord · 2020
Cited alongside, same era.
Dimitris Tsipras, Shibani Santurkar, Logan Engstrom, Andrew Ilyas, and Aleksander Madry · 2020
Later among the works it cites.
To trust or not to trust an explanation: using leaf to evaluate local linear xai methods
Elvio Amparore, Alan Perotti, and Paolo Bajardi · 2021
Closest in time.
Exemplary natural images explain {cnn} activations better than state-of-the-art feature visualization
Judy Borowski, Roland Simon Zimmermann, Judith Schepers, Robert Geirhos, Thomas S. A. Wallis, Matthias Bethge, and Wieland Brendel · 2021
Closest in time.
Curve circuits
Nick Cammarata, Gabriel Goh, Shan Carter, Chelsea Voss, Ludwig Schubert, and Chris Olah · 2021
Closest in time.
Five points to check when comparing visual perception in humans and machines
Christina M Funke, Judy Borowski, Karolina Stosio, Wieland Brendel, Thomas SA Wallis, and Matthias Bethge · 2021
Closest in time.
Multimodal neurons in artificial neural networks
Gabriel Goh, Nick Cammarata, Chelsea Voss, Shan Carter, Michael Petrov, Ludwig Schubert, Alec Radford, and Chris Olah · 2021
Closest in time.
Calibrated prediction in and out-of-domain for state-of-the-art saliency modeling, 2021
Akis Linardos, Matthias Kümmerer, Ori Press, and Matthias Bethge · 2021
Closest in time.
Quantitative evaluation of machine learning explanations: A human-grounded benchmark
Sina Mohseni, Jeremy E Block, and Eric Ragan · 2021
Closest in time.
Weight banding
Michael Petrov, Chelsea Voss, Ludwig Schubert, Nick Cammarata, Gabriel Goh, and Chris Olah · 2021
Closest in time.
High-low frequency detectors
Ludwig Schubert, Chelsea Voss, Nick Cammarata, Gabriel Goh, and Chris Olah · 2021
Closest in time.
Interactive analysis of cnn robustness
Stefan Sietzen, Mathias Lechner, Judy Borowski, Ramin Hasani, and Manuela Waldner · 2021
Closest in time.
Leveraging sparse linear layers for debuggable deep networks, 2021
Eric Wong, Shibani Santurkar, and Aleksander Mądry · 2021
Closest in time.
Polyjuice: Automated, general-purpose counterfactual generation
Tongshuang Wu, Marco Tulio Ribeiro, Jeffrey Heer, and Daniel S Weld · 2021
Closest in time.
Evaluating explanations for reading comprehension with realistic counterfactuals
Xi Ye, Rohan Nair, and Greg Durrett · 2021
Closest in time.
Cause and effect: Concept-based explanation of neural networks
Mohammad Nokhbeh Zaeem and Majid Komeili · 2021
Closest in time.
Do feature attribution methods correctly attribute features?
Yilun Zhou, Serena Booth, Marco Tulio Ribeiro, and Julie Shah · 2021
Closest in time.