Fetching the paper…
Reading the bibliography…
Interpretability techniques are valuable for helping humans understand and oversee AI systems.
Intriguing properties of neural networks, 2014
Christian Szegedy, Wojciech Zaremba, Ilya Sutskever, Joan Bruna, Dumitru Erhan, Ian Goodfellow, and Rob Fergus · 2014
Earlier work this paper cites.
Tom B Brown, Dandelion Mané, Aurko Roy, Martín Abadi, and Justin Gilmer · 2017
Earlier work this paper cites.
Towards a rigorous science of interpretable machine learning
Finale Doshi-Velez and Been Kim · 2017
Earlier work this paper cites.
Feature visualization
Chris Olah, Alexander Mordvintsev, and Ludwig Schubert · 2017
Earlier work this paper cites.
Large scale GAN training for high fidelity natural image synthesis
Andrew Brock, Jeff Donahue, and Karen Simonyan · 2018
Earlier work this paper cites.
Interpretability beyond feature attribution: Quantitative testing with concept activation vectors (TCAV)
Been Kim, Martin Wattenberg, Justin Gilmer, Carrie Cai, James Wexler, Fernanda Viegas, et al · 2018
Earlier work this paper cites.
Exploring neural networks with activation atlases
Shan Carter, Zan Armstrong, Ludwig Schubert, Ian Johnson, and Chris Olah · 2019
Earlier work this paper cites.
Explanation in artificial intelligence: Insights from the social sciences
Tim Miller · 2019
Cited alongside, same era.
Against interpretability: a critical examination of the interpretability problem in machine learning
Maya Krishnan · 2020
Cited alongside, same era.
Compositional explanations of neurons
Jesse Mu and Jacob Andreas · 2020
Cited alongside, same era.
Natural language descriptions of deep visual features
Evan Hernandez, Sarah Schwettmann, David Bau, Teona Bagashvili, Antonio Torralba, and Jacob Andreas · 2021
Cited alongside, same era.
Learning Transferable Visual Models From Natural Language Supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever · 2021
Cited alongside, same era.
Toward transparent ai: A survey on interpreting the inner structures of deep neural networks
Interpreting clip’s image representation via text-based decomposition
Yossi Gandelsman, Alexei A Efros, and Jacob Steinhardt · 2023
Later among the works it cites.
Camels in a changing climate: Enhancing lm adaptation with tulu 2, 2023
Hamish Ivison, Yizhong Wang, Valentina Pyatkin, Nathan Lambert, Matthew Peters, Pradeep Dasigi, Joel Jang, David Wadden, Noah A. Smith, Iz Beltagy, and Hannaneh Hajishirzi · 2023
Later among the works it cites.
Text-To-Concept (and Back) via Cross-Model Alignment
Mazda Moayeri, Keivan Rezaei, Maziar Sanjabi, and Soheil Feizi · 2023
Later among the works it cites.
Prototype generation: Robust feature visualisation for data independent interpretability, 2023
Arush Tagade and Jessica Rumbelow · 2023
Later among the works it cites.
Post-hoc concept bottleneck models
Mert Yuksekgonul, Maggie Wang, and James Zou · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Tilman Räuker, Anson Ho, Stephen Casper, and Dylan Hadfield-Menell · 2022
Cited alongside, same era.
Red teaming deep neural networks with feature synthesis tools, 2023
Stephen Casper, Yuxiao Li, Jiawei Li, Tong Bu, Kevin Zhang, Kaivalya Hariharan, and Dylan Hadfield-Menell · 2023
Cited alongside, same era.
Diagnostics for deep neural networks with automated copy/paste attacks
Stephen Casper, Kaivalya Hariharan, and Dylan Hadfield-Menell
Cited in the paper.
Robust feature-level adversaries are interpretability tools
Stephen Casper, Max Nadeau, Dylan Hadfield-Menell, and Gabriel Kreiman
Cited in the paper.
Jordan Shipard, Arnold Wiliem, Kien Nguyen Thanh, Wei Xiang, and Clinton Fookes · 2024
Closest in time.