Fetching the paper…
Reading the bibliography…
Labeling neural network submodules with human-legible descriptions is useful for many downstream tasks: such descriptions can surface failures, guide interventions, and perhaps even explain important model behaviors.
Cognitron: A self-organizing multilayered neural network
K. Fukushima · 1975
Earlier work this paper cites.
What one intelligence test measures: a theoretical account of the processing in the raven progressive matrices test
P. A. Carpenter, M. A. Just, and P. Shell · 1990
Earlier work this paper cites.
Fluid concepts and creative analogies: Computer models of the fundamental mechanisms of thought
D. R. Hofstadter · 1995
Earlier work this paper cites.
The copycat project: A model of mental fluidity and analogy-making
D. R. Hofstadter, M. Mitchell, et al · 1995
Earlier work this paper cites.
Rectified linear units improve restricted boltzmann machines
V. Nair and G. E. Hinton · 2010
Earlier work this paper cites.
Wikidata: a free collaborative knowledgebase
D. Vrandečić and M. Krötzsch · 2014
Earlier work this paper cites.
Visualizing and understanding convolutional networks
M. D. Zeiler and R. Fergus · 2014
Earlier work this paper cites.
Object detectors emerge in deep scene cnns
B. Zhou, A. Khosla, A. Lapedriza, A. Oliva, and A. Torralba · 2014
Earlier work this paper cites.
Understanding deep image representations by inverting them
A. Mahendran and A. Vedaldi · 2015
Earlier work this paper cites.
Automatic generation of raven’s progressive matrices
K. Wang and Z. Su · 2015
Earlier work this paper cites.
Generating visual explanations
L. A. Hendricks, Z. Akata, M. Rohrbach, J. Donahue, B. Schiele, and T. Darrell · 2016
Earlier work this paper cites.
Network dissection: Quantifying interpretability of deep visual representations
D. Bau, B. Zhou, A. Khosla, A. Oliva, and A. Torralba · 2017
Earlier work this paper cites.
Towards a rigorous science of interpretable machine learning
F. Doshi-Velez and B. Kim · 2017
Earlier work this paper cites.
Interpretable explanations of black boxes by meaningful perturbation
R. C. Fong and A. Vedaldi · 2017
Earlier work this paper cites.
Inconsistency of bayesian inference for misspecified linear models, and a proposal for repairing it
P. Grünwald and T. Van Ommen · 2017
Earlier work this paper cites.
Modeling visual problem solving as analogical reasoning
A. Lovett and K. Forbus · 2017
Earlier work this paper cites.
Feature visualization
C. Olah, A. Mordvintsev, and L. Schubert · 2017
Earlier work this paper cites.
Grad-cam: Visual explanations from deep networks via gradient-based localization
R. R. Selvaraju, M. Cogswell, A. Das, R. Vedantam, D. Parikh, and D. Batra · 2017
Earlier work this paper cites.
Sanity checks for saliency maps
J. Adebayo, J. Gilmer, M. Muelly, I. Goodfellow, M. Hardt, and B. Kim · 2018
Earlier work this paper cites.
Understanding disentangling in β \beta -vae
C. P. Burgess, I. Higgins, A. Pal, L. Matthey, N. Watters, G. Desjardins, and A. Lerchner · 2018
Earlier work this paper cites.
e-snli: Natural language inference with natural language explanations
O.-M. Camburu, T. Rocktäschel, T. Lukasiewicz, and P. Blunsom · 2018
Earlier work this paper cites.
Rationalization: A neural machine translation approach to generating natural language explanations
U. Ehsan, B. Harrison, L. Chan, and M. O. Riedl · 2018
Earlier work this paper cites.
Net2vec: Quantifying and explaining how concepts are encoded by filters in deep neural networks
R. Fong and A. Vedaldi · 2018
Cited alongside, same era.
Interpretability beyond feature attribution: Quantitative testing with concept activation vectors (TCAV)
B. Kim, M. Wattenberg, J. Gilmer, C. Cai, J. Wexler, F. Viegas, et al · 2018
Cited alongside, same era.
The mythos of model interpretability: In machine learning, the concept of interpretability is both important and slippery
Z. C. Lipton · 2018
Cited alongside, same era.
Rise: Randomized input sampling for explanation of black-box models
V. Petsiuk, A. Das, and K. Saenko · 2018
Cited alongside, same era.
Modeling semantics with gated graph neural networks for knowledge base question answering
D. Sorokin and I. Gurevych · 2018
Cited alongside, same era.
Transformer feed-forward layers are key-value memories
M. Geva, R. Schuster, J. Berant, and O. Levy · 2021
Later among the works it cites.
Multimodal neurons in artificial neural networks
G. Goh, N. Cammarata, C. Voss, S. Carter, M. Petrov, L. Schubert, A. Radford, and C. Olah · 2021
Later among the works it cites.
Abstraction and analogy-making in artificial intelligence
M. Mitchell · 2021
Later among the works it cites.
Stylespace analysis: Disentangled controls for stylegan image generation
Z. Wu, D. Lischinski, and E. Shechtman · 2021
Later among the works it cites.
Natural language descriptions of deep visual features
E. Hernandez, S. Schwettmann, D. Bau, T. Bagashvili, A. Torralba, and J. Andreas · 2022
Later among the works it cites.
Hive: evaluating the human interpretability of visual explanations
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
J. Zhang, S. A. Bargal, Z. Lin, J. Brandt, X. Shen, and S. Sclaroff · 2018
Cited alongside, same era.
On the measure of intelligence
F. Chollet · 2019
Cited alongside, same era.
A benchmark for interpretability methods in deep neural networks
S. Hooker, D. Erhan, P.-J. Kindermans, and B. Kim · 2019
Cited alongside, same era.
Explanation in artificial intelligence: Insights from the social sciences
T. Miller · 2019
Cited alongside, same era.
Interpretable and fine-grained visual explanations for convolutional neural networks
J. Wagner, J. M. Kohler, T. Gindele, L. Hetzel, J. T. Wiedemer, and S. Behnke · 2019
Cited alongside, same era.
Benchmarking Attribution Methods with Relative Feature Importance
M. Yang and B. Kim · 2019
Cited alongside, same era.
Understanding the role of individual units in a deep neural network
D. Bau, J.-Y. Zhu, H. Strobelt, A. Lapedriza, B. Zhou, and A. Torralba · 2020
Cited alongside, same era.
S. S. Kim, N. Meister, V. V. Ramaswamy, R. Fong, and O. Russakovsky · 2022
Later among the works it cites.
Locating and editing factual associations in GPT
K. Meng, D. Bau, A. Andonian, and Y. Belinkov · 2022
Later among the works it cites.
Progress measures for grokking via mechanistic interpretability
N. Nanda, L. Chan, T. Lieberum, J. Smith, and J. Steinhardt · 2022
Later among the works it cites.
Interpretability in the wild: a circuit for indirect object identification in gpt-2 small
K. R. Wang, A. Variengien, A. Conmy, B. Shlegeris, and J. Steinhardt · 2022
Later among the works it cites.
Chain-of-thought prompting elicits reasoning in large language models
J. Wei, X. Wang, D. Schuurmans, M. Bosma, F. Xia, E. Chi, Q. V. Le, D. Zhou, et al · 2022
Later among the works it cites.
Language models can explain neurons in language models
S. Bills, N. Cammarata, D. Mossing, H. Tillman, L. Gao, G. Goh, I. Sutskever, J. Leike, J. Wu, and W. Saunders · 2023
Closest in time.
Benchmarking interpretability tools for deep neural networks
S. Casper, Y. Li, J. Li, T. Bu, K. Zhang, and D. Hadfield-Menell · 2023
Closest in time.
Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality, March 2023
W.-L. Chiang, Z. Li, Z. Lin, Y. Sheng, Z. Wu, H. Zhang, L. Zheng, S. Zhuang, Y. Zhuang, J. E. Gonzalez, I. Stoica, and E. P. Xing · 2023
Closest in time.
Towards automated circuit discovery for mechanistic interpretability, 2023
A. Conmy, A. N. Mavor-Parker, A. Lynch, S. Heimersheim, and A. Garriga-Alonso · 2023
Closest in time.
Finding neurons in a haystack: Case studies with sparse probing
W. Gurnee, N. Nanda, M. Pauly, K. Harvey, D. Troitskii, and D. Bertsimas · 2023
Closest in time.
Let’s verify step by step, 2023
H. Lightman, V. Kosaraju, Y. Burda, H. Edwards, B. Baker, T. Lee, J. Leike, J. Schulman, I. Sutskever, and K. Cobbe · 2023
Closest in time.
Clip-dissect: Automatic description of neuron representations in deep vision networks
T. Oikarinen and T.-W. Weng · 2023
Closest in time.
Gpt-4 technical report, 2023
OpenAI · 2023
Closest in time.
Explaining black box text modules in natural language with language models, 2023
C. Singh, A. R. Hsu, R. Antonello, S. Jain, A. G. Huth, B. Yu, and J. Gao · 2023
Closest in time.
Llama 2: Open foundation and fine-tuned chat models, 2023
H. Touvron, L. Martin, K. Stone, P. Albert, A. Almahairi, Y. Babaei, N. Bashlykov, S. Batra, P. Bhargava, S. Bhosale, D. Bikel, L. Blecher, C. C. Ferrer, M. Chen, G. Cucurull, D. Esiobu, J. Fernandes, J. Fu, W. Fu, B. Fuller, C. Gao, V. Goswami, N. Goyal, A. Hartshorn, S. Hosseini, R. Hou, H. Inan, M. Kardas, V. Kerkez, M. Khabsa, I. Kloumann, A. Korenev, P. S. Koura, M.-A. Lachaux, T. Lavril, J. Lee, D. Liskovich, Y. Lu, Y. Mao, X. Martinet, T. Mihaylov, P. Mishra, I. Molybog, Y. Nie, A. Poulton, J. Reizenstein, R. Rungta, K. Saladi, A. Schelten, R. Silva, E. M. Smith, R. Subramanian, X. E. Tan, B. Tang, R. Taylor, A. Williams, J. X. Kuan, P. Xu, Z. Yan, I. Zarov, Y. Zhang, A. Fan, M. Kambadur, S. Narang, A. Rodriguez, R. Stojnic, S. Edunov, and T. Scialom · 2023
Closest in time.
Tree of thoughts: Deliberate problem solving with large language models, 2023
S. Yao, D. Yu, J. Zhao, I. Shafran, T. L. Griffiths, Y. Cao, and K. Narasimhan · 2023
Closest in time.
Judging LLM-as-a-judge with MT-bench and chatbot arena
L. Zheng, W.-L. Chiang, Y. Sheng, S. Zhuang, Z. Wu, Y. Zhuang, Z. Lin, Z. Li, D. Li, E. Xing, et al · 2023
Closest in time.