Fetching the paper…
Reading the bibliography…
Inner Interpretability is a promising emerging field tasked with uncovering the inner mechanisms of AI systems, though how to develop these mechanistic theories is still much debated.
From understanding computation to understanding neural circuitry
Marr, D. and Poggio, T · 1976
Earlier work this paper cites.
Computational models and empirical constraints
Pylyshyn, Z. W · 1978
Earlier work this paper cites.
Computers and intractability
Garey, M. R. and Johnson, D. S · 1979
Earlier work this paper cites.
Systematic Parameterized Complexity Analysis in Computational Phonology
Wareham, H. T · 1998
Earlier work this paper cites.
The Computational Brain
Churchland, P. S. and Sejnowski, T · 1999
Earlier work this paper cites.
Thinking about mechanisms
Machamer, P., Darden, L., and Craver, C. F · 2000
Earlier work this paper cites.
The Human Connectome: A Structural Description of the Human Brain
Sporns, O., Tononi, G., and Kötter, R · 2005
Earlier work this paper cites.
When mechanistic models explain
Craver, C. F · 2006
Earlier work this paper cites.
Can cognitive processes be inferred from neuroimaging data?
Poldrack, R. A · 2006
Earlier work this paper cites.
Mental mechanisms: Philosophical perspectives on cognitive neuroscience
Bechtel, W · 2007
Earlier work this paper cites.
The cortical organization of speech processing
Hickok, G. and Poeppel, D · 2007
Earlier work this paper cites.
Computational complexity: a modern approach
Arora, S. and Barak, B · 2009
Earlier work this paper cites.
Probabilistic models of cognition: Exploring representations and inductive biases
Griffiths, T. L., Chater, N., Kemp, C., Perfors, A., and Tenenbaum, J. B · 2010
Earlier work this paper cites.
The explanatory force of dynamical and mathematical models in neuroscience: A mechanistic perspective
Kaplan, D. M. and Craver, C. F · 2011
Earlier work this paper cites.
Cortical oscillations and speech processing: emerging computational principles and operations
Giraud, A.-L. and Poeppel, D · 2012
Earlier work this paper cites.
The levels of understanding framework, revised
Poggio, T · 2012
Earlier work this paper cites.
Moving from levels & reduction to dimensions & constraints
Danks, D · 2013
Earlier work this paper cites.
Fundamentals of parameterized complexity
Downey, R. G. and Fellows, M. R · 2013
Earlier work this paper cites.
Severe tests in neuroimaging: what we can learn and how we can learn it
Aktunç, M. E · 2014
Earlier work this paper cites.
Marr’s levels revisited: understanding how brains break
Hardcastle, V. G. and Hardcastle, K · 2015
Earlier work this paper cites.
The maps problem and the mapping problem: two challenges for a cognitive neuroscience of speech and language
Poeppel, D · 2016
Earlier work this paper cites.
Causal feature learning: an overview
Chalupka, K., Eberhardt, F., and Perona, P · 2017
Earlier work this paper cites.
Towards a rigorous science of interpretable machine learning
Doshi-Velez, F. and Kim, B · 2017
Earlier work this paper cites.
Causal circuit explanations of behavior: Are necessity and sufficiency necessary and sufficient?
Gomez-Marin, A · 2017
Earlier work this paper cites.
Variations on a theme: species differences in synaptic connectivity do not predict central pattern generator activity
Gunaratne, C. A., Sakurai, A., and Katz, P. S · 2017
Earlier work this paper cites.
Abstraction hierarchy in deep learning neural networks
Ilin, R., Watson, T., and Kozma, R · 2017
Earlier work this paper cites.
Neuroscience needs behavior: correcting a reductionist bias
Krakauer, J. W., Ghazanfar, A. A., Gomez-Marin, A., MacIver, M. A., and Poeppel, D · 2017
Earlier work this paper cites.
Causal consistency of structural equation models
Rubenstein, P. K., Weichwald, S., Bongers, S., Mooij, J. M., Janzing, D., Grosse-Wentrup, M., and Schölkopf, B · 2017
Earlier work this paper cites.
Marr’s Computational Level and Delineating Phenomena
Shagrir, O. and Bechtel, W · 2017
Earlier work this paper cites.
Function-theoretic explanation and the search for neural mechanisms
Egan, F · 2018
Earlier work this paper cites.
Interpretability beyond feature attribution: Quantitative testing with concept activation vectors (tcav)
Kim, B., Wattenberg, M., Gilmer, J., Cai, C., Wexler, J., Viegas, F., et al · 2018
Earlier work this paper cites.
Statistical inference as severe testing: How to get beyond the statistics wars
Mayo, D. G · 2018
Earlier work this paper cites.
Abstracting causal models
Beckers, S. and Halpern, J. Y · 2019
Earlier work this paper cites.
A deep learning framework for neuroscience
Richards, B. A., Lillicrap, T. P., Beaudoin, P., Bengio, Y., Bogacz, R., Christensen, A., Clopath, C., Costa, R. P., de Berker, A., Ganguli, S., et al · 2019
Cited alongside, same era.
Deep learning for cognitive neuroscience
Storrs, K. R. and Kriegeskorte, N · 2019
Cited alongside, same era.
Augmenting self-attention with persistent memory
Sukhbaatar, S., Grave, E., Lample, G., Jegou, H., and Joulin, A · 2019
Cited alongside, same era.
Ten simple rules for the computational modeling of behavioral data
Wilson, R. C. and Collins, A. G · 2019
Cited alongside, same era.
The Brain–Cognitive Behavior Problem: A Retrospective
Buzsáki, G · 2020
Cited alongside, same era.
On the importance of severely testing deep learning models of cognition
Bowers, J. S., Malhotra, G., Adolfi, F., Dujmović, M., Montero, M. L., Biscione, V., Puebla, G., Hummel, J. H., and Heaton, R. F · 2023
Later among the works it cites.
Summing up the facts: Additive mechanisms behind factual recall in LLMs
Chughtai, B., Cooney, A., and Nanda, N · 2023
Later among the works it cites.
Towards automated circuit discovery for mechanistic interpretability
Conmy, A., Mavor-Parker, A. N., Lynch, A., Heimersheim, S., and Garriga-Alonso, A · 2023
Later among the works it cites.
Analyzing transformers in embedding space
Dar, G., Geva, M., Gupta, A., and Berant, J · 2023
Later among the works it cites.
Causal abstraction for faithful model interpretation
Geiger, A., Potts, C., and Icard, T · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Hamrick, J. and Mohamed, S · 2020
Cited alongside, same era.
How can we know what language models know?
Jiang, Z., Xu, F. F., Araki, J., and Neubig, G · 2020
Cited alongside, same era.
Against interpretability: a critical examination of the interpretability problem in machine learning
Krishnan, M · 2020
Cited alongside, same era.
Towards falsifiable interpretability research
Leavitt, M. L. and Morcos, A · 2020
Cited alongside, same era.
Resource-rational analysis: Understanding human cognition as the optimal use of limited computational resources
Lieder, F. and Griffiths, T. L · 2020
Cited alongside, same era.
Compositional explanations of neurons
Mu, J. and Andreas, J · 2020
Cited alongside, same era.
Against the epistemological primacy of the hardware: The brain from inside out, turned upside down
Poeppel, D. and Adolfi, F · 2020
Cited alongside, same era.
Geva, M., Bastings, J., Filippova, K., and Globerson, A · 2023
Later among the works it cites.
Finding neurons in a haystack: Case studies with sparse probing
Gurnee, W., Nanda, N., Pauly, M., Harvey, K., Troitskii, D., and Bertsimas, D · 2023
Later among the works it cites.
Inspecting and editing knowledge representations in language models
Hernandez, E., Li, B. Z., and Andreas, J · 2023
Later among the works it cites.
Uncovering intermediate variables in transformers using circuit probing
Lepori, M. A., Serre, T., and Pavlick, E · 2023
Later among the works it cites.
Testing methods of neural systems understanding
Lindsay, G. W. and Bau, D · 2023
Later among the works it cites.
Counting Carbon: A Survey of Factors Influencing the Emissions of Machine Learning, February 2023
Luccioni, A. S. and Hernandez-Garcia, A · 2023
Later among the works it cites.
Power Hungry Processing: Watts Driving the Cost of AI Deployment?, November 2023
Luccioni, A. S., Jernite, Y., and Strubell, E · 2023
Later among the works it cites.
Copy suppression: Comprehensively understanding an attention head
McDougall, C., Conmy, A., Rushing, C., McGrath, T., and Nanda, N · 2023
Later among the works it cites.
The debate over understanding in AI’s large language models
Mitchell, M. and Krakauer, D. C · 2023
Later among the works it cites.
The ConceptARC Benchmark: Evaluating Understanding and Generalization in the ARC Domain, May 2023
Moskvichev, A., Odouard, V. V., and Mitchell, M · 2023
Later among the works it cites.
Interpretability dreams, 2023
Olah, C · 2023
Later among the works it cites.
Toward transparent ai: A survey on interpreting the inner structures of deep neural networks
Räuker, T., Ho, A., Casper, S., and Hadfield-Menell, D · 2023
Later among the works it cites.
Average-Hard Attention Transformers are Constant-Depth Uniform Threshold Circuits, August 2023
Strobl, L · 2023
Later among the works it cites.
Transformers as Recognizers of Formal Languages: A Survey on Expressivity, October 2023
Strobl, L., Merrill, W., Weiss, G., Chiang, D., and Angluin, D · 2023
Later among the works it cites.
Getting aligned on representational alignment
Sucholutsky, I., Muttenthaler, L., Weller, A., Peng, A., Bobu, A., Kim, B., Love, B. C., Grant, E., Achterberg, J., Tenenbaum, J. B., et al · 2023
Later among the works it cites.
Analyzing vision transformers for image classification in class embedding space
Vilas, M. G., Schaumlöffel, T., and Roig, G · 2023
Later among the works it cites.
Characterizing mechanisms for factual recall in language models
Yu, Q., Merullo, J., and Pavlick, E · 2023
Later among the works it cites.
Instilling inductive biases with subnetworks
Zhang, E., Lepori, M. A., and Pavlick, E · 2023
Later among the works it cites.
The clock and the pizza: Two stories in mechanistic explanation of neural networks
Zhong, Z., Liu, Z., Tegmark, M., and Andreas, J · 2023
Later among the works it cites.
Representation engineering: A top-down approach to ai transparency
Zou, A., Phan, L., Chen, S., Campbell, J., Guo, P., Ren, R., Pan, A., Yin, X., Mazeika, M., Dombrowski, A.-K., et al · 2023
Later among the works it cites.
Complexity-theoretic limits on the promises of artificial neural network reverse-engineering
Adolfi, F., Vilas, M., and Wareham, T · 2024
Closest in time.
In-context language learning: Architectures and algorithms, 2024
Akyürek, E., Wang, B., Kim, Y., and Andreas, J · 2024
Closest in time.
Into the LAION’s den: Investigating hate in multimodal datasets
Birhane, A., Han, S., Boddeti, V., Luccioni, S., et al · 2024
Closest in time.
Black-box access is insufficient for rigorous ai audits
Casper, S., Ezell, C., Siegmann, C., Kolt, N., Curtis, T. L., Bucknall, B., Haupt, A., Wei, K., Scheurer, J., Hobbhahn, M., Sharkey, L., Krishna, S., Hagen, M. V., Alberti, S., Chan, A., Sun, Q., Gerovitch, M., Bau, D., Tegmark, M., Krueger, D., and Hadfield-Menell, D · 2024
Closest in time.
Does localization inform editing? surprising differences in causality-based localization vs. knowledge editing in language models
Hase, P., Bansal, M., Kim, B., and Ghandeharioun, A · 2024
Closest in time.
Causation in neuroscience: keeping mechanism meaningful
Ross, L. N. and Bassett, D. S · 2024
Closest in time.
Function vectors in large language models
Todd, E., Li, M. L., Sharma, A. S., Mueller, A., Wallace, B. C., and Bau, D · 2024
Closest in time.