Fetching the paper…
Reading the bibliography…
Many proposed applications of neural networks in machine learning, cognitive/brain science, and society hinge on the feasibility of inner interpretability via circuit discovery.
Computational Complexity of Probabilistic Turing Machines
John Gill · 1977
Earlier work this paper cites.
Computers and intractability
Michael R Garey and David S Johnson · 1979
Earlier work this paper cites.
The Complexity of Enumeration and Reliability Problems
Leslie G. Valiant · 1979
Earlier work this paper cites.
The Complexity of Counting Cuts and of Computing the Probability that a Graph is Connected
J. Scott Provan and Michael O. Ball · 1983
Earlier work this paper cites.
A new polynomial-time algorithm for linear programming
N. Karmarkar · 1984
Earlier work this paper cites.
The complexity of optimization functions
MW Krentel · 1988
Earlier work this paper cites.
Optimization, approximation, and complexity classes
C.H. Papadimitriou and M. Yannakakis · 1991
Earlier work this paper cites.
PP is as Hard as the Polynomial-Time Hierarchy
Seinosuke Toda · 1991
Earlier work this paper cites.
OptP as the normal behavior of NP-complete problems
William I. Gasarch, Mark W. Krentel, and Kevin J. Rappoport · 1995
Earlier work this paper cites.
Randomized Algorithms
Rajeev Motwani and Prabhakar Raghavan · 1995
Earlier work this paper cites.
Proof verification and the hardness of approximation problems
Sanjeev Arora, Carsten Lund, Rajeev Motwani, Madhu Sudan, and Mario Szegedy · 1998
Earlier work this paper cites.
Complexity and Approximation
Giorgio Ausiello, Alberto Marchetti-Spaccamela, Pierluigi Crescenzi, Giorgio Gambosi, Marco Protasi, and Viggo Kann · 1999
Earlier work this paper cites.
Parameterized Complexity
R. G. Downey and M. R. Fellows · 1999
Earlier work this paper cites.
Systematic Parameterized Complexity Analysis in Computational Phonology
Harold T. Wareham · 1999
Earlier work this paper cites.
Completeness in the polynomial-time hierarchy: A compendium
Marcus Schaefer and Christopher Umans · 2002
Earlier work this paper cites.
Parameterized Complexity Theory
Jörg Flum and Martin Grohe · 2006
Earlier work this paper cites.
Computational Complexity: A Modern Approach
Sanjeev Arora and Boaz Barak · 2009
Earlier work this paper cites.
The status of the P versus NP problem
Lance Fortnow · 2009
Earlier work this paper cites.
SIGACT News Complexity Theory Column 76: An atypical survey of typical-case heuristic algorithms
Lane A. Hemaspaandra and Ryan Williams · 2012
Earlier work this paper cites.
Fundamentals of Parameterized Complexity
Rod G. Downey and Michael R. Fellows · 2013
Earlier work this paper cites.
On the Computational Efficiency of Training Neural Networks
Roi Livni, Shai Shalev-Shwartz, and Ohad Shamir · 2014
Earlier work this paper cites.
On the NP-Completeness of the Minimum Circuit Size Problem
John M. Hitchcock and A. Pavan · 2015
Earlier work this paper cites.
Parameterized complexity classes beyond para-NP
Ronald de Haan and Stefan Szeider · 2017
Earlier work this paper cites.
Feature Visualization
Chris Olah, Alexander Mordvintsev, and Ludwig Schubert · 2017
Earlier work this paper cites.
On the Complexity of Learning Neural Networks
Le Song, Santosh Vempala, John Wilmes, and Bo Xie · 2017
Earlier work this paper cites.
Attention is All you Need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
The Constant Inapproximability of the Parameterized Dominating Set Problem
Yijia Chen and Bingkai Lin · 2019
Earlier work this paper cites.
An Algorithmic Barrier to Neural Circuit Understanding
Venkatakrishnan Ramaswamy · 2019
Earlier work this paper cites.
Model Interpretability through the lens of Computational Complexity
Pablo Barceló, Mikaël Monet, Jorge Pérez, and Bernardo Subercaseaux · 2020
Earlier work this paper cites.
Understanding the role of individual units in a deep neural network
David Bau, Jun-Yan Zhu, Hendrik Strobelt, Agata Lapedriza, Bolei Zhou, and Antonio Torralba · 2020
Earlier work this paper cites.
Thread: Circuits
Nick Cammarata, Shan Carter, Gabriel Goh, Chris Olah, Michael Petrov, Ludwig Schubert, Chelsea Voss, Ben Egan, and Swee Kiat Lim · 2020
Earlier work this paper cites.
Learning Deep ReLU Networks Is Fixed-Parameter Tractable
Sitan Chen, Adam R. Klivans, and Raghu Meka · 2020
Cited alongside, same era.
Are there any ‘object detectors’ in the hidden layers of CNNs trained to identify objects or scenes?
Ella M. Gale, Nicholas Martin, Ryan Blything, Anh Nguyen, and Jeffrey S. Bowers · 2020
Cited alongside, same era.
Towards Faithfully Interpretable NLP Systems: How Should We Define and Evaluate Faithfulness?
Alon Jacovi and Yoav Goldberg · 2020
Cited alongside, same era.
Zoom In: An Introduction to Circuits
Chris Olah, Nick Cammarata, Ludwig Schubert, Gabriel Goh, Michael Petrov, and Shan Carter · 2020
Cited alongside, same era.
Model Interpretability through the Lens of Computational Complexity
Bernardo Aníbal Subercaseaux · 2020
Cited alongside, same era.
What does the Knowledge Neuron Thesis Have to do with Knowledge?
Jingcheng Niu, Andrew Liu, Zining Zhu, and Gerald Penn · 2023
Later among the works it cites.
Speech language models lack important brain-relevant semantics
Subba Reddy Oota, Emin Çelik, Fatma Deniz, and Mariya Toneva · 2023
Later among the works it cites.
The Parameterized Complexity of Finding Concise Local Explanations
Sebastian Ordyniak, Giacomo Paesani, and Stefan Szeider · 2023
Later among the works it cites.
Symbols and grounding in large language models
Ellie Pavlick · 2023
Later among the works it cites.
Toward Transparent AI: A Survey on Interpreting the Inner Structures of Deep Neural Networks
Tilman Räuker, Anson Ho, Stephen Casper, and Dylan Hadfield-Menell · 2023
Later among the works it cites.
Attribution Patching Outperforms Automated Circuit Discovery
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Armin Biere, Marijn J. H. Heule, Hans van Maaren, and Toby Walsh (eds.) · 2021
Cited alongside, same era.
Transformer Feed-Forward Layers Are Key-Value Memories
Mor Geva, Roei Schuster, Jonathan Berant, and Omer Levy · 2021
Cited alongside, same era.
What Do Compressed Deep Neural Networks Forget?
Sara Hooker, Aaron Courville, Gregory Clark, Yann Dauphin, and Andrea Frome · 2021
Cited alongside, same era.
MLP-Mixer: An all-MLP Architecture for Vision
Ilya O Tolstikhin, Neil Houlsby, Alexander Kolesnikov, Lucas Beyer, Xiaohua Zhai, Thomas Unterthiner, Jessica Yung, Andreas Steiner, Daniel Keysers, Jakob Uszkoreit, Mario Lucic, and Alexey Dosovitskiy · 2021
Cited alongside, same era.
Visualizing Weights
Chelsea Voss, Nick Cammarata, Gabriel Goh, Michael Petrov, Ludwig Schubert, Ben Egan, Swee Kiat Lim, and Chris Olah · 2021
Cited alongside, same era.
The computational complexity of understanding binary classifier decisions
Stephan Wäldchen, Jan Macdonald, Sascha Hauch, and Gitta Kutyniok · 2021
Cited alongside, same era.
Knowledge Neurons in Pretrained Transformers
Damai Dai, Li Dong, Yaru Hao, Zhifang Sui, Baobao Chang, and Furu Wei · 2022
Cited alongside, same era.
Aaquib Syed, Can Rager, and Arthur Conmy · 2023
Later among the works it cites.
Towards Best Practices of Activation Patching in Language Models: Metrics and Methods
Fred Zhang and Neel Nanda · 2023
Later among the works it cites.
Complexity-Theoretic Limits on the Promises of Artificial Neural Network Reverse-Engineering
Federico Adolfi, Martina G. Vilas, and Todd Wareham · 2024
Closest in time.
Local vs. Global Interpretability: A Computational Complexity Perspective
Shahaf Bassan, Guy Amir, and Guy Katz · 2024
Closest in time.
Mechanistic interpretability for AI safety - a review
Leonard Bereska and Stratis Gavves · 2024
Closest in time.
What Are Large Language Models Mapping to in the Brain? A Case Against Over-Reliance on Brain Scores
Ebrahim Feghhi, Nima Hadidi, Bryan Song, Idan A. Blank, and Jonathan C. Kao · 2024
Closest in time.
Information Flow Routes: Automatically Interpreting Language Models at Scale
Javier Ferrando and Elena Voita · 2024
Closest in time.
Interpretability Illusions in the Generalization of Simplified Models
Dan Friedman, Andrew Lampinen, Lucas Dixon, Danqi Chen, and Asma Ghandeharioun · 2024
Closest in time.
Detecting and Understanding Vulnerabilities in Language Models via Mechanistic Interpretability
Jorge García-Carrasco, Alejandro Maté, and Juan Trujillo · 2024
Closest in time.
Causal Abstraction: A Theoretical Foundation for Mechanistic Interpretability
Atticus Geiger, Duligur Ibeling, Amir Zur, Maheep Chaudhary, Sonakshi Chauhan, Jing Huang, Aryaman Arora, Zhengxuan Wu, Noah Goodman, Christopher Potts, and Thomas Icard · 2024
Closest in time.
Patchscopes: A Unifying Framework for Inspecting Hidden Representations of Language Models
Asma Ghandeharioun, Avi Caciularu, Adam Pearce, Lucas Dixon, and Mor Geva · 2024
Closest in time.
Have Faith in Faithfulness: Going Beyond Circuit Overlap When Finding Model Mechanisms
Michael Hanna, Sandro Pezzelle, and Yonatan Belinkov · 2024
Closest in time.
LLM-assisted Concept Discovery: Automatically Identifying and Explaining Neuron Functions
Nhat Hoang-Xuan, Minh Vu, and My T. Thai · 2024
Closest in time.
AtP*: An efficient and scalable method for localizing LLM behaviour to components
János Kramár, Tom Lieberum, Rohin Shah, and Neel Nanda · 2024
Closest in time.
Learned feature representations are biased by complexity, learning order, position, and more
Andrew Kyle Lampinen, Stephanie C. Y. Chan, and Katherine Hermann · 2024
Closest in time.
Grounding neuroscience in behavioral changes using artificial neural networks
Grace W Lindsay · 2024
Closest in time.
Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language Models
Samuel Marks, Can Rager, Eric J. Michaud, Yonatan Belinkov, David Bau, and Aaron Mueller · 2024
Closest in time.
Locating and editing factual associations in GPT
Kevin Meng, David Bau, Alex Andonian, and Yonatan Belinkov · 2024
Closest in time.
Circuit Component Reuse Across Tasks in Transformer Language Models
Jack Merullo, Carsten Eickhoff, and Ellie Pavlick · 2024
Closest in time.
Causation in neuroscience: Keeping mechanism meaningful
Lauren N. Ross and Dani S. Bassett · 2024
Closest in time.
When Representations Align: Universality in Representation Learning Dynamics
Loek Van Rossem and Andrew M. Saxe · 2024
Closest in time.
Hypothesis Testing the Circuit Hypothesis in LLMs
Claudia Shi, Nicolas Beltran-Velez, Achille Nazaret, Carolina Zheng, Adrià Garriga-Alonso, Andrew Jesson, Maggie Makar, and David Blei · 2024
Closest in time.
LLM Circuit Analyses Are Consistent Across Training and Scale
Curt Tigges, Michael Hanna, Qinan Yu, and Stella Biderman · 2024
Closest in time.
Does Editing Provide Evidence for Localization?
Zihao Wang and Victor Veitch · 2024
Closest in time.
Instilling Inductive Biases with Subnetworks
Enyan Zhang, Michael A. Lepori, and Ellie Pavlick · 2024
Closest in time.