Fetching the paper…
Reading the bibliography…
Causal abstraction provides a theoretical foundation for mechanistic interpretability, the field concerned with providing intelligible algorithms that are faithful simplifications of the known, but opaque low-level details of black box AI models.
Explaining classifiers with causal concept effect (cace)
Yash Goyal, Uri Shalit, and Been Kim · 1907
Earlier work this paper cites.
Aggregation of variables in dynamic systems
Herbert A. Simon and Albert Ando · 1961
Earlier work this paper cites.
Abstract interpretation: a unified lattice model for static analysis of programs by construction or approximation of fixpoints
P. Cousot and R. Cousot · 1977
Earlier work this paper cites.
Parallel Distributed Processing. Volume 2: Psychological and Biological Models
J. L. McClelland, D. E. Rumelhart, and PDP Research Group, editors · 1986
Earlier work this paper cites.
Parallel Distributed Processing. Volume 1: Foundations
D. E. Rumelhart, J. L. McClelland, and PDP Research Group, editors · 1986
Earlier work this paper cites.
Neural and conceptual interpretation of PDP models
Paul Smolensky · 1986
Earlier work this paper cites.
Local vs. distributed coding
Simon Thorpe · 1989
Earlier work this paper cites.
Mental causation
Stephen Yablo · 1992
Earlier work this paper cites.
Causality and model abstraction
Yumi Iwasaki and Herbert A. Simon · 1994
Earlier work this paper cites.
Rule learning by seven-month-old infants
Gary F Marcus, Sugumaran Vijayan, S Bandi Rao, and Peter M Vishton · 1999
Earlier work this paper cites.
Causation, Prediction, and Search
Peter Spirtes, Clark Glymour, and Richard Scheines · 2000
Earlier work this paper cites.
Making Things Happen: A Theory of Causal Explanation
James Woodward · 2003
Earlier work this paper cites.
Interventions and causal inference
Frederick Eberhardt and Richard Scheines · 2007
Earlier work this paper cites.
Identifying dynamic sequential plans
Jin Tian · 2008
Earlier work this paper cites.
Causality
Judea Pearl · 2009
Earlier work this paper cites.
A general approach to causal mediation analysis
Kosuke Imai, Luke Keele, and Dustin Tingley · 2010
Earlier work this paper cites.
Causal mediation analysis
Raymond Hicks and Dustin Tingley · 2011
Earlier work this paper cites.
Linguistic regularities in continuous space word representations
Tomas Mikolov, Wen-tau Yih, and Geoffrey Zweig · 2013
Earlier work this paper cites.
Striving for simplicity: The all convolutional net
Jost Springenberg, Alexey Dosovitskiy, Thomas Brox, and Martin Riedmiller · 2014
Earlier work this paper cites.
Visualizing and understanding convolutional networks
Matthew D. Zeiler and Rob Fergus · 2014
Earlier work this paper cites.
Visual causal feature learning
Krzysztof Chalupka, Pietro Perona, and Frederick Eberhardt · 2015
Earlier work this paper cites.
Layer-wise relevance propagation for neural networks with local renormalization layers
Alexander Binder, Grégoire Montavon, Sebastian Bach, Klaus-Robert Müller, and Wojciech Samek · 2016
Earlier work this paper cites.
Man is to computer programmer as woman is to homemaker? Debiasing word embeddings
Tolga Bolukbasi, Kai-Wei Chang, James Y Zou, Venkatesh Saligrama, and Adam T Kalai · 2016
Earlier work this paper cites.
Unsupervised discovery of el nino using causal feature learning on microlevel climate data
Krzysztof Chalupka, Tobias Bischoff, Frederick Eberhardt, and Pietro Perona · 2016
Earlier work this paper cites.
”why should i trust you?”: Explaining the predictions of any classifier
Marco Tulio Ribeiro, Sameer Singh, and Carlos Guestrin · 2016
Earlier work this paper cites.
Not just a black box: Learning important features through propagating activation differences
Avanti Shrikumar, Peyton Greenside, Anna Shcherbina, and Anshul Kundaje · 2016
Earlier work this paper cites.
Causal feature learning: an overview
Krzysztof Chalupka, Frederick Eberhardt, and Pietro Perona · 2017
Earlier work this paper cites.
CausaLM: Causal Model Explanation Through Counterfactual Language Models
Amir Feder, Nadav Oved, Uri Shalit, and Roi Reichart · 2017
Earlier work this paper cites.
A unified approach to interpreting model predictions
Scott M Lundberg and Su-In Lee · 2017
Earlier work this paper cites.
Causal consistency of structural equation models
Paul K. Rubenstein, Sebastian Weichwald, Stephan Bongers, Joris M. Mooij, Dominik Janzing, Moritz Grosse-Wentrup, and Bernhard Schölkopf · 2017
Earlier work this paper cites.
Axiomatic attribution for deep networks
Mukund Sundararajan, Ankur Taly, and Qiqi Yan · 2017
Earlier work this paper cites.
What you can cram into a single $&!#* vector: Probing sentence embeddings for linguistic properties
Alexis Conneau, German Kruszewski, Guillaume Lample, Loïc Barrault, and Marco Baroni · 2018
Earlier work this paper cites.
Under the hood: Using diagnostic classifiers to investigate and improve how language models track agreement information
Mario Giulianelli, Jack Harding, Florian Mohnert, Dieuwke Hupkes, and Willem H. Zuidema · 2018
Earlier work this paper cites.
Causal learning and explanation of deep neural networks via autoencoded activations
Michael Harradon, Jeff Druce, and Brian E. Ruttenberg · 2018
Earlier work this paper cites.
Visualisation and ’diagnostic classifiers’ reveal how recurrent and recursive neural networks process hierarchical structure
Dieuwke Hupkes, Sara Veldhoen, and Willem H. Zuidema · 2018
Earlier work this paper cites.
Interpretability beyond feature attribution: Quantitative testing with concept activation vectors (tcav), 2018
Been Kim, Martin Wattenberg, Justin Gilmer, Carrie Cai, James Wexler, Fernanda Viegas, and Rory Sayres · 2018
Earlier work this paper cites.
The mythos of model interpretability
Zachary C. Lipton · 2018
Earlier work this paper cites.
Explaining deep learning models using causal inference
Tanmayee Narendra, Anush Sankaran, Deepak Vijaykeerthy, and Senthil Mani · 2018
Earlier work this paper cites.
Dissecting contextual word embeddings: Architecture and representation
Matthew E. Peters, Mark Neumann, Luke Zettlemoyer, and Wen-tau Yih · 2018
Earlier work this paper cites.
A review of computational models of basic rule learning: The neural-symbolic debate and beyond
Raquel G Alhama and Willem Zuidema · 2019
Earlier work this paper cites.
Visualizing and understanding gans
David Bau, Jun-Yan Zhu, Hendrik Strobelt, Bolei Zhou, Joshua B. Tenenbaum, William T. Freeman, and Antonio Torralba · 2019
Earlier work this paper cites.
Abstracting causal models
Sander Beckers and Joseph Halpern · 2019
Earlier work this paper cites.
Approximate causal abstractions
Sander Beckers, Frederick Eberhardt, and Joseph Y. Halpern · 2019
Earlier work this paper cites.
Neural network attributions: A causal perspective
Aditya Chattopadhyay, Piyushi Manupriya, Anirban Sarkar, and Vineeth N Balasubramanian · 2019
Earlier work this paper cites.
What does BERT look at? An analysis of BERT’s attention
Kevin Clark, Urvashi Khandelwal, Omer Levy, and Christopher D. Manning · 2019
Earlier work this paper cites.
Designing and interpreting probes with control tasks
John Hewitt and Percy Liang · 2019
Earlier work this paper cites.
Parameter-efficient transfer learning for NLP
Neil Houlsby, Andrei Giurgiu, Stanislaw Jastrzebski, Bruna Morrone, Quentin de Laroussilhe, Andrea Gesmundo, Mona Attariyan, and Sylvain Gelly · 2019
Earlier work this paper cites.
On the explanatory depth and pragmatic value of coarse-grained, probabilistic, causal explanations
David Kinney · 2019
Earlier work this paper cites.
Consistent individualized feature attribution for tree ensembles, 2019
Scott M. Lundberg, Gabriel G. Erion, and Su-In Lee · 2019
Earlier work this paper cites.
Are sixteen heads really better than one?, 2019
Paul Michel, Omer Levy, and Graham Neubig · 2019
Earlier work this paper cites.
The limitations of opaque learning machines
Judea Pearl · 2019
Earlier work this paper cites.
BERT rediscovers the classical NLP pipeline
Ian Tenney, Dipanjan Das, and Ellie Pavlick · 2019
Earlier work this paper cites.
Attention is not not explanation
Sarah Wiegreffe and Yuval Pinter · 2019
Earlier work this paper cites.
A calculus for stochastic interventions: Causal effect identification and surrogate experiments
Elias Bareinboim and Juan D. Correa · 2020
Earlier work this paper cites.
Counterfactuals uncover the modular structure of deep generative models
Michel Besserve, Arash Mehrjou, Rémy Sun, and Bernhard Schölkopf · 2020
Cited alongside, same era.
Thread: Circuits
Nick Cammarata, Shan Carter, Gabriel Goh, Chris Olah, Michael Petrov, Ludwig Schubert, Chelsea Voss, Ben Egan, and Swee Kiat Lim · 2020
Cited alongside, same era.
Transparency in complex computational systems
Kathleen A. Creel · 2020
Cited alongside, same era.
How do decisions emerge across layers in neural models? interpretation with differentiable masking
Nicola De Cao, Michael Sejr Schlichtkrull, Wilker Aziz, and Ivan Titov · 2020
Cited alongside, same era.
Amnesic probing: Behavioral explanation with amnesic counterfactuals
Yanai Elazar, Shauli Ravfogel, Alon Jacovi, and Yoav Goldberg · 2020
Cited alongside, same era.
Neural natural language inference models partially embed theories of lexical entailment and negation
Dissecting recall of factual associations in auto-regressive language models
Mor Geva, Jasmijn Bastings, Katja Filippova, and Amir Globerson · 2023
Closest in time.
Localizing model behavior with path patching, 2023
Nicholas Goldowsky-Dill, Chris MacLeod, Lucas Sato, and Aryaman Arora · 2023
Closest in time.
Finding neurons in a haystack: Case studies with sparse probing
Wes Gurnee, Neel Nanda, Matthew Pauly, Katherine Harvey, Dmitrii Troitskii, and Dimitris Bertsimas · 2023
Closest in time.
How does GPT-2 compute greater-than?: Interpreting mathematical abilities in a pre-trained language model
Michael Hanna, Ollie Liu, and Alexandre Variengien · 2023
Closest in time.
Does localization inform editing? surprising differences in causality-based localization vs. knowledge editing in language models
Peter Hase, Mohit Bansal, Been Kim, and Asma Ghandeharioun · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Atticus Geiger, Kyle Richardson, and Christopher Potts · 2020
Cited alongside, same era.
Towards faithfully interpretable NLP systems: How should we define and evaluate faithfulness?
Alon Jacovi and Yoav Goldberg · 2020
Cited alongside, same era.
Concept bottleneck models
Pang Wei Koh, Thao Nguyen, Yew Siang Tang, Stephen Mussmann, Emma Pierson, Been Kim, and Percy Liang · 2020
Cited alongside, same era.
Zoom in: An introduction to circuits
Chris Olah, Nick Cammarata, Ludwig Schubert, Gabriel Goh, Michael Petrov, and Shan Carter · 2020
Cited alongside, same era.
Information-theoretic probing for linguistic structure
Tiago Pimentel, Josef Valvoda, Rowan Hall Maudslay, Ran Zmigrod, Adina Williams, and Ryan Cotterell · 2020
Cited alongside, same era.
Null it out: Guarding protected attributes by iterative nullspace projection
Shauli Ravfogel, Yanai Elazar, Hila Gonen, Michael Twiton, and Yoav Goldberg · 2020
Cited alongside, same era.
Probing the probing paradigm: Does probing accuracy entail task relevance?, 2020
Abhilasha Ravichander, Yonatan Belinkov, and Eduard Hovy · 2020
Cited alongside, same era.
John Hewitt, John Thickstun, Christopher D. Manning, and Percy Liang · 2023
Closest in time.
Rigorously assessing natural language explanations of neurons
Jing Huang, Atticus Geiger, Karel D’Oosterlinck, Zhengxuan Wu, and Christopher Potts · 2023
Closest in time.
Inducing character-level structure in subword-based language models with type-level interchange intervention training
Jing Huang, Zhengxuan Wu, Kyle Mahowald, and Christopher Potts · 2023
Closest in time.
A survey of algorithmic recourse: Contrastive explanations and consequential recommendations
Amir-Hossein Karimi, Gilles Barthe, Bernhard Schölkopf, and Isabel Valera · 2023
Closest in time.
Targeted reduction of causal models
Armin Kekic, Bernhard Schölkopf, and Michel Besserve · 2023
Closest in time.
Break it down: Evidence for structural compositionality in neural networks
Michael A. Lepori, Thomas Serre, and Ellie Pavlick · 2023
Closest in time.
Inference-time intervention: Eliciting truthful answers from a language model
Kenneth Li, Oam Patel, Fernanda Viégas, Hanspeter Pfister, and Martin Wattenberg · 2023
Closest in time.
Tom Lieberum, Matthew Rahtz, János Kramár, Neel Nanda, Geoffrey Irving, Rohin Shah, and Vladimir Mikulik · 2023
Closest in time.
The geometry of truth: Emergent linear structure in large language model representations of true/false datasets, 2023
Samuel Marks and Max Tegmark · 2023
Closest in time.
Causal abstraction with soft interventions
Riccardo Massidda, Atticus Geiger, Thomas Icard, and Davide Bacciu · 2023
Closest in time.
Mass-editing memory in a transformer
Kevin Meng, Arnab Sen Sharma, Alex J. Andonian, Yonatan Belinkov, and David Bau · 2023
Closest in time.
Progress measures for grokking via mechanistic interpretability
Neel Nanda, Lawrence Chan, Tom Lieberum, Jess Smith, and Jacob Steinhardt · 2023
Closest in time.
Emergent linear representations in world models of self-supervised sequence models
Neel Nanda, Andrew Lee, and Martin Wattenberg · 2023
Closest in time.
The linear representation hypothesis and the geometry of large language models
Kiho Park, Yo Joong Choe, and Victor Veitch · 2023
Closest in time.
Log-linear guardedness and its implications
Shauli Ravfogel, Yoav Goldberg, and Ryan Cotterell · 2023
Closest in time.
A mechanistic interpretation of arithmetic reasoning in language models using causal mediation analysis
Alessandro Stolfo, Yonatan Belinkov, and Mrinmaya Sachan · 2023
Closest in time.
Linear representations of sentiment in large language models, 2023
Curt Tigges, Oskar John Hollinsworth, Atticus Geiger, and Neel Nanda · 2023
Closest in time.
Activation addition: Steering language models without optimization, 2023
Alexander Matt Turner, Lisa Thiergart, David Udell, Gavin Leech, Ulisse Mini, and Monte MacDiarmid · 2023
Closest in time.
Interpretability in the wild: a circuit for indirect object identification in GPT-2 small
Kevin Ro Wang, Alexandre Variengien, Arthur Conmy, Buck Shlegeris, and Jacob Steinhardt · 2023
Closest in time.
Interpretability at scale: Identifying causal mechanisms in alpaca
Zhengxuan Wu, Atticus Geiger, Thomas Icard, Christopher Potts, and Noah Goodman · 2023
Closest in time.
Post-hoc concept bottleneck models
Mert Yüksekgönül, Maggie Wang, and James Zou · 2023
Closest in time.
Jointly learning consistent causal abstractions over multiple interventional distributions
Fabio Massimo Zennaro, Máté Drávucz, Geanina Apachitei, Widanalage Dhammika Widanage, and Theodoros Damoulas · 2023
Closest in time.
Quantifying consistency and information loss for causal abstraction learning
Fabio Massimo Zennaro, Paolo Turrini, and Theodoros Damoulas · 2023
Closest in time.
Representation engineering: A top-down approach to AI transparency
Andy Zou, Long Phan, Sarah Chen, James Campbell, Phillip Guo, Richard Ren, Alexander Pan, Xuwang Yin, Mantas Mazeika, Ann-Kathrin Dombrowski, et al · 2023
Closest in time.
CausalGym: Benchmarking causal interpretability methods on linguistic tasks
Aryaman Arora, Dan Jurafsky, and Christopher Potts · 2024
Closest in time.
Recurrent neural networks learn to store and generate sequences using non-linear representations
Róbert Csordás, Christopher Potts, Christopher D Manning, and Atticus Geiger · 2024
Closest in time.
How do language models bind entities in context?
Jiahai Feng and Jacob Steinhardt · 2024
Closest in time.
Monitoring latent world states in language models with propositional probes
Jiahai Feng, Stuart Russell, and Jacob Steinhardt · 2024
Closest in time.
Patchscopes: A unifying framework for inspecting hidden representations of language models
Asma Ghandeharioun, Avi Caciularu, Adam Pearce, Lucas Dixon, and Mor Geva · 2024
Closest in time.
Interchange intervention training applied to post-meal glucose prediction for type 1 diabetes mellitus patients
Ana Esponera Gómez and Giovanni Cinà · 2024
Closest in time.
How to use and interpret activation patching, 2024
Stefan Heimersheim and Neel Nanda · 2024
Closest in time.
Ravel: Evaluating interpretability methods on disentangling language model representations, 2024
Jing Huang, Zhengxuan Wu, Christopher Potts, Mor Geva, and Atticus Geiger · 2024
Closest in time.
Sparse autoencoders find highly interpretable features in language models
Robert Huben, Hoagy Cunningham, Logan Riggs Smith, Aidan Ewart, and Lee Sharkey · 2024
Closest in time.
On the origins of linear representations in large language models
Yibo Jiang, Goutham Rajendran, Pradeep Ravikumar, Bryon Aragam, and Victor Veitch · 2024
Closest in time.
KAN: kolmogorov-arnold networks
Ziming Liu, Yixuan Wang, Sachin Vaidya, Fabian Ruehle, James Halverson, Marin Soljacic, Thomas Y. Hou, and Max Tegmark · 2024
Closest in time.
Sparse feature circuits: Discovering and editing interpretable causal graphs in language models, 2024
Samuel Marks, Can Rager, Eric J. Michaud, Yonatan Belinkov, David Bau, and Aaron Mueller · 2024
Closest in time.
Learning causal abstractions of linear structural causal models
Riccardo Massidda, Sara Magliacane, and Davide Bacciu · 2024
Closest in time.
Controllable context sensitivity and the knob behind it, 2024
Julian Minder, Kevin Du, Niklas Stoehr, Giovanni Monea, Chris Wendler, Robert West, and Ryan Cotterell · 2024
Closest in time.
Aaron Mueller, Jannik Brinkmann, Millicent Li, Samuel Marks, Koyena Pal, Nikhil Prakash, Can Rager, Aruna Sankaranarayanan, Arnab Sen Sharma, Jiuding Sun, Eric Todd, David Bau, and Yonatan Belinkov · 2024
Closest in time.
Fine-tuning enhances existing mechanisms: A case study on entity tracking
Nikhil Prakash, Tamar Rott Shaham, Tal Haklay, Yonatan Belinkov, and David Bau · 2024
Closest in time.
Characterizing the role of similarity in the property inferences of language models
Juan Diego Rodriguez, Aaron Mueller, and Kanishka Misra · 2024
Closest in time.
Naomi Saphra and Sarah Wiegreffe · 2024
Closest in time.
Codebook features: Sparse and discrete interpretability for neural networks, 2024
Alex Tamkin, Mohammad Taufeeque, and Noah Goodman · 2024
Closest in time.
Function vectors in large language models
Eric Todd, Millicent L. Li, Arnab Sen Sharma, Aaron Mueller, Byron C. Wallace, and David Bau · 2024
Closest in time.
repeng, 2024
Theia Vogel · 2024
Closest in time.
Towards best practices of activation patching in language models: Metrics and methods
Fred Zhang and Neel Nanda · 2024
Closest in time.
Updating CLIP to prefer descriptions over captions
Amir Zur, Elisa Kreiss, Karel D’Oosterlinck, Christopher Potts, and Atticus Geiger · 2024
Closest in time.
Emergent symbol-like number variables in artificial neural networks, 2025
Satchel Grant, Noah D. Goodman, and James L. McClelland · 2025
Closest in time.