Fetching the paper…
Reading the bibliography…
Localizing behaviors of neural networks to a subset of the network's components or a subset of interactions between components is a natural first step towards analyzing network mechanisms and possible failure modes.
Judea Pearl · 2013
Earlier work this paper cites.
Residual networks behave like ensembles of relatively shallow networks
Andreas Veit, Michael J Wilber, and Serge Belongie · 2016
Earlier work this paper cites.
This looks like that: deep learning for interpretable image recognition
Chaofan Chen, Oscar Li, Daniel Tao, Alina Barnett, Cynthia Rudin, and Jonathan K Su · 2019
Earlier work this paper cites.
Approximate causal abstractions
Sander Beckers, Frederick Eberhardt, and Joseph Y Halpern · 2020
Earlier work this paper cites.
Neural natural language inference models partially embed theories of lexical entailment and negation
Atticus Geiger, Kyle Richardson, and Christopher Potts · 2020
Earlier work this paper cites.
Zoom in: An introduction to circuits
Chris Olah, Nick Cammarata, Ludwig Schubert, Gabriel Goh, Michael Petrov, and Shan Carter · 2020
Earlier work this paper cites.
Shortformer: Better language modeling using shorter inputs
Ofir Press, Noah A Smith, and Mike Lewis · 2020
Earlier work this paper cites.
Investigating gender bias in language models using causal mediation analysis
Jesse Vig, Sebastian Gehrmann, Yonatan Belinkov, Sharon Qian, Daniel Nevo, Yaron Singer, and Stuart Shieber · 2020
Earlier work this paper cites.
An interpretability illusion for bert
Tolga Bolukbasi, Adam Pearce, Ann Yuan, Andy Coenen, Emily Reif, Fernanda Viégas, and Martin Wattenberg · 2021
Earlier work this paper cites.
A mathematical framework for transformer circuits
Nelson Elhage, Neel Nanda, Catherine Olsson, Tom Henighan, Nicholas Joseph, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Nova DasSarma, Dawn Drain, Deep Ganguli, Zac Hatfield-Dodds, Danny Hernandez, Andy Jones, Jackson Kernion, Liane Lovitt, Kamal Ndousse, Dario Amodei, Tom Brown, Jack Clark, Jared Kaplan, Sam McCandlish, and Chris Olah · 2021
Cited alongside, same era.
Causal analysis of syntactic agreement mechanisms in neural language models
Matthew Finlayson, Aaron Mueller, Sebastian Gehrmann, Stuart Shieber, Tal Linzen, and Yonatan Belinkov · 2021
Cited alongside, same era.
Causal abstractions of neural networks
Atticus Geiger, Hanson Lu, Thomas Icard, and Christopher Potts · 2021
Cited alongside, same era.
Causal scrubbing: a method for rigorously testing interpretability hypotheses., 2022
Lawrence Chan, Adrià Garriga-Alonso, Nicholas Goldowsky-Dill, Ryan Greenblatt, Jenny Nitishinskaya, Ansh Radhakrishnan, Buck Shlegeris, and Nate Thomas · 2022
Cited alongside, same era.
Softmax linear units
Nelson Elhage, Tristan Hume, Catherine Olsson, Neel Nanda, Tom Henighan, Scott Johnston, Sheer ElShowk, Nicholas Joseph, Nova DasSarma, Ben Mann, Danny Hernandez, Amanda Askell, Kamal Ndousse, Andy Jones, Dawn Drain, Anna Chen, Yuntao Bai, Deep Ganguli, Liane Lovitt, Zac Hatfield-Dodds, Jackson Kernion, Tom Conerly, Shauna Kravec, Stanislav Fort, Saurav Kadavath, Josh Jacobson, Eli Tran-Johnson, Jared Kaplan, Jack Clark, Tom Brown, Sam McCandlish, Dario Amodei, and Christopher Olah · 2022
Scale efficiently: Insights from pre-training and fine-tuning transformers, 2022
Yi Tay, Mostafa Dehghani, Jinfeng Rao, William Fedus, Samira Abnar, Hyung Won Chung, Sharan Narang, Dani Yogatama, Ashish Vaswani, and Donald Metzler · 2022
Later among the works it cites.
Interpretability in the wild: a circuit for indirect object identification in gpt-2 small, 2022
Kevin Wang, Alexandre Variengien, Arthur Conmy, Buck Shlegeris, and Jacob Steinhardt · 2022
Later among the works it cites.
Eliciting latent predictions from transformers with the tuned lens, 2023
Nora Belrose, Zach Furman, Logan Smith, Danny Halawi, Igor Ostrovsky, Lev McKinney, Stella Biderman, and Jacob Steinhardt · 2023
Closest in time.
A toy model of universality: Reverse engineering how networks learn group operations, 2023
Bilal Chughtai, Lawrence Chan, and Neel Nanda · 2023
Closest in time.
Towards automated circuit discovery for mechanistic interpretability, 2023
Arthur Conmy, Augustine N. Mavor-Parker, Aengus Lynch, Stefan Heimersheim, and Adrià Garriga-Alonso · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Inducing causal structure for interpretable neural networks
Atticus Geiger, Zhengxuan Wu, Hanson Lu, Josh Rozner, Elisa Kreiss, Thomas Icard, Noah Goodman, and Christopher Potts · 2022
Cited alongside, same era.
Locating and editing factual associations in gpt
Kevin Meng, David Bau, Alex Andonian, and Yonatan Belinkov · 2022
Cited alongside, same era.
In-context learning and induction heads
Catherine Olsson, Nelson Elhage, Neel Nanda, Nicholas Joseph, Nova DasSarma, T. J. Henighan, Benjamin Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Dawn Drain, Deep Ganguli, Zac Hatfield-Dodds, Danny Hernandez, Scott Johnston, Andy Jones, John Kernion, Liane Lovitt, Kamal Ndousse, Dario Amodei, Tom B. Brown, Jack Clark, Jared Kaplan, Sam McCandlish, and Christopher Olah · 2022
Cited alongside, same era.
Causal abstraction for faithful model interpretation
Atticus Geiger, Chris Potts, and Thomas Icard
Cited in the paper.
Finding alignments between interpretable causal variables and distributed neural representations
Atticus Geiger, Zhengxuan Wu, Christopher Potts, Thomas Icard, and Noah D Goodman
Cited in the paper.
Tracr: Compiled transformers as a laboratory for interpretability, 2023
David Lindner, János Kramár, Matthew Rahtz, Thomas McGrath, and Vladimir Mikulik · 2023
Closest in time.
Toward transparent ai: A survey on interpreting the inner structures of deep neural networks, 2023
Tilman Räuker, Anson Ho, Stephen Casper, and Dylan Hadfield-Menell · 2023
Closest in time.
Practical pitfalls of causal scrubbing, 2023
Jérémy Scheurer Scheurer, Phil3, tony, Jacques Thibodeau, and David Lindner · 2023
Closest in time.