Fetching the paper…
Reading the bibliography…
We respond to the recent paper by Makelov et al.
Parallel Distributed Processing. Volume 2: Psychological and Biological Models
J. L. McClelland, D. E. Rumelhart, and PDP Research Group (eds.) · 1986
Earlier work this paper cites.
Parallel Distributed Processing. Volume 1: Foundations
D. E. Rumelhart, J. L. McClelland, and PDP Research Group (eds.) · 1986
Earlier work this paper cites.
Neural and conceptual interpretation of PDP models
P. Smolensky · 1986
Earlier work this paper cites.
Neural natural language inference models partially embed theories of lexical entailment and negation
Atticus Geiger, Kyle Richardson, and Chris Potts · 2004
Earlier work this paper cites.
Investigating gender bias in language models using causal mediation analysis
Jesse Vig, Sebastian Gehrmann, Yonatan Belinkov, Sharon Qian, Daniel Nevo, Yaron Singer, and Stuart Shieber · 2004
Earlier work this paper cites.
Amnesic probing: Behavioral explanation with amnesic counterfactuals
Yanai Elazar, Shauli Ravfogel, Alon Jacovi, and Yoav Goldberg · 2006
Earlier work this paper cites.
Linguistic regularities in continuous space word representations
Tomas Mikolov, Wen-tau Yih, and Geoffrey Zweig · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Zoom in: An introduction to circuits
Chris Olah, Nick Cammarata, Ludwig Schubert, Gabriel Goh, Michael Petrov, and Shan Carter · 2020
Earlier work this paper cites.
Null it out: Guarding protected attributes by Iterative Nullspace Projection
Shauli Ravfogel, Yanai Elazar, Hila Gonen, Michael Twiton, and Yoav Goldberg · 2020
Earlier work this paper cites.
Are neural nets modular? inspecting functional modularity through differentiable weight masks
Róbert Csordás, Sjoerd van Steenkiste, and Jürgen Schmidhuber · 2021
Cited alongside, same era.
CausaLM: Causal Model Explanation Through Counterfactual Language Models
Amir Feder, Nadav Oved, Uri Shalit, and Roi Reichart · 2021
Cited alongside, same era.
Causal abstractions of neural networks
Atticus Geiger, Hanson Lu, Thomas Icard, and Christopher Potts · 2021
Cited alongside, same era.
Causal scrubbing: a method for rigorously testing interpretability hypotheses, 2022
Lawrence Chan, Adrià Garriga-Alonso, Nicholas Goldowsky-Dill, Ryan Greenblatt, Jenny Nitishinskaya, Ansh Radhakrishnan, Buck Shlegeris, and Nate Thomas · 2022
Cited alongside, same era.
Sparse interventions in language models with differentiable masking
Nicola De Cao, Leon Schmid, Dieuwke Hupkes, and Ivan Titov · 2022
Cited alongside, same era.
Towards monosemanticity: Decomposing language models with dictionary learning
Trenton Bricken, Adly Templeton, Joshua Batson, Brian Chen, Adam Jermyn, Tom Conerly, Nick Turner, Cem Anil, Carson Denison, Amanda Askell, Robert Lasenby, Yifan Wu, Shauna Kravec, Nicholas Schiefer, Tim Maxwell, Nicholas Joseph, Zac Hatfield-Dodds, Alex Tamkin, Karina Nguyen, Brayden McLean, Josiah E Burke, Tristan Hume, Shan Carter, Tom Henighan, and Christopher Olah · 2023
Later among the works it cites.
Atticus Geiger, Zhengxuan Wu, Christopher Potts, Thomas Icard, and Noah D. Goodman · 2023
Later among the works it cites.
Rigorously assessing natural language explanations of neurons
Jing Huang, Atticus Geiger, Karel D’Oosterlinck, Zhengxuan Wu, and Christopher Potts · 2023
Later among the works it cites.
An interpretability illusion for activation patching of arbitrary subspaces
Georg Lange, Alex Makelov, and Neel Nanda · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Nelson Elhage, Tristan Hume, Catherine Olsson, Nicholas Schiefer, Tom Henighan, Shauna Kravec, Zac Hatfield-Dodds, Robert Lasenby, Dawn Drain, Carol Chen, Roger Grosse, Sam McCandlish, Jared Kaplan, Dario Amodei, Martin Wattenberg, and Christopher Olah · 2022
Cited alongside, same era.
Inducing causal structure for interpretable neural networks
Atticus Geiger, Zhengxuan Wu, Hanson Lu, Josh Rozner, Elisa Kreiss, Thomas Icard, Noah Goodman, and Christopher Potts · 2022
Cited alongside, same era.
CEBaB: Estimating the causal effects of real-world concepts on NLP model behavior
Eldar David Abraham, Karel D’Oosterlinck, Amir Feder, Yair Ori Gat, Atticus Geiger, Christopher Potts, Roi Reichart, and Zhengxuan Wu · 2023
Cited alongside, same era.
Language models can explain neurons in language models
Steven Bills, Nick Cammarata, Dan Mossing, Henk Tillman, Leo Gao, Gabriel Goh, Ilya Sutskever, Jan Leike, Jeff Wu, and William Saunders · 2023
Cited alongside, same era.
Aleksandar Makelov, Georg Lange, and Neel Nanda · 2023
Later among the works it cites.
Emergent linear representations in world models of self-supervised sequence models
Neel Nanda, Andrew Lee, and Martin Wattenberg · 2023
Later among the works it cites.
The linear representation hypothesis and the geometry of large language models, 2023
Kiho Park, Yo Joong Choe, and Victor Veitch · 2023
Later among the works it cites.
Interpretability in the wild: A circuit for indirect object identification in GPT-2 small
Kevin Wang, Alexandre Variengien, Arthur Conmy, Buck Shlegeris, and Jacob Steinhardt · 2023
Later among the works it cites.
Interpretability at scale: Identifying causal mechanisms in Alpaca
Zhengxuan Wu, Atticus Geiger, Thomas Icard, Christopher Potts, and Noah Goodman · 2023
Later among the works it cites.