Fetching the paper…
Reading the bibliography…
Automated mechanistic interpretation research has attracted great interest due to its potential to scale explanations of neural network internals to large models.
Causal mediation analysis for interpreting neural nlp: The case of gender bias
Jesse Vig, Sebastian Gehrmann, Yonatan Belinkov, Sharon Qian, Daniel Nevo, Yaron Singer, and Stuart M. Shieber · 2004
Earlier work this paper cites.
Language models are few-shot learners
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeff Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Ma teusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei · 2005
Earlier work this paper cites.
Causality: Models, Reasoning and Inference
Judea Pearl · 2009
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Beyond word importance: Contextual decomposition to extract interactions from LSTMs
W. James Murdoch, Peter J. Liu, and Bin Yu · 2018
Earlier work this paper cites.
Hierarchical interpretations for neural network predictions
Chandan Singh, W. James Murdoch, and Bin Yu · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Earlier work this paper cites.
Analysing neural language models: Contextual decomposition reveals default reasoning in number and gender assignment
Jaap Jumelet, Willem H. Zuidema, and Dieuwke Hupkes · 2019
Earlier work this paper cites.
Hierarchical interpretations for neural network predictions
Chandan Singh, W. James Murdoch, and Bin Yu · 2019
Earlier work this paper cites.
Causalm: Causal model explanation through counterfactual language models
Amir Feder, Nadav Oved, Uri Shalit, and Roi Reichart · 2020
Earlier work this paper cites.
Transformer feed-forward layers are key-value memories
Mor Geva, R. Schuster, Jonathan Berant, and Omer Levy · 2020
Earlier work this paper cites.
Zoom in: An introduction to circuits
Christopher Olah, Nick Cammarata, Ludwig Schubert, Gabriel Goh, Michael Petrov, and Shan Carter · 2020
Earlier work this paper cites.
A mathematical framework for transformer circuits
Nelson Elhage, Neel Nanda, Catherine Olsson, Tom Henighan, Nicholas Joseph, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Nova DasSarma, Dawn Drain, Deep Ganguli, Zac Hatfield-Dodds, Danny Hernandez, Andy Jones, Jackson Kernion, Liane Lovitt, Kamal Ndousse, Dario Amodei, Tom Brown, Jack Clark, Jared Kaplan, Sam McCandlish, and Chris Olah · 2021
Earlier work this paper cites.
Causal abstractions of neural networks
Atticus Geiger, Hanson Lu, Thomas F. Icard, and Christopher Potts · 2021
Cited alongside, same era.
Highly accurate protein structure prediction with alphafold
John M. Jumper, Richard Evans, Alexander Pritzel, Tim Green, Michael Figurnov, Olaf Ronneberger, Kathryn Tunyasuvunakool, Russ Bates, Augustin Zídek, Anna Potapenko, Alex Bridgland, Clemens Meyer, Simon A A Kohl, Andy Ballard, Andrew Cowie, Bernardino Romera-Paredes, Stanislav Nikolov, Rishub Jain, Jonas Adler, Trevor Back, Stig Petersen, David A. Reiman, Ellen Clancy, Michal Zielinski, Martin Steinegger, Michalina Pacholska, Tamas Berghammer, Sebastian Bodenstein, David Silver, Oriol Vinyals, Andrew W. Senior, Koray Kavukcuoglu, Pushmeet Kohli, and Demis Hassabis · 2021
Cited alongside, same era.
Causal distillation for language models
Zhengxuan Wu, Atticus Geiger, Josh Rozner, Elisa Kreiss, Hanson Lu, Thomas F. Icard, Christopher Potts, and Noah D. Goodman · 2021
Cited alongside, same era.
Beyond atoms and bonds: contextual explainability via molecular graphical depictions
Marco Bertolini, Linlin Zhao, Djork-Arné Clevert, and Floriane Montanari · 2022
Cited alongside, same era.
Toy models of superposition
Tracr: Compiled transformers as a laboratory for interpretability
David Lindner, J’anos Kram’ar, Matthew Rahtz, Tom McGrath, and Vladimir Mikulik · 2023
Later among the works it cites.
Locating and editing factual associations in gpt, 2023
Kevin Meng, David Bau, Alex Andonian, and Yonatan Belinkov · 2023
Later among the works it cites.
Progress measures for grokking via mechanistic interpretability
Neel Nanda, Lawrence Chan, Tom Lieberum, Jess Smith, and Jacob Steinhardt · 2023
Later among the works it cites.
Toward transparent ai: A survey on interpreting the inner structures of deep neural networks, 2023
Tilman Räuker, Anson Ho, Stephen Casper, and Dylan Hadfield-Menell · 2023
Later among the works it cites.
A mechanistic interpretation of arithmetic reasoning in language models using causal mediation analysis
Alessandro Stolfo, Yonatan Belinkov, and Mrinmaya Sachan · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Nelson Elhage, Tristan Hume, Catherine Olsson, Nicholas Schiefer, Tom Henighan, Shauna Kravec, Zac Hatfield-Dodds, Robert Lasenby, Dawn Drain, Carol Chen, Roger Grosse, Sam McCandlish, Jared Kaplan, Dario Amodei, Martin Wattenberg, and Christopher Olah · 2022
Cited alongside, same era.
Transformers in medical image analysis
Kelei He, Chen Gan, Zhuoyuan Li, Islem Rekik, Zihao Yin, Wen Ji, Yang Gao, Qian Wang, Junfeng Zhang, and Dinggang Shen · 2022
Cited alongside, same era.
Causal machine learning: A survey and open problems
Jean Kaddour, Aengus Lynch, Qi Liu, Matt J. Kusner, and Ricardo Silva · 2022
Cited alongside, same era.
Towards automated circuit discovery for mechanistic interpretability
Arthur Conmy, Augustine N. Mavor-Parker, Aengus Lynch, Stefan Heimersheim, and Adrià Garriga-Alonso · 2023
Cited alongside, same era.
How do language models bind entities in context?
Jiahai Feng and Jacob Steinhardt · 2023
Cited alongside, same era.
Localizing model behavior with path patching
Nicholas W. Goldowsky-Dill, Chris MacLeod, Lucas Jun Koba Sato, and Aryaman Arora · 2023
Cited alongside, same era.
Michael Hanna, Ollie Liu, and Alexandre Variengien · 2023
Cited alongside, same era.
A circuit for python docstrings in a 4-layer attention-only transformer
Stefan Heimersheim and Jett Janiak · 2023
Cited alongside, same era.
Later among the works it cites.
Attribution patching outperforms automated circuit discovery, 2023
Aaquib Syed, Can Rager, and Arthur Conmy · 2023
Later among the works it cites.
Interpretability in the wild: a circuit for indirect object identification in GPT-2 small
Kevin Ro Wang, Alexandre Variengien, Arthur Conmy, Buck Shlegeris, and Jacob Steinhardt · 2023
Later among the works it cites.
Have faith in faithfulness: Going beyond circuit overlap when finding model mechanisms, 2024
Michael Hanna, Sandro Pezzelle, and Yonatan Belinkov · 2024
Closest in time.
Zhengfu He, Xuyang Ge, Qiong Tang, Tianxiang Sun, Qinyuan Cheng, and Xipeng Qiu · 2024
Closest in time.
Sparse feature circuits: Discovering and editing interpretable causal graphs in language models
Samuel Marks, Can Rager, Eric J. Michaud, Yonatan Belinkov, David Bau, and Aaron Mueller · 2024
Closest in time.
Sparse autoencoders enable scalable and reliable circuit identification in language models, 2024
Charles O’Neill and Thang Bui · 2024
Closest in time.
Hypothesis testing the circuit hypothesis in LLMs
Claudia Shi, Nicolas Beltran-Velez, Achille Nazaret, Carolina Zheng, Adrià Garriga-Alonso, Andrew Jesson, Maggie Makar, and David Blei · 2024
Closest in time.