Fetching the paper…
Reading the bibliography…
We present a single attention head in GPT-2 Small that has one main role across the entire training distribution.
Compositional explanations of neurons
Jesse Mu and Jacob Andreas · 2006
Earlier work this paper cites.
Network dissection: Quantifying interpretability of deep visual representations, 2017
David Bau, Bolei Zhou, Aditya Khosla, Aude Oliva, and Antonio Torralba · 2017
Earlier work this paper cites.
Highway and residual networks learn unrolled iterative estimation, 2017
Klaus Greff, Rupesh K. Srivastava, and Jürgen Schmidhuber · 2017
Earlier work this paper cites.
Learning to generate reviews and discovering sentiment, 2017
Alec Radford, Rafal Jozefowicz, and Ilya Sutskever · 2017
Earlier work this paper cites.
Activation atlas
Shan Carter, Zan Armstrong, Ludwig Schubert, Ian Johnson, and Chris Olah · 2019
Earlier work this paper cites.
Openwebtext corpus, 2019
Aaron Gokaslan, Vanya Cohen, Ellie Pavlick, and Stefanie Tellex · 2019
Earlier work this paper cites.
A multiscale visualization of attention in the transformer model
Jesse Vig · 2019
Earlier work this paper cites.
Analyzing multi-head self-attention: Specialized heads do the heavy lifting, the rest can be pruned, 2019
Elena Voita, David Talbot, Fedor Moiseev, Rico Sennrich, and Ivan Titov · 2019
Earlier work this paper cites.
Thread: Circuits
Nick Cammarata, Shan Carter, Gabriel Goh, Chris Olah, Michael Petrov, Ludwig Schubert, Chelsea Voss, Ben Egan, and Swee Kiat Lim · 2020
Earlier work this paper cites.
Towards faithfully interpretable nlp systems: How should we define and evaluate faithfulness?, 2020
Alon Jacovi and Yoav Goldberg · 2020
Earlier work this paper cites.
interpreting gpt: the logit lens, 2020
nostalgebraist · 2020
Earlier work this paper cites.
An interpretability illusion for bert
Tolga Bolukbasi, Adam Pearce, Ann Yuan, Andy Coenen, Emily Reif, Fernanda Viégas, and Martin Wattenberg · 2021
Cited alongside, same era.
A mathematical framework for transformer circuits
Nelson Elhage, Neel Nanda, Catherine Olsson, Tom Henighan, Nicholas Joseph, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Nova DasSarma, Dawn Drain, Deep Ganguli, Zac Hatfield-Dodds, Danny Hernandez, Andy Jones, Jackson Kernion, Liane Lovitt, Kamal Ndousse, Dario Amodei, Tom Brown, Jack Clark, Jared Kaplan, Sam McCandlish, and Chris Olah · 2021
Cited alongside, same era.
Multimodal neurons in artificial neural networks
Gabriel Goh, Nick Cammarata, Chelsea Voss, Shan Carter, Michael Petrov, Ludwig Schubert, Alec Radford, and Chris Olah · 2021
Cited alongside, same era.
Causal scrubbing: A method for rigorously testing interpretability hypotheses
Lawrence Chan, Adria Garriga-Alonso, Nix Goldowsky-Dill, Ryan Greenblatt, Jenny Nitishinskaya, Ansh Radhakrishnan, Buck Shlegeris, and Nate Thomas · 2022
Cited alongside, same era.
Towards automated circuit discovery for mechanistic interpretability, 2023
Arthur Conmy, Augustine N. Mavor-Parker, Aengus Lynch, Stefan Heimersheim, and Adrià Garriga-Alonso · 2023
Closest in time.
Localizing model behavior with path patching, 2023
Nicholas Goldowsky-Dill, Chris MacLeod, Lucas Sato, and Aryaman Arora · 2023
Closest in time.
Finding neurons in a haystack: Case studies with sparse probing, 2023
Wes Gurnee, Neel Nanda, Matthew Pauly, Katherine Harvey, Dmitrii Troitskii, and Dimitris Bertsimas · 2023
Closest in time.
Overthinking the truth: Understanding how language models process false demonstrations, 2023
Danny Halawi, Jean-Stanislas Denain, and Jacob Steinhardt · 2023
Closest in time.
How does gpt-2 compute greater-than?: Interpreting mathematical abilities in a pre-trained language model, 2023
Michael Hanna, Ollie Liu, and Alexandre Variengien · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Guy Dar, Mor Geva, Ankit Gupta, and Jonathan Berant · 2022
Cited alongside, same era.
Transformerlens, 2022
Neel Nanda and Joseph Bloom · 2022
Cited alongside, same era.
Mechanistic interpretability, variables, and the importance of interpretable bases
Chris Olah · 2022
Cited alongside, same era.
In-context learning and induction heads, 2022
Catherine Olsson, Nelson Elhage, Neel Nanda, Nicholas Joseph, Nova DasSarma, Tom Henighan, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, et al · 2022
Cited alongside, same era.
Interpretability at scale: Identifying causal mechanisms in alpaca, 2023
Zhengxuan Wu, Atticus Geiger, Christopher Potts, and Noah D. Goodman · 2022
Cited alongside, same era.
Eliciting latent predictions from transformers with the tuned lens, 2023
Nora Belrose, Zach Furman, Logan Smith, Danny Halawi, Igor Ostrovsky, Lev McKinney, Stella Biderman, and Jacob Steinhardt · 2023
Cited alongside, same era.
Language models can explain neurons in language models
Steven Bills, Nick Cammarata, Dan Mossing, Henk Tillman, Leo Gao, Gabriel Goh, Ilya Sutskever, Jan Leike, Jeff Wu, and William Saunders · 2023
Cited alongside, same era.
Stefan Heimersheim and Jett Janiak · 2023
Closest in time.
Uncertainty in natural language processing: Sources, quantification, and applications, 2023
Mengting Hu, Zhen Zhang, Shiwan Zhao, Minlie Huang, and Bingzhe Wu · 2023
Closest in time.
The hydra effect: Emergent self-repair in language model computations, 2023
Thomas McGrath, Matthew Rahtz, Janos Kramar, Vladimir Mikulik, and Shane Legg · 2023
Closest in time.
Toward transparent ai: A survey on interpreting the inner structures of deep neural networks, 2023
Tilman Räuker, Anson Ho, Stephen Casper, and Dylan Hadfield-Menell · 2023
Closest in time.
Neurons in large language models: Dead, n-gram, positional, 2023
Elena Voita, Javier Ferrando, and Christoforos Nalmpantis · 2023
Closest in time.
Interpretability in the wild: a circuit for indirect object identification in GPT-2 small
Kevin Ro Wang, Alexandre Variengien, Arthur Conmy, Buck Shlegeris, and Jacob Steinhardt · 2023
Closest in time.