Fetching the paper…
Reading the bibliography…
Understanding the function of individual neurons within language models is essential for mechanistic interpretability research.
Compositional explanations of neurons
Jesse Mu and Jacob Andreas · 2006
Earlier work this paper cites.
Concrete problems in ai safety
Dario Amodei, Chris Olah, Jacob Steinhardt, Paul Christiano, John Schulman, and Dan Mané · 2016
Earlier work this paper cites.
Real time image saliency for black box classifiers
Piotr Dabkowski and Yarin Gal · 2017
Earlier work this paper cites.
Learning to generate reviews and discovering sentiment
Alec Radford, Rafal Jozefowicz, and Ilya Sutskever · 2017
Earlier work this paper cites.
Identifying and controlling important neurons in neural machine translation
Anthony Bau, Yonatan Belinkov, Hassan Sajjad, Nadir Durrani, Fahim Dalvi, and James R. Glass · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding, 2018
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Earlier work this paper cites.
Nlize: A perturbation-driven visual interrogation tool for analyzing and interpreting natural language inference models
Shusen Liu, Zhimin Li, Tao Li, Vivek Srikumar, Valerio Pascucci, and Peer-Timo Bremer · 2018
Earlier work this paper cites.
Interpretable textual neuron representations for NLP
Nina Poerner, Benjamin Roth, and Hinrich Schütze · 2018
Earlier work this paper cites.
What is one grain of sand in the desert? analyzing individual neurons in deep NLP models
Fahim Dalvi, Nadir Durrani, Hassan Sajjad, Yonatan Belinkov, Anthony Bau, and James Glass · 2019
Earlier work this paper cites.
Nlp augmentation
Edward Ma · 2019
Cited alongside, same era.
Analyzing individual neurons in pre-trained language models
Nadir Durrani, Hassan Sajjad, Fahim Dalvi, and Yonatan Belinkov · 2020
Cited alongside, same era.
The Pile: An 800gb dataset of diverse text for language modeling
Leo Gao, Stella Biderman, Sid Black, Laurence Golding, Travis Hoppe, Charles Foster, Jason Phang, Horace He, Anish Thite, Noa Nabeshima, Shawn Presser, and Connor Leahy · 2020
Cited alongside, same era.
Transformer feed-forward layers are key-value memories
Mor Geva, Roei Schuster, Jonathan Berant, and Omer Levy · 2020
Cited alongside, same era.
Zoom in: An introduction to circuits
Chris Olah, Nick Cammarata, Ludwig Schubert, Gabriel Goh, Michael Petrov, and Shan Carter · 2020
Cited alongside, same era.
Knowledge neurons in pretrained transformers
Damai Dai, Li Dong, Yaru Hao, Zhifang Sui, and Furu Wei · 2021
Later among the works it cites.
Multimodal neurons in artificial neural networks
Gabriel Goh, Nick Cammarata, Chelsea Voss, Shan Carter, Michael Petrov, Ludwig Schubert, Alec Radford, and Chris Olah · 2021
Later among the works it cites.
Unsolved problems in ml safety
Dan Hendrycks, Nicholas Carlini, John Schulman, and Jacob Steinhardt · 2021
Later among the works it cites.
High-low frequency detectors
Ludwig Schubert, Chelsea Voss, Nick Cammarata, Gabriel Goh, and Chris Olah · 2021
Later among the works it cites.
Implicit representations of event properties within contextual language models: Searching for “causativity neurons”
Esther Seyffarth, Younes Samih, Laura Kallmeyer, and Hassan Sajjad · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Lucas Torroba Hennigen, Adina Williams, and Ryan Cotterell · 2020
Cited alongside, same era.
Investigating Gender Bias in Language Models Using Causal Mediation Analysis
Jesse Vig, Sebastian Gehrmann, Yonatan Belinkov, Sharon Qian, Daniel Nevo, Yaron Singer, and Stuart Shieber · 2020
Cited alongside, same era.
Similarity analysis of contextual word representation models
John Wu, Yonatan Belinkov, Hassan Sajjad, Nadir Durrani, Fahim Dalvi, and James Glass · 2020
Cited alongside, same era.
An interpretability illusion for bert
Tolga Bolukbasi, Adam Pearce, Ann Yuan, Andy Coenen, Emily Reif, Fernanda Viégas, and Martin Wattenberg
Cited in the paper.
An interpretability illusion for BERT
Tolga Bolukbasi, Adam Pearce, Ann Yuan, Andy Coenen, Emily Reif, Fernanda B. Viégas, and Martin Wattenberg
Cited in the paper.
Toy models of superposition, 2022b
Nelson Elhage, Tristan Hume, Catherine Olsson, Nicholas Schiefer, Tom Henighan, Shauna Kravec, Zac Hatfield-Dodds, Robert Lasenby, Dawn Drain, Carol Chen, Roger Grosse, Sam McCandlish, Jared Kaplan, Dario Amodei, Martin Wattenberg, and Christopher Olah
Cited in the paper.
Neuroscope, 2022a
Neel Nanda
Cited in the paper.
System III: Learning with domain knowledge for safety constraints
Fazl Barez, Hosein Hasanbeig, and Alessandro Abate · 2022
Later among the works it cites.
Softmax linear units
Nelson Elhage, Tristan Hume, Catherine Olsson, Neel Nanda, Tom Henighan, Scott Johnston, Sheer ElShowk, Nicholas Joseph, Nova DasSarma, Ben Mann, Danny Hernandez, Amanda Askell, Kamal Ndousse, And Jones, Dawn Drain, Anna Chen, Yuntao Bai, Deep Ganguli, Liane Lovitt, Zac Hatfield-Dodds, Jackson Kernion, Tom Conerly, Shauna Kravec, Stanislav Fort, Saurav Kadavath, Josh Jacobson, Eli Tran-Johnson, Jared Kaplan, Jack Clark, Tom Brown, Sam McCandlish, Dario Amodei, and Christopher Olah · 2022
Later among the works it cites.
Mechanistic interpretability, variables, and the importance of interpretable bases
Chris Olah · 2022
Later among the works it cites.