The building blocks of interpretability
Chris Olah, Arvind Satyanarayan, Ian Johnson, Shan Carter, Ludwig Schubert, Katherine Ye, and Alexander Mordvintsev · 2018
Later among the works it cites.
Towards the neural population doctrine
Shreya Saxena and John P Cunningham · 2019
Later among the works it cites.
Towards the neural population doctrine
Shreya Saxena and John P Cunningham · 2019
Later among the works it cites.
Approximation Theory and Approximation Practice, Extended Edition
Lloyd N. Trefethen · 2019
Later among the works it cites.
Reverse-engineering deep ReLU networks
David Rolnick and Konrad Kording · 2020
Later among the works it cites.
An Interpretability Illusion for BERT, April 2021
Original
Tolga Bolukbasi, Adam Pearce, Ann Yuan, Andy Coenen, Emily Reif, Fernanda Viégas, and Martin Wattenberg · 2021
Later among the works it cites.
Curve circuits
Nick Cammarata, Gabriel Goh, Shan Carter, Chelsea Voss, Ludwig Schubert, and Chris Olah · 2021
Later among the works it cites.
A mathematical framework for transformer circuits
Nelson Elhage, Neel Nanda, Catherine Olsson, Tom Henighan, Nicholas Joseph, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Nova DasSarma, Dawn Drain, Deep Ganguli, Zac Hatfield-Dodds, Danny Hernandez, Andy Jones, Jackson Kernion, Liane Lovitt, Kamal Ndousse, Dario Amodei, Tom Brown, Jack Clark, Jared Kaplan, Sam McCandlish, and Chris Olah · 2021
Later among the works it cites.
Multimodal neurons in artificial neural networks
Gabriel Goh, Nick Cammarata †, Chelsea Voss †, Shan Carter, Michael Petrov, Ludwig Schubert, Alec Radford, and Chris Olah · 2021
Later among the works it cites.
Knowledge Neurons in Pretrained Transformers, March 2022
Original
Damai Dai, Li Dong, Yaru Hao, Zhifang Sui, Baobao Chang, and Furu Wei · 2022
Closest in time.
Toy models of superposition
Nelson Elhage, Tristan Hume, Catherine Olsson, Nicholas Schiefer, Tom Henighan, Shauna Kravec, Zac Hatfield-Dodds, Robert Lasenby, Dawn Drain, Carol Chen, Roger Grosse, Sam McCandlish, Jared Kaplan, Dario Amodei, Martin Wattenberg, and Christopher Olah · 2022
Closest in time.
Origami in N dimensions: How feed-forward networks manufacture linear separability, March 2022
Original
Christian Keup and Moritz Helias · 2022
Closest in time.
Traversing the local polytopes of ReLU neural networks
Shaojie Xu, Joel Vaughan, Jie Chen, Aijun Zhang, and Agus Sudjianto · 2022
Closest in time.