Finding neurons in a haystack: Case studies with sparse probing
Wes Gurnee, Neel Nanda, Matthew Pauly, Katherine Harvey, Dmitrii Troitskii, and Dimitris Bertsimas. 2023 · 2023
Later among the works it cites.
How does GPT-2 compute greater-than?: Interpreting mathematical abilities in a pre-trained language model
Michael Hanna, Ollie Liu, and Alexandre Variengien. 2023 · 2023
Later among the works it cites.
Does localization inform editing? surprising differences in causality-based localization vs. knowledge editing in language models
Peter Hase, Mohit Bansal, Been Kim, and Asma Ghandeharioun. 2023 · 2023
Later among the works it cites.
In-context learning creates task vectors
Roee Hendel, Mor Geva, and Amir Globerson. 2023 · 2023
Later among the works it cites.
Rigorously assessing natural language explanations of neurons
Jing Huang, Atticus Geiger, Karel D’Oosterlinck, Zhengxuan Wu, and Christopher Potts. 2023 · 2023
Later among the works it cites.
Does circuit analysis interpretability scale? Evidence from multiple choice capabilities in chinchilla
Original
Tom Lieberum, Matthew Rahtz, János Kramár, Neel Nanda, Geoffrey Irving, Rohin Shah, and Vladimir Mikulik. 2023 · 2023
Later among the works it cites.
The geometry of truth: Emergent linear structure in large language model representations of true/false datasets
Original
Samuel Marks and Max Tegmark. 2023 · 2023
Later among the works it cites.
A mechanism for solving relational tasks in transformer language models
Original
Jack Merullo, Carsten Eickhoff, and Ellie Pavlick. 2023 · 2023
Later among the works it cites.
Almanacs: A simulatability benchmark for language model explainability
Original
Edmund Mills, Shiye Su, Stuart Russell, and Scott Emmons. 2023 · 2023
Later among the works it cites.
The linear representation hypothesis and the geometry of large language models
Original
Kiho Park, Yo Joong Choe, and Victor Veitch. 2023 · 2023
Later among the works it cites.
Find: A function description benchmark for evaluating interpretability methods
Sarah Schwettmann, Tamar Rott Shaham, Joanna Materzynska, Neil Chowdhury, Shuang Li, Jacob Andreas, David Bau, and Antonio Torralba. 2023 · 2023
Later among the works it cites.
Codebook features: Sparse and discrete interpretability for neural networks
Original
Alex Tamkin, Mohammad Taufeeque, and Noah D. Goodman. 2023 · 2023
Later among the works it cites.
Linear representations of sentiment in large language models
Original
Curt Tigges, Oskar John Hollinsworth, Atticus Geiger, and Neel Nanda. 2023 · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
Original
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Dan Bikel, Lukas Blecher, Cristian Canton Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fernandes, Jeremy Fu, Wenyin Fu, Brian Fuller, Cynthia Gao, Vedanuj Goswami, Naman Goyal, Anthony Hartshorn, Saghar Hosseini, Rui Hou, Hakan Inan, Marcin Kardas, Viktor Kerkez, Madian Khabsa, Isabel Kloumann, Artem Korenev, Punit Singh Koura, Marie-Anne Lachaux, Thibaut Lavril, Jenya Lee, Diana Liskovich, Yinghai Lu, Yuning Mao, Xavier Martinet, Todor Mihaylov, Pushkar Mishra, Igor Molybog, Yixin Nie, Andrew Poulton, Jeremy Reizenstein, Rashi Rungta, Kalyan Saladi, Alan Schelten, Ruan Silva, Eric Michael Smith, Ranjan Subramanian, Xiaoqing Ellen Tan, Binh Tang, Ross Taylor, Adina Williams, Jian Xiang Kuan, Puxin Xu, Zheng Yan, Iliyan Zarov, Yuchen Zhang, Angela Fan, Melanie Kambadur, Sharan Narang, Aurelien Rodriguez, Robert Stojnic, Sergey Edunov, and Thomas Scialom. 2023 · 2023
Later among the works it cites.
Interpretability in the wild: a circuit for indirect object identification in GPT-2 small
Kevin Ro Wang, Alexandre Variengien, Arthur Conmy, Buck Shlegeris, and Jacob Steinhardt. 2023 · 2023
Later among the works it cites.
Interpretability at scale: Identifying causal mechanisms in alpaca
Original
Zhengxuan Wu, Atticus Geiger, Thomas Icard, Christopher Potts, and Noah D. Goodman. 2023 · 2023
Later among the works it cites.
MQuAKE: Assessing knowledge editing in language models via multi-hop questions
Zexuan Zhong, Zhengxuan Wu, Christopher Manning, Christopher Potts, and Danqi Chen. 2023 · 2023
Later among the works it cites.
Evaluating the ripple effects of knowledge editing in language models
Original
Roi Cohen, Eden Biran, Ori Yoran, Amir Globerson, and Mor Geva. 2024 · 2024
Closest in time.
Sparse autoencoders find highly interpretable features in language models
Original
Hoagy Cunningham, Aidan Ewart, Logan Riggs, Robert Huben, and Lee Sharkey. 2024 · 2024
Closest in time.
How do language models bind entities in context?
Jiahai Feng and Jacob Steinhardt. 2024 · 2024
Closest in time.
Patchscopes: A unifying framework for inspecting hidden representations of language models
Original
Asma Ghandeharioun, Avi Caciularu, Adam Pearce, Lucas Dixon, and Mor Geva. 2024 · 2024
Closest in time.
Linearity of relation decoding in transformer language models
Evan Hernandez, Arnab Sen Sharma, Tal Haklay, Kevin Meng, Martin Wattenberg, Jacob Andreas, Yonatan Belinkov, and David Bau. 2024 · 2024
Closest in time.
Function vectors in large language models
Original
Eric Todd, Millicent L. Li, Arnab Sen Sharma, Aaron Mueller, Byron C. Wallace, and David Bau. 2024 · 2024
Closest in time.