Fetching the paper…
Reading the bibliography…
Language models learn a great quantity of factual information during pretraining, and recent work localizes this information to specific model weights like mid-layer MLP weights.
Towards automatic concept-based explanations
Amirata Ghorbani, James Wexler, James Y Zou, and Been Kim · 1902
Earlier work this paper cites.
The emergence of number and syntax units in lstm language models
Yair Lakretz, German Kruszewski, Theo Desbordes, Dieuwke Hupkes, Stanislas Dehaene, and Marco Baroni · 1903
Earlier work this paper cites.
Causal mediation analysis for interpreting neural nlp: The case of gender bias
Jesse Vig, Sebastian Gehrmann, Yonatan Belinkov, Sharon Qian, Daniel Nevo, Simas Sakenis, Jason Huang, Yaron Singer, and Stuart Shieber · 2004
Earlier work this paper cites.
Explaining neural networks by decoding layer activations
Johannes Schneider and Michalis Vlachos · 2005
Earlier work this paper cites.
Mechanisms for handling nested dependencies in neural-network language models and humans
Yair Lakretz, Dieuwke Hupkes, Alessandra Vergallito, Marco Marelli, Marco Baroni, and Stanislas Dehaene · 2006
Earlier work this paper cites.
Invertible concept-based explanations for cnn models with non-negative concept activation vectors
Ruihan Zhang, Prashan Madumal, Tim Miller, Krista A Ehinger, and Benjamin IP Rubinstein · 2006
Earlier work this paper cites.
Rewriting a deep generative model
David Bau, Steven Liu, Tongzhou Wang, Jun-Yan Zhu, and Antonio Torralba · 2007
Earlier work this paper cites.
Are neural nets modular? inspecting functional modularity through differentiable weight masks
Róbert Csordás, Sjoerd van Steenkiste, and Jürgen Schmidhuber · 2010
Earlier work this paper cites.
Transformer feed-forward layers are key-value memories
Mor Geva, Roei Schuster, Jonathan Berant, and Omer Levy · 2012
Earlier work this paper cites.
Modifying memories in transformer models
Chen Zhu, Ankit Singh Rawat, Manzil Zaheer, Srinadh Bhojanapalli, Daliang Li, Felix Yu, and Sanjiv Kumar · 2012
Earlier work this paper cites.
Efficient estimation of word representations in vector space
Tomas Mikolov, Kai Chen, Greg Corrado, and Jeffrey Dean · 2013
Earlier work this paper cites.
Visualizing and understanding convolutional networks
Matthew D Zeiler and Rob Fergus · 2014
Earlier work this paper cites.
Zero-shot relation extraction via reading comprehension
Omer Levy, Minjoon Seo, Eunsol Choi, and Luke Zettlemoyer · 2017
Earlier work this paper cites.
Learning to generate reviews and discovering sentiment
Alec Radford, Rafal Jozefowicz, and Ilya Sutskever · 2017
Cited alongside, same era.
Interpretability beyond feature attribution: Quantitative testing with concept activation vectors (TCAV)
Been Kim, Martin Wattenberg, Justin Gilmer, Carrie Cai, James Wexler, Fernanda Viegas, and Rory sayres · 2018
Cited alongside, same era.
The building blocks of interpretability
Chris Olah, Arvind Satyanarayan, Ian Johnson, Shan Carter, Ludwig Schubert, Katherine Ye, and Alexander Mordvintsev · 2018
Cited alongside, same era.
Language models as knowledge bases?
Fabio Petroni, Tim Rocktäschel, Sebastian Riedel, Patrick Lewis, Anton Bakhtin, Yuxiang Wu, and Alexander Miller · 2019
Cited alongside, same era.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al · 2019
Cited alongside, same era.
Editing a classifier by rewriting its prediction rules
Shibani Santurkar, Dimitris Tsipras, Mahalaxmi Elango, David Bau, Antonio Torralba, and Aleksander Madry · 2021
Later among the works it cites.
GPT-J-6B: A 6 Billion Parameter Autoregressive Language Model
Ben Wang and Aran Komatsuzaki · 2021
Later among the works it cites.
Of non-linearity and commutativity in bert
Sumu Zhao, Damián Pascual, Gino Brunner, and Roger Wattenhofer · 2021
Later among the works it cites.
Graphical clusterability and local specialization in deep neural networks
Stephen Casper, Shlomi Hod, Daniel Filan, Cody Wild, Andrew Critch, and Stuart Russell · 2022
Later among the works it cites.
Disentangled explanations of neural network predictions by finding relevant subspaces
Pattarawat Chormai, Jan Herrmann, Klaus-Robert Müller, and Grégoire Montavon · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Compositional explanations of neurons
Jesse Mu and Jacob Andreas · 2020
Cited alongside, same era.
A primer in BERTology: What we know about how BERT works
Anna Rogers, Olga Kovaleva, and Anna Rumshisky · 2020
Cited alongside, same era.
An interpretability illusion for bert
Tolga Bolukbasi, Adam Pearce, Ann Yuan, Andy Coenen, Emily Reif, Fernanda Viégas, and Martin Wattenberg · 2021
Cited alongside, same era.
Editing factual knowledge in language models
Nicola De Cao, Wilker Aziz, and Ivan Titov · 2021
Cited alongside, same era.
A mathematical framework for transformer circuits
Nelson Elhage, Neel Nanda, Catherine Olsson, Tom Henighan, Nicholas Joseph, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Nova DasSarma, Dawn Drain, Deep Ganguli, Zac Hatfield-Dodds, Danny Hernandez, Andy Jones, Jackson Kernion, Liane Lovitt, Kamal Ndousse, Dario Amodei, Tom Brown, Jack Clark, Jared Kaplan, Sam McCandlish, and Chris Olah · 2021
Cited alongside, same era.
Do language models have beliefs? methods for detecting, updating, and visualizing model beliefs
Peter Hase, Mona Diab, Asli Celikyilmaz, Xian Li, Zornitsa Kozareva, Veselin Stoyanov, Mohit Bansal, and Srinivasan Iyer · 2021
Cited alongside, same era.
Eric Mitchell, Charles Lin, Antoine Bosselut, Chelsea Finn, and Christopher D Manning · 2021
Cited alongside, same era.
Later among the works it cites.
Local relighting of real scenes
Audrey Cui, Ali Jahanian, Agata Lapedriza, Antonio Torralba, Shahin Mahdizadehaghdam, Rohit Kumar, and David Bau · 2022
Later among the works it cites.
Knowledge neurons in pretrained transformers
Damai Dai, Li Dong, Yaru Hao, Zhifang Sui, and Furu Wei · 2022
Later among the works it cites.
Transformer feed-forward layers build predictions by promoting concepts in the vocabulary space
Mor Geva, Avi Caciularu, Kevin Ro Wang, and Yoav Goldberg · 2022
Later among the works it cites.
Natural language descriptions of deep visual features
Evan Hernandez, Sarah Schwettmann, David Bau, Teona Bagashvili, Antonio Torralba, and Jacob Andreas · 2022
Later among the works it cites.
Locating and editing factual knowledge in gpt
Kevin Meng, David Bau, Alex Andonian, and Yonatan Belinkov · 2022
Later among the works it cites.
Memory-based model editing at scale
Eric Mitchell, Charles Lin, Antoine Bosselut, Christopher D Manning, and Chelsea Finn · 2022
Later among the works it cites.
Finding skill neurons in pre-trained transformer-based language models
Xiaozhi Wang, Kaiyue Wen, Zhengyan Zhang, Lei Hou, Zhiyuan Liu, and Juanzi Li · 2022
Later among the works it cites.