A mathematical framework for transformer circuits
Nelson Elhage, Neel Nanda, Catherine Olsson, Tom Henighan, Nicholas Joseph, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Nova DasSarma, Dawn Drain, Deep Ganguli, Zac Hatfield-Dodds, Danny Hernandez, Andy Jones, Jackson Kernion, Liane Lovitt, Kamal Ndousse, Dario Amodei, Tom Brown, Jack Clark, Jared Kaplan, Sam McCandlish, and Chris Olah. 2021 · 2021
Cited alongside, same era.
Attention flows are shapley value explanations
Original
Kawin Ethayarajh and Dan Jurafsky. 2021 · 2021
Cited alongside, same era.
Visqa: X-raying vision and language reasoning in transformers
Theo Jaunet, Corentin Kervadec, Romain Vuillemot, Grigory Antipov, Moez Baccouche, and Christian Wolf. 2021 · 2021
Cited alongside, same era.
How can we know when language models know? on the calibration of language models for question answering
Zhengbao Jiang, Jun Araki, Haibo Ding, and Graham Neubig. 2021 · 2021
Cited alongside, same era.
Analyzing transformers in embedding space
Original
Guy Dar, Mor Geva, Ankit Gupta, and Jonathan Berant. 2022 · 2022
Cited alongside, same era.
Toy models of superposition
Original
Nelson Elhage, Tristan Hume, Catherine Olsson, Nicholas Schiefer, Tom Henighan, Shauna Kravec, Zac Hatfield-Dodds, Robert Lasenby, Dawn Drain, Carol Chen, et al. 2022 · 2022
Cited alongside, same era.
Transformer feed-forward layers build predictions by promoting concepts in the vocabulary space
Original
Mor Geva, Avi Caciularu, Kevin Ro Wang, and Yoav Goldberg. 2022 · 2022
Cited alongside, same era.
A rigorous study of integrated gradients method and extensions to internal neuron attributions
Daniel D Lundstrom, Tianjian Huang, and Meisam Razaviyayn. 2022 · 2022
Cited alongside, same era.
Locating and editing factual associations in gpt
Kevin Meng, David Bau, Alex Andonian, and Yonatan Belinkov. 2022 · 2022
Cited alongside, same era.
Mechanistic interpretability, variables, and the importance of interpretable bases
Chris Olah. 2022 · 2022
Cited alongside, same era.
In-context learning and induction heads
Original
Catherine Olsson, Nelson Elhage, Neel Nanda, Nicholas Joseph, Nova DasSarma, Tom Henighan, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, et al. 2022 · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. 2022 · 2022
Cited alongside, same era.