A mathematical framework for transformer circuits
Nelson Elhage, Neel Nanda, Catherine Olsson, Tom Henighan, Nicholas Joseph, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Nova DasSarma, Dawn Drain, Deep Ganguli, Zac Hatfield-Dodds, Danny Hernandez, Andy Jones, Jackson Kernion, Liane Lovitt, Kamal Ndousse, Dario Amodei, Tom Brown, Jack Clark, Jared Kaplan, Sam McCandlish, and Chris Olah · 2021
Cited alongside, same era.
Causal abstractions of neural networks
Atticus Geiger, Hanson Lu, Thomas Icard, and Christopher Potts · 2021
Cited alongside, same era.
Unsolved problems in ml safety
Original
Dan Hendrycks, Nicholas Carlini, John Schulman, and Jacob Steinhardt · 2021
Cited alongside, same era.
Natural language descriptions of deep visual features
Evan Hernandez, Sarah Schwettmann, David Bau, Teona Bagashvili, Antonio Torralba, and Jacob Andreas · 2021
Cited alongside, same era.
Implicit representations of meaning in neural language models
Original
Belinda Z Li, Maxwell Nye, and Jacob Andreas · 2021
Cited alongside, same era.
Mapping language models to grounded conceptual spaces
Roma Patel and Ellie Pavlick · 2021
Cited alongside, same era.
A primer in bertology: What we know about how bert works
Anna Rogers, Olga Kovaleva, and Anna Rumshisky · 2021
Cited alongside, same era.
Roformer: Enhanced transformer with rotary position embedding
Original
Jianlin Su, Yu Lu, Shengfeng Pan, Ahmed Murtadha, Bo Wen, and Yunfeng Liu · 2021
Cited alongside, same era.
Calibrate before use: Improving few-shot performance of language models
Zihao Zhao, Eric Wallace, Shi Feng, Dan Klein, and Sameer Singh · 2021
Cited alongside, same era.
Emergent world representations: Exploring a sequence model trained on a synthetic task
Original
Kenneth Li, Aspen K Hopkins, David Bau, Fernanda Viégas, Hanspeter Pfister, and Martin Wattenberg · 2022
Cited alongside, same era.
Locating and editing factual associations in gpt
Kevin Meng, David Bau, Alex Andonian, and Yonatan Belinkov · 2022
Cited alongside, same era.
Measuring and manipulating knowledge representations in language models
Original
Evan Hernandez, Belinda Z Li, and Jacob Andreas
Cited in the paper.