Discovering variable binding circuitry with desiderata, 2023
Xander Davies, Max Nadeau, Nikhil Prakash, Tamar Rott Shaham, and David Bau · 2023
Later among the works it cites.
How do language models bind entities in context?, 2023
Jiahai Feng and Jacob Steinhardt · 2023
Later among the works it cites.
Dissecting recall of factual associations in auto-regressive language models
Mor Geva, Jasmijn Bastings, Katja Filippova, and Amir Globerson · 2023
Later among the works it cites.
Localizing model behavior with path patching, 2023
Nicholas Goldowsky-Dill, Chris MacLeod, Lucas Sato, and Aryaman Arora · 2023
Later among the works it cites.
Knowledge is a region in weight space for fine-tuned language models
Original
Almog Gueta, Elad Venezian, Colin Raffel, Noam Slonim, Yoav Katz, and Leshem Choshen · 2023
Later among the works it cites.
Mechanistically analyzing the effects of fine-tuning on procedurally defined tasks, 2023
Samyak Jain, Robert Kirk, Ekdeep Singh Lubana, Robert P. Dick, Hidenori Tanaka, Edward Grefenstette, Tim Rocktäschel, and David Scott Krueger · 2023
Later among the works it cites.
Entity tracking in language models
Original
Najoung Kim and Sebastian Schuster · 2023
Later among the works it cites.
Does circuit analysis interpretability scale? evidence from multiple choice capabilities in chinchilla, 2023
Tom Lieberum, Matthew Rahtz, János Kramár, Neel Nanda, Geoffrey Irving, Rohin Shah, and Vladimir Mikulik · 2023
Later among the works it cites.
Goat: Fine-tuned llama outperforms gpt-4 on arithmetic tasks, 2023
Tiedong Liu and Bryan Kian Hsiang Low · 2023
Later among the works it cites.
Progress measures for grokking via mechanistic interpretability
Original
Neel Nanda, Lawrence Chan, Tom Liberum, Jess Smith, and Jacob Steinhardt · 2023
Later among the works it cites.
Task-specific skill localization in fine-tuned language models
Original
Abhishek Panigrahi, Nikunj Saunshi, Haoyu Zhao, and Sanjeev Arora · 2023
Later among the works it cites.
Stanford alpaca: An instruction-following llama model, 2023
Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B Hashimoto · 2023
Later among the works it cites.
Judging llm-as-a-judge with mt-bench and chatbot arena
Original
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, et al · 2023
Later among the works it cites.
LIMA: Less is more for alignment
Chunting Zhou, Pengfei Liu, Puxin Xu, Srini Iyer, Jiao Sun, Yuning Mao, Xuezhe Ma, Avia Efrat, Ping Yu, LILI YU, Susan Zhang, Gargi Ghosh, Mike Lewis, Luke Zettlemoyer, and Omer Levy · 2023
Later among the works it cites.
Towards automated circuit discovery for mechanistic interpretability
Arthur Conmy, Augustine Mavor-Parker, Aengus Lynch, Stefan Heimersheim, and Adrià Garriga-Alonso · 2024
Closest in time.
Tracr: Compiled transformers as a laboratory for interpretability
David Lindner, János Kramár, Sebastian Farquhar, Matthew Rahtz, Tom McGrath, and Vladimir Mikulik · 2024
Closest in time.
Interpretability at scale: Identifying causal mechanisms in alpaca
Zhengxuan Wu, Atticus Geiger, Thomas Icard, Christopher Potts, and Noah Goodman · 2024
Closest in time.