Fetching the paper…
Reading the bibliography…
In this paper, we delve into several mechanisms employed by Transformer-based language models (LLMs) for factual recall tasks.
On information and sufficiency
Solomon Kullback and Richard A Leibler · 1951
Earlier work this paper cites.
Direct and indirect effects
Judea Pearl · 2001
Earlier work this paper cites.
Causality
Judea Pearl · 2009
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Openwebtext corpus
Aaron Gokaslan and Vanya Cohen · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever · 2019
Earlier work this paper cites.
Fine-tuning language models from human preferences
Daniel M. Ziegler, Nisan Stiennon, Jeffrey Wu, Tom B. Brown, Alec Radford, Dario Amodei, Paul Christiano, and Geoffrey Irving · 2019
Earlier work this paper cites.
Language models are few-shot learners, 2020
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei · 2020
Earlier work this paper cites.
Zoom in: An introduction to circuits
Chris Olah, Nick Cammarata, Ludwig Schubert, Gabriel Goh, Michael Petrov, and Shan Carter · 2020
Earlier work this paper cites.
Investigating gender bias in language models using causal mediation analysis
Jesse Vig, Sebastian Gehrmann, Yonatan Belinkov, Sharon Qian, Daniel Nevo, Yaron Singer, and Stuart Shieber · 2020
Earlier work this paper cites.
A mathematical framework for transformer circuits
Nelson Elhage, Neel Nanda, Catherine Olsson, Tom Henighan, Nicholas Joseph, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Nova DasSarma, Dawn Drain, Deep Ganguli, Zac Hatfield-Dodds, Danny Hernandez, Andy Jones, Jackson Kernion, Liane Lovitt, Kamal Ndousse, Dario Amodei, Tom Brown, Jack Clark, Jared Kaplan, Sam McCandlish, and Chris Olah · 2021
Earlier work this paper cites.
Finetuned language models are zero-shot learners
Jason Wei, Maarten Bosma, Vincent Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M. Dai, and Quoc V. Le · 2021
Earlier work this paper cites.
Softmax linear units
Nelson Elhage, Tristan Hume, Catherine Olsson, Neel Nanda, Tom Henighan, Scott Johnston, Sheer ElShowk, Nicholas Joseph, Nova DasSarma, Ben Mann, Danny Hernandez, Amanda Askell, Kamal Ndousse, Andy Jones, Dawn Drain, Anna Chen, Yuntao Bai, Deep Ganguli, Liane Lovitt, Zac Hatfield-Dodds, Jackson Kernion, Tom Conerly, Shauna Kravec, Stanislav Fort, Saurav Kadavath, Josh Jacobson, Eli Tran-Johnson, Jared Kaplan, Jack Clark, Tom Brown, Sam McCandlish, Dario Amodei, and Christopher Olah · 2022
Earlier work this paper cites.
Toy models of superposition
Nelson Elhage, Tristan Hume, Catherine Olsson, Nicholas Schiefer, Tom Henighan, Shauna Kravec, Zac Hatfield-Dodds, Robert Lasenby, Dawn Drain, Carol Chen, Roger Grosse, Sam McCandlish, Jared Kaplan, Dario Amodei, Martin Wattenberg, and Christopher Olah · 2022
Earlier work this paper cites.
Transformer feed-forward layers build predictions by promoting concepts in the vocabulary space, 2022
Mor Geva, Avi Caciularu, Kevin Ro Wang, and Yoav Goldberg · 2022
Earlier work this paper cites.
Locating and editing factual associations in GPT
Kevin Meng, David Bau, Alex J Andonian, and Yonatan Belinkov · 2022
Earlier work this paper cites.
In-context learning and induction heads
Catherine Olsson, Nelson Elhage, Neel Nanda, Nicholas Joseph, Nova DasSarma, Tom Henighan, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Dawn Drain, Deep Ganguli, Zac Hatfield-Dodds, Danny Hernandez, Scott Johnston, Andy Jones, Jackson Kernion, Liane Lovitt, Kamal Ndousse, Dario Amodei, Tom Brown, Jack Clark, Jared Kaplan, Sam McCandlish, and Chris Olah · 2022
Cited alongside, same era.
OpenAI: Introducing ChatGPT, 2022
OpenAI · 2022
Cited alongside, same era.
Opt: Open pre-trained transformer language models, 2022
Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li, Xi Victoria Lin, Todor Mihaylov, Myle Ott, Sam Shleifer, Kurt Shuster, Daniel Simig, Punit Singh Koura, Anjali Sridhar, Tianlu Wang, and Luke Zettlemoyer · 2022
Cited alongside, same era.
Towards automated circuit discovery for mechanistic interpretability
Arthur Conmy, Augustine N. Mavor-Parker, Aengus Lynch, Stefan Heimersheim, and Adrià Garriga-Alonso · 2023
Cited alongside, same era.
An adversarial example for direct logit attribution: Memory management in gelu-4l, 2023
James Dao, Yeu-Tong Lau, Can Rager, and Jett Janiak · 2023
Cited alongside, same era.
Interpretability in the wild: a circuit for indirect object identification in GPT-2 small
Kevin Ro Wang, Alexandre Variengien, Arthur Conmy, Buck Shlegeris, and Jacob Steinhardt · 2023
Later among the works it cites.
Efficient streaming language models with attention sinks, 2023
Guangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han, and Mike Lewis · 2023
Later among the works it cites.
Characterizing mechanisms for factual recall in language models
Qinan Yu, Jack Merullo, and Ellie Pavlick · 2023
Later among the works it cites.
Siren’s song in the ai ocean: A survey on hallucination in large language models, 2023
Yue Zhang, Yafu Li, Leyang Cui, Deng Cai, Lemao Liu, Tingchen Fu, Xinting Huang, Enbo Zhao, Yu Zhang, Yulong Chen, Longyue Wang, Anh Tuan Luu, Wei Bi, Freda Shi, and Shuming Shi · 2023
Later among the works it cites.
A survey of large language models, 2023
Wayne Xin Zhao, Kun Zhou, Junyi Li, Tianyi Tang, Xiaolei Wang, Yupeng Hou, Yingqian Min, Beichen Zhang, Junjie Zhang, Zican Dong, Yifan Du, Chen Yang, Yushuo Chen, Zhipeng Chen, Jinhao Jiang, Ruiyang Ren, Yifan Li, Xinyu Tang, Zikang Liu, Peiyu Liu, Jian-Yun Nie, and Ji-Rong Wen · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Dissecting recall of factual associations in auto-regressive language models
Mor Geva, Jasmijn Bastings, Katja Filippova, and Amir Globerson · 2023
Cited alongside, same era.
Inspecting and editing knowledge representations in language models, 2023
Evan Hernandez, Belinda Z. Li, and Jacob Andreas · 2023
Cited alongside, same era.
Transformer language models handle word frequency in prediction head
Goro Kobayashi, Tatsuki Kuribayashi, Sho Yokoi, and Kentaro Inui · 2023
Cited alongside, same era.
Does circuit analysis interpretability scale? evidence from multiple choice capabilities in chinchilla, 2023
Tom Lieberum, Matthew Rahtz, János Kramár, Neel Nanda, Geoffrey Irving, Rohin Shah, and Vladimir Mikulik · 2023
Cited alongside, same era.
The hydra effect: Emergent self-repair in language model computations, 2023
Thomas McGrath, Matthew Rahtz, Janos Kramar, Vladimir Mikulik, and Shane Legg · 2023
Cited alongside, same era.
A mechanism for solving relational tasks in transformer language models, 2023
Jack Merullo, Carsten Eickhoff, and Ellie Pavlick · 2023
Cited alongside, same era.
Attention is off by one, 2023
Evan Miller · 2023
Cited alongside, same era.
Fortify the shortest stave in attention: Enhancing context awareness of large language models for effective tool use, 2024
Yuhan Chen, Ang Lv, Ting-En Lin, Changyu Chen, Yuchuan Wu, Fei Huang, Yongbin Li, and Rui Yan · 2024
Closest in time.
How to think step-by-step: A mechanistic understanding of chain-of-thought reasoning, 2024
Subhabrata Dutta, Joykirat Singh, Soumen Chakrabarti, and Tanmoy Chakraborty · 2024
Closest in time.
Monotonic representation of numeric properties in language models, 2024
Benjamin Heinzerling and Kentaro Inui · 2024
Closest in time.
Carrying over algorithm in transformers, 2024
Jorrit Kruthoff · 2024
Closest in time.
Is this the subspace you are looking for? an interpretability illusion for subspace activation patching
Aleksandar Makelov, Georg Lange, Atticus Geiger, and Neel Nanda · 2024
Closest in time.
Circuit component reuse across tasks in transformer language models
Jack Merullo, Carsten Eickhoff, and Ellie Pavlick · 2024
Closest in time.
Competition of mechanisms: Tracing how language models handle facts and counterfactuals, 2024
Francesco Ortu, Zhijing Jin, Diego Doimo, Mrinmaya Sachan, Alberto Cazzaniga, and Bernhard Schölkopf · 2024
Closest in time.
The mechanistic basis of data dependence and abrupt learning in an in-context classification task
Gautam Reddy · 2024
Closest in time.
Do llamas work in english? on the latent language of multilingual transformers, 2024
Chris Wendler, Veniamin Veselovsky, Giovanni Monea, and Robert West · 2024
Closest in time.
Locating factual knowledge in large language models: Exploring the residual stream and analyzing subvalues in vocabulary space, 2024
Zeping Yu and Sophia Ananiadou · 2024
Closest in time.