Fetching the paper…
Reading the bibliography…
Do large language models (LLMs) solve reasoning tasks by learning robust generalizable algorithms, or do they memorize training data? To investigate this question, we use arithmetic reasoning as a representative task.
Direct and indirect effects
Judea Pearl · 2001
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Interpreting GPT: The logit lens, 2020
nostalgebraist · 2020
Earlier work this paper cites.
Investigating gender bias in language models using causal mediation analysis
Jesse Vig, Sebastian Gehrmann, Yonatan Belinkov, Sharon Qian, Daniel Nevo, Yaron Singer, and Stuart Shieber · 2020
Earlier work this paper cites.
A mathematical framework for transformer circuits
Nelson Elhage, Neel Nanda, Catherine Olsson, Tom Henighan, Nicholas Joseph, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Nova DasSarma, Dawn Drain, Deep Ganguli, Zac Hatfield-Dodds, Danny Hernandez, Andy Jones, Jackson Kernion, Liane Lovitt, Kamal Ndousse, Dario Amodei, Tom Brown, Jack Clark, Jared Kaplan, Sam McCandlish, and Chris Olah · 2021
Earlier work this paper cites.
Causal abstractions of neural networks
Atticus Geiger, Hanson Lu, Thomas Icard, and Christopher Potts · 2021
Earlier work this paper cites.
Transformer feed-forward layers are key-value memories
Mor Geva, Roei Schuster, Jonathan Berant, and Omer Levy · 2021
Earlier work this paper cites.
Pay attention to MLPs
Hanxiao Liu, Zihang Dai, David So, and Quoc V Le · 2021
Earlier work this paper cites.
GPT-J-6B: A 6 Billion Parameter Autoregressive Language Model
Ben Wang and Aran Komatsuzaki · 2021
Earlier work this paper cites.
Understanding deep learning (still) requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2021
Earlier work this paper cites.
Measures of information reflect memorization patterns
Rachit Bansal, Danish Pruthi, and Yonatan Belinkov · 2022
Earlier work this paper cites.
Probing classifiers: Promises, shortcomings, and advances
Yonatan Belinkov · 2022
Earlier work this paper cites.
Emergent world representations: Exploring a sequence model trained on a synthetic task
Kenneth Li, Aspen K Hopkins, David Bau, Fernanda Viégas, Hanspeter Pfister, and Martin Wattenberg · 2022
Earlier work this paper cites.
Attribution patching: Activation patching at industrial scale, 2022
Neel Nanda · 2022
Earlier work this paper cites.
TransformerLens
Neel Nanda and Joseph Bloom · 2022
Cited alongside, same era.
Extracting latent steering vectors from pretrained language models
Nishant Subramani, Nivedita Suresh, and Matthew E Peters · 2022
Cited alongside, same era.
Memorisation versus generalisation in pre-trained language models
Michael Tänzer, Sebastian Ruder, and Marek Rei · 2022
Cited alongside, same era.
Interpretability in the wild: A circuit for indirect object identification in GPT-2 small
Kevin Ro Wang, Alexandre Variengien, Arthur Conmy, Buck Shlegeris, and Jacob Steinhardt · 2022
Cited alongside, same era.
Pythia: A suite for analyzing large language models across training and scaling
Stella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley, Kyle O’Brien, Eric Hallahan, Mohammad Aflah Khan, Shivanshu Purohit, USVSN Sai Prashanth, Edward Raff, et al · 2023
Cited alongside, same era.
Quantifying memorization across neural language models
Attribution patching outperforms automated circuit discovery
Aaquib Syed, Can Rager, and Arthur Conmy · 2023
Later among the works it cites.
Activation addition: Steering language models without optimization
Alexander Matt Turner, Lisa Thiergart, Gavin Leech, David Udell, Juan J Vazquez, Ulisse Mini, and Monte MacDiarmid · 2023
Later among the works it cites.
Explaining grokking through circuit efficiency
Vikrant Varma, Rohin Shah, Zachary Kenton, János Kramár, and Ramana Kumar · 2023
Later among the works it cites.
Generalization vs. memorization: Tracing language models’ capabilities back to pretraining data
Antonis Antoniades, Xinyi Wang, Yanai Elazar, Alfonso Amayuelas, Alon Albalak, Kexun Zhang, and William Yang Wang · 2024
Closest in time.
Generalisation first, memorisation second? Memorisation localisation for natural language classification tasks
Verna Dankers and Ivan Titov · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Nicholas Carlini, Daphne Ippolito, Matthew Jagielski, Katherine Lee, Florian Tramer, and Chiyuan Zhang · 2023
Cited alongside, same era.
Analyzing transformers in embedding space
Guy Dar, Mor Geva, Ankit Gupta, and Jonathan Berant · 2023
Cited alongside, same era.
Superposition, memorization, and double descent
Tom Henighan, Shan Carter, Tristan Hume, Nelson Elhage, Robert Lasenby, Stanislav Fort, Nicholas Schiefer, and Christopher Olah · 2023
Cited alongside, same era.
Towards a mechanistic interpretation of multi-step reasoning capabilities of language models
Yifan Hou, Jiaoda Li, Yu Fei, Alessandro Stolfo, Wangchunshu Zhou, Guangtao Zeng, Antoine Bosselut, and Mrinmaya Sachan · 2023
Cited alongside, same era.
Arithmetic with language models: From memorization to computation
Davide Maltoni and Matteo Ferrara · 2023
Cited alongside, same era.
Copy suppression: Comprehensively understanding an attention head
Callum McDougall, Arthur Conmy, Cody Rushing, Thomas McGrath, and Neel Nanda · 2023
Cited alongside, same era.
Progress measures for grokking via mechanistic interpretability
Neel Nanda, Lawrence Chan, Tom Lieberum, Jess Smith, and Jacob Steinhardt · 2023
Cited alongside, same era.
Survival of the fittest representation: A case study with modular addition
Xiaoman Delores Ding, Zifan Carl Guo, Eric J Michaud, Ziming Liu, and Max Tegmark · 2024
Closest in time.
Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al · 2024
Closest in time.
Successor heads: Recurring, interpretable attention heads in the wild
Rhys Gould, Euan Ong, George Ogden, and Arthur Conmy · 2024
Closest in time.
Othellogpt learned a bag of heuristics, 2024
jylin, JackS, Adam Karvonen, and Can Rager · 2024
Closest in time.
Fine-tuning enhances existing mechanisms: A case study on entity tracking
Nikhil Prakash, Tamar Rott Shaham, Tal Haklay, Yonatan Belinkov, and David Bau · 2024
Closest in time.
Interpreting and improving large language models in arithmetic calculation
Wei Zhang, Chaoqun Wan, Yonggang Zhang, Yiu ming Cheung, Xinmei Tian, Xu Shen, and Jieping Ye · 2024
Closest in time.
The clock and the pizza: Two stories in mechanistic explanation of neural networks
Ziqian Zhong, Ziming Liu, Max Tegmark, and Jacob Andreas · 2024
Closest in time.
Pre-trained large language models use Fourier features to compute addition
Tianyi Zhou, Deqing Fu, Vatsal Sharan, and Robin Jia · 2024
Closest in time.