Fetching the paper…
Reading the bibliography…
Recent research has explored the memorization capacity of multi-head attention, but these findings are constrained by unrealistic limitations on the context size.
How Much Knowledge Can You Pack Into the Parameters of a Language Model?, October 2020
Adam Roberts, Colin Raffel, and Noam Shazeer · 2002
Earlier work this paper cites.
Low-rank bottleneck in multi-head attention models
Srinadh Bhojanapalli, Chulhee Yun, Ankit Singh Rawat, Sashank Reddi, and Sanjiv Kumar · 2020
Earlier work this paper cites.
Network size and size of the weights in memorization with two-layers neural networks
Sébastien Bubeck, Ronen Eldan, Yin Tat Lee, and Dan Mikulincer · 2020
Earlier work this paper cites.
Dense passage retrieval for open-domain question answering
Vladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis, Ledell Wu, Sergey Edunov, Danqi Chen, and Wen-tau Yih · 2020
Earlier work this paper cites.
A mathematical framework for transformer circuits
Nelson Elhage, Neel Nanda, Catherine Olsson, Tom Henighan, Nicholas Joseph, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Nova DasSarma, Dawn Drain, Deep Ganguli, Zac Hatfield-Dodds, Danny Hernandez, Andy Jones, Jackson Kernion, Liane Lovitt, Kamal Ndousse, Dario Amodei, Tom Brown, Jack Clark, Jared Kaplan, Sam McCandlish, and Chris Olah · 2021
Earlier work this paper cites.
Locating and editing factual associations in gpt
Kevin Meng, David Bau, Alex Andonian, and Yonatan Belinkov · 2022
Earlier work this paper cites.
On the optimal memorization power of relu neural networks
Gal Vardi, Gilad Yehudai, and Ohad Shamir · 2022
Cited alongside, same era.
Birth of a transformer: A memory viewpoint
Alberto Bietti, Vivien Cabannes, Diane Bouchacourt, Herve Jegou, and Leon Bottou · 2023
Cited alongside, same era.
Attention-only transformers and implementing mlps with attention heads, 2023
Robert Huben and Valerie Morris · 2023
Cited alongside, same era.
Provable memorization capacity of transformers
Junghwan Kim, Michelle Kim, and Barzan Mozafari · 2023
Cited alongside, same era.
Mass-editing memory in a transformer
Kevin Meng, Arnab Sen Sharma, Alex J Andonian, Yonatan Belinkov, and David Bau · 2023
Cited alongside, same era.
Fact Finding: Attempting to Reverse-Engineer Factual Recall on the Neuron Level, December 2023
Neel Nanda, Senthooran Rajamanoharan, János Kramár, and Rohin Shah · 2023
Later among the works it cites.
Interpretability in the wild: a circuit for indirect object identification in GPT-2 small
Kevin Ro Wang, Alexandre Variengien, Arthur Conmy, Buck Shlegeris, and Jacob Steinhardt · 2023
Later among the works it cites.
Scaling laws for associative memories
Vivien Cabannes, Elvis Dohmatob, and Alberto Bietti · 2024
Closest in time.
Memorization capacity of multi-head attention in transformers
Sadegh Mahdavi, Renjie Liao, and Christos Thrampoulidis · 2024
Closest in time.
Look before you leap: A universal emergent decomposition of retrieval tasks in language models
Alexandre Variengien and Eric Winsor · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…