Fetching the paper…
Reading the bibliography…
Large language models have demonstrated an impressive ability to perform factual recall.
Non-holographic associative memory
David J Willshaw, O Peter Buneman, and Hugh Christopher Longuet-Higgins · 1969
Earlier work this paper cites.
Correlation matrix memories
Teuvo Kohonen · 1972
Earlier work this paper cites.
Inequalities in fourier analysis
William Beckner · 1975
Earlier work this paper cites.
Neural networks and physical systems with emergent collective computational abilities
John J Hopfield · 1982
Earlier work this paper cites.
The capacity of the hopfield associative memory
R. McEliece, E. Posner, E. Rodemich, and S. Venkatesh · 1987
Earlier work this paper cites.
Sobolev inequalities, the poisson semigroup, and analysis on the sphere sn
William Beckner · 1992
Earlier work this paper cites.
Some applications of hypercontractive inequalities in quantum information theory
Ashley Montanaro · 2012
Earlier work this paper cites.
Dense associative memory for pattern recognition
Dmitry Krotov and John J Hopfield · 2016
Earlier work this paper cites.
On a model of associative memory with huge storage capacity
Mete Demircigil, Judith Heusel, Matthias Löwe, Sven Upgang, and Franck Vermet · 2017
Earlier work this paper cites.
Neural tangent kernel: Convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel, and Clément Hongler · 2018
Earlier work this paper cites.
Language models as knowledge bases?
Fabio Petroni, Tim Rocktäschel, Patrick Lewis, Anton Bakhtin, Yuxiang Wu, Alexander H Miller, and Sebastian Riedel · 2019
Earlier work this paper cites.
Network size and size of the weights in memorization with two-layers neural networks
Sébastien Bubeck, Ronen Eldan, Yin Tat Lee, and Dan Mikulincer · 2020
Earlier work this paper cites.
How can we know what language models know?
Zhengbao Jiang, Frank F Xu, Jun Araki, and Graham Neubig · 2020
Earlier work this paper cites.
Self-attentive associative memory
Hung Le, Truyen Tran, and Svetha Venkatesh · 2020
Earlier work this paper cites.
Overparameterized neural networks implement associative memory
Adityanarayanan Radhakrishnan, Mikhail Belkin, and Caroline Uhler · 2020
Earlier work this paper cites.
Hopfield networks is all you need
Hubert Ramsauer, Bernhard Schäfl, Johannes Lehner, Philipp Seidl, Michael Widrich, Thomas Adler, Lukas Gruber, Markus Holzleitner, Milena Pavlović, Geir Kjetil Sandve, et al · 2020
Earlier work this paper cites.
How much knowledge can you pack into the parameters of a language model?
Adam Roberts, Colin Raffel, and Noam Shazeer · 2020
Earlier work this paper cites.
Transformer feed-forward layers are key-value memories
Mor Geva, Roei Schuster, Jonathan Berant, and Omer Levy · 2021
Cited alongside, same era.
Linear transformers are secretly fast weight programmers
Imanol Schlag, Kazuki Irie, and Jürgen Schmidhuber · 2021
Cited alongside, same era.
Approximating how single head attention learns
Charlie Snell, Ruiqi Zhong, Dan Klein, and Jacob Steinhardt · 2021
Cited alongside, same era.
On the optimal memorization power of relu neural networks
Gal Vardi, Gilad Yehudai, and Ohad Shamir · 2021
Cited alongside, same era.
Gpt3. int8 (): 8-bit matrix multiplication for transformers at scale
Tim Dettmers, Mike Lewis, Younes Belkada, and Luke Zettlemoyer · 2022
Cited alongside, same era.
Scan and snap: Understanding training dynamics and token composition in 1-layer transformer
Yuandong Tian, Yiping Wang, Beidi Chen, and Simon Du · 2023
Later among the works it cites.
Transformers learn in-context by gradient descent
Johannes Von Oswald, Eyvind Niklasson, Ettore Randazzo, João Sacramento, Alexander Mordvintsev, Andrey Zhmoginov, and Max Vladymyrov · 2023
Later among the works it cites.
Transformers learn to implement preconditioned gradient descent for in-context learning
Kwangjun Ahn, Xiang Cheng, Hadi Daneshmand, and Suvrit Sra · 2024
Closest in time.
Physics of language models: Part 3.3, knowledge capacity scaling laws
Zeyuan Allen-Zhu and Yuanzhi Li · 2024
Closest in time.
Scaling laws for associative memories
Vivien Cabannes, Elvis Dohmatob, and Alberto Bietti · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Nelson Elhage, Tristan Hume, Catherine Olsson, Nicholas Schiefer, Tom Henighan, Shauna Kravec, Zac Hatfield-Dodds, Robert Lasenby, Dawn Drain, Carol Chen, Roger Grosse, Sam McCandlish, Jared Kaplan, Dario Amodei, Martin Wattenberg, and Christopher Olah · 2022
Cited alongside, same era.
Training compute-optimal large language models
Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, et al · 2022
Cited alongside, same era.
Vision transformers provably learn spatial structure
Samy Jelassi, Michael Sander, and Yuanzhi Li · 2022
Cited alongside, same era.
Locating and editing factual associations in gpt
Kevin Meng, David Bau, Alex Andonian, and Yonatan Belinkov · 2022
Cited alongside, same era.
Birth of a transformer: A memory viewpoint
Alberto Bietti, Vivien Cabannes, Diane Bouchacourt, Herve Jegou, and Leon Bottou · 2023
Cited alongside, same era.
Dissecting recall of factual associations in auto-regressive language models
Mor Geva, Jasmijn Bastings, Katja Filippova, and Amir Globerson · 2023
Cited alongside, same era.
Energy transformer
Benjamin Hoover, Yuchen Liang, Bao Pham, Rameswar Panda, Hendrik Strobelt, Duen Horng Chau, Mohammed Zaki, and Dmitry Krotov · 2023
Cited alongside, same era.
Gaurav Ghosal, Tatsunori Hashimoto, and Aditi Raghunathan · 2024
Closest in time.
Mixture of parrots: Experts improve memorization more than reasoning
Samy Jelassi, Clara Mohri, David Brandfonbrener, Alex Gu, Nikhil Vyas, Nikhil Anand, David Alvarez-Melis, Yuanzhi Li, Sham M Kakade, and Eran Malach · 2024
Closest in time.
Yibo Jiang, Goutham Rajendran, Pradeep Ravikumar, and Bryon Aragam · 2024
Closest in time.
Optimal memorization capacity of transformers
Tokio Kajitsuka and Issei Sato · 2024
Closest in time.
Exponential capacity of dense associative memories
Carlo Lucibello and Marc Mézard · 2024
Closest in time.
Interpreting key mechanisms of factual recall in transformer-based language models
Ang Lv, Kaiyi Zhang, Yuhan Chen, Yulong Wang, Lifeng Liu, Ji-Rong Wen, Jian Xie, and Rui Yan · 2024
Closest in time.
Memory capacity of two layer neural networks with smooth activations
Liam Madden and Christos Thrampoulidis · 2024
Closest in time.
Upper and lower memory capacity bounds of transformers for next-token prediction
Liam Madden, Curtis Fox, and Christos Thrampoulidis · 2024
Closest in time.
One step of gradient descent is provably the optimal in-context learner with one layer of linear self-attention
Arvind Mahankali, Tatsunori B Hashimoto, and Tengyu Ma · 2024
Closest in time.
How transformers learn causal structure with gradient descent
Eshaan Nichani, Alex Damian, and Jason D Lee · 2024
Closest in time.
Trained transformers learn linear models in-context
Ruiqi Zhang, Spencer Frei, and Peter L Bartlett · 2024
Closest in time.