Fetching the paper…
Reading the bibliography…
Learning arguably involves the discovery and memorization of abstract rules.
Die Lernmatrix
Karl Steinbuch · 1961
Earlier work this paper cites.
Probability inequalities for sums of bounded random variables
Wassily Hoeffding · 1963
Earlier work this paper cites.
Non-holographic associative memory
David Willshaw, Peter Buneman, and Christopher Longuet-Higgins · 1969
Earlier work this paper cites.
Theories of associative recall
Christopher Longuet-Higgins, David. Willshaw, and Peter Buneman · 1970
Earlier work this paper cites.
Learning patterns and pattern sequences by self-organizing nets of threshold elements
Shun-Ichi Amari · 1972
Earlier work this paper cites.
Correlation matrix memories
Teuvo Kohonen · 1972
Earlier work this paper cites.
The existence of persistent states in the brain
William Little · 1974
Earlier work this paper cites.
Neural networks and physical systems with emergent collective computational abilities
John Hopfield · 1982
Earlier work this paper cites.
Tensor product variable binding and the representation of symbolic structures in connectionist systems
Paul Smolensky · 1990
Earlier work this paper cites.
Elements of Information Theory
Thomas Cover and Joy Thomas · 1991
Earlier work this paper cites.
Mesures dominantes et théorème de sanov
Ian Dinwoodie · 1992
Earlier work this paper cites.
Zipf’s word frequency law in natural language: A critical review and future directions
Steven Piantadosi · 2014
Earlier work this paper cites.
Dense associative memory for pattern recognition
Dmitry Krotov and John Hopfield · 2016
Earlier work this paper cites.
Dissecting adam: The sign, magnitude and variance of stochastic gradients
Lukas Balles and Philipp Hennig · 2018
Earlier work this paper cites.
Don’t decay the learning rate, increase the batch size
Samuel Smith, Pieter-Jan Kindermans, Chris Ying, and Quoc V. Le · 2018
Cited alongside, same era.
Pytorch: An imperative style, high-performance deep learning library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Köpf, Edward Yang, Zach DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala · 2019
Cited alongside, same era.
Augmenting self-attention with persistent memory
Sainbayar Sukhbaatar, Edouard Grave, Guillaume Lample, Herve Jegou, and Armand Joulin · 2019
Cited alongside, same era.
Does learning require memorization? a short tale about a long tail
Vitaly Feldman · 2020
Cited alongside, same era.
What neural networks memorize and why: Discovering the long tail via influence estimation
Vitaly Feldman and Chiyuan Zhang · 2020
Cited alongside, same era.
Tensor programs v: Tuning large neural networks via zero-shot hyperparameter transfer
Greg Yang, Edward Hu, Igor Babuschkin, Szymon Sidor, Xiaodong Liu, David Farhi, Nick Ryder, Jakub Pachocki, Weizhu Chen, and Jianfeng Gao · 2021
Later among the works it cites.
Toy models of superposition
Nelson Elhage, Tristan Hume, Catherine Olsson, Nicholas Schiefer, Tom Henighan, Shauna Kravec, Zac Hatfield-Dodds, Robert Lasenby, Dawn Drain, Carol Chen, Roger Grosse, Sam McCandlish, Jared Kaplan, Dario Amodei, Martin Wattenberg, and Christopher Olah · 2022
Later among the works it cites.
Training compute-optimal large language models
Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, et al · 2022
Later among the works it cites.
A solvable model of neural scaling laws
Alexander Maloney, Daniel Roberts, and James Sully · 2022
Later among the works it cites.
Locating and editing factual associations in GPT
Kevin Meng, David Bau, Alex Andonian, and Yonatan Belinkov · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Scaling laws for neural language models
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei · 2020
Cited alongside, same era.
Why are adaptive methods good for attention models?
Jingzhao Zhang, Sairaneeth Karimireddy, Andreas Veit, Seungyeon Kim, Sashank Reddi, Sanjiv Kumar, and Suvrit Sra · 2020
Cited alongside, same era.
Explaining neural scaling laws
Yasaman Bahri, Ethan Dyer, Jared Kaplan, Jaehoon Lee, and Utkarsh Sharma · 2021
Cited alongside, same era.
A mathematical framework for transformer circuits
Nelson Elhage, Neel Nanda, Catherine Olsson, Tom Henighan, Nicholas Joseph, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Nova DasSarma, Dawn Drain, Deep Ganguli, Zac Hatfield-Dodds, Danny Hernandez, Andy Jones, Jackson Kernion, Liane Lovitt, Kamal Ndousse, Dario Amodei, Tom Brown, Jack Clark, Jared Kaplan, Sam McCandlish, and Chris Olah · 2021
Cited alongside, same era.
Transformer feed-forward layers are key-value memories
Mor Geva, Roei Schuster, Jonathan Berant, and Omer Levy · 2021
Cited alongside, same era.
A loss curvature perspective on training instabilities of deep learning models
Justin Gilmer, Behrooz Ghorbani, Ankush Garg, Sneha Kudugunta, Behnam Neyshabur, David Cardoze, George Edward Dahl, Zachary Nado, and Orhan Firat · 2021
Cited alongside, same era.
Marcus Hutter · 2021
Cited alongside, same era.
Beyond neural scaling laws: beating power law scaling via data pruning
Ben Sorscher, Robert Geirhos, Shashank Shekhar, Surya Ganguli, and Ari Morcos · 2022
Later among the works it cites.
Memorizing transformers
Yuhuai Wu, Markus Rabe, DeLesley Hutchins, and Christian Szegedy · 2022
Later among the works it cites.
Birth of a transformer: A memory viewpoint
Alberto Bietti, Vivien Cabannes, Diane Bouchacourt, Herve Jegou, and Leon Bottou · 2023
Closest in time.
A simplistic model of neural scaling laws: Multiperiodic santa fe processes
Lukasz Debowski · 2023
Closest in time.
Noise is not the main factor behind the gap between sgd and adam on transformers, but sign descent might be
Frederik Kunstner, Jacques Chen, Jonathan Wilder Lavington, and Mark Schmidt · 2023
Closest in time.
The quantization model of neural scaling
Eric Michaud, Ziming Liu, Uzay Girit, and Max Tegmark · 2023
Closest in time.
Multiplicative processing in the modeling of cognitive activities in large neural networks
Juan Valle-Lisboa, Andrés Pomi, and Eduardo Mizraji · 2023
Closest in time.
Tensor programs ivb: Adaptive optimization in the infinite-width limit
Greg Yang and Etai Littwin · 2023
Closest in time.