Fetching the paper…
Reading the bibliography…
This work focuses on the training dynamics of one associative memory module storing outer products of token embeddings.
Non-holographic associative memory
Willshaw, D., Buneman, P., and Longuet-Higgins, C · 1969
Earlier work this paper cites.
Theories of associative recall
Longuet-Higgins, C., Willshaw, D., and Buneman, P · 1970
Earlier work this paper cites.
Neural networks and physical systems with emergent collective computational abilities
Hopfield, J · 1982
Earlier work this paper cites.
Neural computation of decisions in optimization problems
Hopfield, J. and Tank, D · 1985
Earlier work this paper cites.
Support-vector networks
Cortes, C. and Vapnik, V · 1995
Earlier work this paper cites.
Inequalities on the lambert w function and hyperpower function
Hoorfar, A. and Hassani, M · 2008
Earlier work this paper cites.
Distributed representations of words and phrases and their compositionality
Mikolov, T., Sutskever, I., Chen, K., Corrado, G. S., and Dean, J · 2013
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Ioffe, S. and Szegedy, C · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization, 2015
Kingma, D. and Ba, J · 2015
Earlier work this paper cites.
Ba, J. L., Kiros, J. R., and Hinton, G. E · 2016
Earlier work this paper cites.
Stein variational gradient descent: A general purpose bayesian inference algorithm
Liu, Q. and Wang, D · 2016
Earlier work this paper cites.
On the global convergence of gradient descent for over-parameterized models using optimal transport
Chizat, L. and Bach, F · 2018
Earlier work this paper cites.
A mean field view of the landscape of two-layer neural networks
Mei, S., Montanari, A., and Nguyen, P.-M · 2018
Earlier work this paper cites.
Trainability and accuracy of neural networks: An interacting particle system approach
Rotskoff, G. M. and Vanden-Eijnden, E · 2018
Cited alongside, same era.
The implicit bias of gradient descent on separable data
Soudry, D., Hoffer, E., Nacson, M. S., Gunasekar, S., and Srebro, N · 2018
Cited alongside, same era.
Maximum mean discrepancy gradient flow
Arbel, M., Korba, A., Salim, A., and Gretton, A · 2019
Cited alongside, same era.
What is the effect of importance weighting in deep learning?
Byrd, J. and Lipton, Z · 2019
Cited alongside, same era.
The implicit bias of gradient descent on nonseparable data
Ji, Z. and Telgarsky, M · 2019
Cited alongside, same era.
Gradient descent maximizes the margin of homogeneous neural networks
Lyu, K. and Li, J · 2019
Hopfield networks is all you need
Ramsauer, H., Schäfl, B., Lehner, J., Seidl, P., Widrich, M., Adler, T., Gruber, L., Holzleitner, M., Pavlović, M., Sandve, G. K., Greiff, V., Kreil, D., Kopp, M., Klambauer, G., Brandstetter, J., and Hochreiter, S · 2021
Later among the works it cites.
Linear transformers are secretly fast weight programmers
Schlag, I., Irie, K., and Schmidhuber, J · 2021
Later among the works it cites.
On the benefits of large learning rates for kernel methods
Beugnot, G., Rudi, A., and Mairal, J · 2022
Later among the works it cites.
Infinite-width limit of deep linear neural networks
Chizat, L., Colombo, M., Fernández-Real, X., and Figalli, A · 2022
Later among the works it cites.
In-context learning and induction heads
Olsson, C., Elhage, N., Nanda, N., Joseph, N., DasSarma, N., Henighan, T., Mann, B., Askell, A., Bai, Y., Chen, A., Conerly, T., Drain, D., Ganguli, D., Hatfield-Dodds, Z., Hernandez, D., Johnston, S., Jones, A., Kernion, J., Lovitt, L., Ndousse, K., Amodei, D., Brown, T., Clark, J., Kaplan, J., McCandlish, S., and Olah, C · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Gradient descent on neural networks typically occurs at the edge of stability
Cohen, J., Kaur, S., Li, Y., Kolter, J. Z., and Talwalkar, A · 2020
Cited alongside, same era.
Learning rate annealing can provably help generalization, even for convex problems
Nakkiran, P · 2020
Cited alongside, same era.
Dual training of energy-based models with overparametrized shallow neural networks
Domingo-Enrich, C., Bietti, A., Gabrié, M., Bruna, J., and Vanden-Eijnden, E · 2021
Cited alongside, same era.
A mathematical framework for transformer circuits
Elhage, N., Nanda, N., Olsson, C., Henighan, T., Joseph, N., Mann, B., Askell, A., Bai, Y., Chen, A., Conerly, T., DasSarma, N., Drain, D., Ganguli, D., Hatfield-Dodds, Z., Hernandez, D., Jones, A., Kernion, J., Lovitt, L., Ndousse, K., Amodei, D., Brown, T., Clark, J., Kaplan, J., McCandlish, S., and Olah, C · 2021
Cited alongside, same era.
Transformer feed-forward layers are key-value memories
Geva, M., Schuster, R., Berant, J., and Levy, O · 2021
Cited alongside, same era.
Jacot, A., Ged, F., Şimşek, B., Hongler, C., and Gabriel, F · 2021
Cited alongside, same era.
Second-order regression models exhibit progressive sharpening to the edge of stability
Agarwala, A., Pedregosa, F., and Pennington, J · 2023
Later among the works it cites.
The dynamics of sharpness-aware minimization: Bouncing across ravines and drifting towards wide minima
Bartlett, P. L., Long, P. M., and Bousquet, O · 2023
Later among the works it cites.
Birth of a transformer: A memory viewpoint
Bietti, A., Cabannes, V., Bouchacourt, D., Jegou, H., and Bottou, L · 2023
Later among the works it cites.
Beyond the edge of stability via two-step gradient updates
Chen, L. and Bruna, J · 2023
Later among the works it cites.
Outliers with opposing signals have an outsized effect on neural network optimization
Rosenfeld, E. and Risteski, A · 2023
Later among the works it cites.
Implicit bias of gradient descent for logistic regression at the edge of stability
Wu, J., Braverman, V., and Lee, J. D · 2023
Later among the works it cites.
Scaling laws for associative memories
Cabannes, V., Dohmatob, E., and Bietti, A · 2024
Closest in time.
Large stepsize gradient descent for logistic loss: Non-monotonicity of the loss improves optimization efficiency, 2024
Wu, J., Bartlett, P., Telgarsky, M., and Yu, B · 2024
Closest in time.