Fetching the paper…
Reading the bibliography…
Linear layers in neural networks (NNs) trained by gradient descent can be expressed as a key-value memory system which stores all training datapoints and the initial weights, and produces outputs using unnormalised dot attention over the entire training experience.
The perceptron: a probabilistic model for information storage and organization in the brain
Rosenblatt, F · 1958
Earlier work this paper cites.
Theoretical foundations of potential function method in pattern recognition
Aizerman, M. A., Braverman, E. M., and Rozonoer, L. I · 1964
Earlier work this paper cites.
Learning to control fast-weight memories: An alternative to recurrent nets
Schmidhuber, J · 1991
Earlier work this paper cites.
A training algorithm for optimal margin classifiers
Boser, B. E., Guyon, I., and Vapnik, V · 1992
Earlier work this paper cites.
Reducing the ratio between learning complexity and number of time varying variables in fully recurrent nets
Schmidhuber, J · 1993
Earlier work this paper cites.
Long short-term memory
Hochreiter, S. and Schmidhuber, J · 1997
Earlier work this paper cites.
A tutorial on support vector machines for pattern recognition
Burges, C. J · 1998
Earlier work this paper cites.
The MNIST database of handwritten digits
LeCun, Y., Cortes, C., and Burges, C. J · 1998
Earlier work this paper cites.
Catastrophic forgetting in connectionist networks
French, R. M · 1999
Earlier work this paper cites.
Learning with kernels: support vector machines, regularization, optimization, and beyond
Schölkopf, B. and Smola, A. J · 2002
Earlier work this paper cites.
Pattern Recognition and Machine Learning
Bishop, C. M · 2006
Earlier work this paper cites.
Supervised sequence labelling with recurrent neural networks
Graves, A · 2008
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A · 2009
Cited alongside, same era.
Distributed representations of words and phrases and their compositionality
Mikolov, T., Sutskever, I., Chen, K., Corrado, G. S., and Dean, J · 2013
Cited alongside, same era.
Graves, A., Wayne, G., and Danihelka, I · 2014
Cited alongside, same era.
Visualizing and understanding convolutional networks
Zeiler, M. D. and Fergus, R · 2014
Cited alongside, same era.
Neural machine translation by jointly learning to align and translate
Bahdanau, D., Cho, K., and Bengio, Y · 2015
Cited alongside, same era.
Effective approaches to attention-based neural machine translation
Luong, M.-T., Pham, H., and Manning, C. D · 2015
Language modeling with deep Transformers
Irie, K., Zeyer, A., Schlüter, R., and Ney, H · 2019
Later among the works it cites.
Augmenting self-attention with persistent memory
Sukhbaatar, S., Grave, E., Lample, G., Jegou, H., and Joulin, A · 2019
Later among the works it cites.
Transformer dissection: An unified understanding for transformer’s attention via the lens of kernel
Tsai, Y.-H. H., Bai, S., Yamada, M., Morency, L.-P., and Salakhutdinov, R · 2019
Later among the works it cites.
Language models are few-shot learners
Brown, T. B. et al · 2020
Later among the works it cites.
Every model learned by gradient descent is approximately a kernel machine
Domingos, P · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
End-to-end memory networks
Sukhbaatar, S., Szlam, A., Weston, J., and Fergus, R · 2015
Cited alongside, same era.
Using fast weights to attend to the recent past
Ba, J., Hinton, G. E., Mnih, V., Leibo, J. Z., and Ionescu, C · 2016
Cited alongside, same era.
Key-value memory networks for directly reading documents
Miller, A. H., Fisch, A., Dodge, J., Karimi, A., Bordes, A., and Weston, J · 2016
Cited alongside, same era.
Pointer sentinel mixture models
Merity, S., Xiong, C., Bradbury, J., and Socher, R · 2017
Cited alongside, same era.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Cited alongside, same era.
Fashion-MNIST: a novel image dataset for benchmarking machine learning algorithms
Xiao, H., Rasul, K., and Vollgraf, R · 2017
Cited alongside, same era.
Transformers are RNNs: Fast autoregressive transformers with linear attention
Katharopoulos, A., Vyas, A., Pappas, N., and Fleuret, F · 2020
Later among the works it cites.
Learning in high dimension always amounts to extrapolation
Balestriero, R., Pesenti, J., and LeCun, Y · 2021
Later among the works it cites.
Rethinking attention with performers
Choromanski, K., Likhosherstov, V., Dohan, D., Song, X., Gane, A., Sarlos, T., Hawkins, P., Davis, J., Mohiuddin, A., Kaiser, L., et al · 2021
Later among the works it cites.
Transformer feed-forward layers are key-value memories
Geva, M., Schuster, R., Berant, J., and Levy, O · 2021
Later among the works it cites.
Random feature attention
Peng, H., Pappas, N., Yogatama, D., Schwartz, R., Smith, N. A., and Kong, L · 2021
Later among the works it cites.
Zero-shot text-to-image generation
Ramesh, A., Pavlov, M., Goh, G., Gray, S., Voss, C., Radford, A., Chen, M., and Sutskever, I · 2021
Later among the works it cites.
Linear Transformers are secretly fast weight programmers
Schlag, I., Irie, K., and Schmidhuber, J · 2021
Later among the works it cites.