Fetching the paper…
Reading the bibliography…
We show the formal equivalence of linearised self-attention mechanisms and fast weight controllers from the early '90s, where a ``slow" neural net learns by gradient descent to program the ``fast weights" of another net through sequences of elementary programming instructions which are additive outer products of self-invented activation patterns (today called keys and values).
The organization of behavior: a neuropsycholocigal theory
Hebb, D. O · 1949
Earlier work this paper cites.
Adaptive switching circuits
Widrow, B. and Hoff, M. E · 1960
Earlier work this paper cites.
Die lernmatrix
Steinbuch, K · 1961
Earlier work this paper cites.
Learning matrices and their applications
Steinbuch, K. and Piske, U. A. W · 1963
Earlier work this paper cites.
Correlation matrix memories
Kohonen, T · 1972
Earlier work this paper cites.
The existence of persistent states in the brain
Little, W. A · 1974
Earlier work this paper cites.
On associative memory
Palm, G · 1980
Earlier work this paper cites.
The correlation theory of brain function
von der Malsburg, C · 1981
Earlier work this paper cites.
Dynamic connections in neural networks
Feldman, J. A · 1982
Earlier work this paper cites.
Neural networks and physical systems with emergent collective computational abilities
Hopfield, J. J · 1982
Earlier work this paper cites.
Using fast weights to deblur old memories
Hinton, G. E. and Plaut, D. C · 1987
Earlier work this paper cites.
Bidirectional associative memories
Kosko, B · 1988
Earlier work this paper cites.
A stochastic version of the delta rule
Hanson, S. J · 1990
Earlier work this paper cites.
Tensor product variable binding and the representation of symbolic structures in connectionist systems
Smolensky, P · 1990
Earlier work this paper cites.
Learning to control fast-weight memories: An alternative to recurrent nets
Schmidhuber, J · 1991
Earlier work this paper cites.
26 March 1991: Neural nets learn to program neural nets with fast weights—like today’s Transformer variants. 2021: New stuff!, AI Blog, 2021
Schmidhuber, J · 1991
Earlier work this paper cites.
Learning to control fast-weight memories: An alternative to dynamic recurrent networks
Schmidhuber, J · 1992
Earlier work this paper cites.
Reducing the ratio between learning complexity and number of time varying variables in fully recurrent nets
Schmidhuber, J · 1993
Earlier work this paper cites.
Long short-term memory
Hochreiter, S. and Schmidhuber, J · 1997
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Papineni, K., Roukos, S., Ward, T., and Zhu, W.-J · 2002
Earlier work this paper cites.
Generating sequences with recurrent neural networks
Graves, A · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2014
Earlier work this paper cites.
Dropout: a simple way to prevent neural networks from overfitting
Srivastava, N., Hinton, G., Krizhevsky, A., Sutskever, I., and Salakhutdinov, R · 2014
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Bahdanau, D., Cho, K., and Bengio, Y · 2015
Cited alongside, same era.
A dynamic convolutional layer for short rangeweather prediction
Klein, B., Wolf, L., and Afek, Y · 2015
Cited alongside, same era.
Using fast weights to attend to the recent past
Ba, J., Hinton, G. E., Mnih, V., Leibo, J. Z., and Ionescu, C · 2016
Cited alongside, same era.
Long short-term memory-networks for machine reading
Cheng, J., Dong, L., and Lapata, M · 2016
Cited alongside, same era.
Fast and accurate deep network learning by exponential linear units (ELUs)
Clevert, D.-A., Unterthiner, T., and Hochreiter, S · 2016
Cited alongside, same era.
Dynamic filter networks
Jia, X., De Brabandere, B., Tuytelaars, T., and Gool, L. V · 2016
Cited alongside, same era.
Efficient attention: Attention with linear complexities
Shen, Z., Zhang, M., Zhao, H., Yi, S., and Li, H · 2018
Later among the works it cites.
Character-level language modeling with deeper self-attention
Al-Rfou, R., Choe, D., Constant, N., Guo, M., and Jones, L · 2019
Later among the works it cites.
Adaptive input representations for neural language modeling
Baevski, A. and Auli, M · 2019
Later among the works it cites.
Transformer-XL: Attentive language models beyond a fixed-length context
Dai, Z., Yang, Z., Yang, Y., Cohen, W. W., Carbonell, J., Le, Q. V., and Salakhutdinov, R · 2019
Later among the works it cites.
BERT: pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M., Lee, K., and Toutanova, K · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Dense associative memory for pattern recognition
Krotov, D. and Hopfield, J. J · 2016
Cited alongside, same era.
Image question answering using convolutional neural network with dynamic parameter prediction
Noh, H., Seo, P. H., and Han, B · 2016
Cited alongside, same era.
A decomposable attention model for natural language inference
Parikh, A. P., Täckström, O., Das, D., and Uszkoreit, J · 2016
Cited alongside, same era.
Neural machine translation of rare words with subword units
Sennrich, R., Haddow, B., and Birch, A · 2016
Cited alongside, same era.
On a model of associative memory with huge storage capacity
Demircigil, M., Heusel, J., Löwe, M., Upgang, S., and Vermet, F · 2017
Cited alongside, same era.
Hypernetworks
Ha, D., Dai, A., and Le, Q. V · 2017
Cited alongside, same era.
Miconi, T., Rawal, A., Clune, J., and Stanley, K. O · 2019
Later among the works it cites.
Metalearned neural memory
Munkhdalai, T., Sordoni, A., Wang, T., and Trischler, A · 2019
Later among the works it cites.
fairseq: A fast, extensible toolkit for sequence modeling
Ott, M., Edunov, S., Baevski, A., Fan, A., Gross, S., Ng, N., Grangier, D., and Auli, M · 2019
Later among the works it cites.
Pytorch: An imperative style, high-performance deep learning library
Paszke, A. et al · 2019
Later among the works it cites.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., and Sutskever, I · 2019
Later among the works it cites.
Transformer dissection: An unified understanding for transformer’s attention via the lens of kernel
Tsai, Y.-H. H., Bai, S., Yamada, M., Morency, L.-P., and Salakhutdinov, R · 2019
Later among the works it cites.
On the modularity of hypernetworks
Galanti, T. and Wolf, L · 2020
Later among the works it cites.
On the binding problem in artificial neural networks
Greff, K., van Steenkiste, S., and Schmidhuber, J · 2020
Later among the works it cites.
How much self-attention do we need? Trading attention for feed-forward layers
Irie, K., Gerstenberger, A., Schlüter, R., and Ney, H · 2020
Later among the works it cites.
Transformers are rnns: Fast autoregressive transformers with linear attention
Katharopoulos, A., Vyas, A., Pappas, N., and Fleuret, F · 2020
Later among the works it cites.
Meta learning backpropagation and improving it
Kirsch, L. and Schmidhuber, J · 2020
Later among the works it cites.
Compressive transformers for long-range sequence modelling
Rae, J. W., Potapenko, A., Jayakumar, S. M., Hillier, C., and Lillicrap, T. P · 2020
Later among the works it cites.
Lambdanetworks: Modeling long-range interactions without attention
Bello, I · 2021
Closest in time.
Rethinking attention with performers
Choromanski, K., Likhosherstov, V., Dohan, D., Song, X., Gane, A., Sarlos, T., Hawkins, P., Davis, J., Mohiuddin, A., Kaiser, L., et al · 2021
Closest in time.
Random feature attention
Peng, H., Pappas, N., Yogatama, D., Schwartz, R., Smith, N. A., and Kong, L · 2021
Closest in time.
Hopfield networks is all you need
Ramsauer, H., Schäfl, B., Lehner, J., Seidl, P., Widrich, M., Gruber, L., Holzleitner, M., Adler, T., Kreil, D., Kopp, M. K., Klambauer, G., Brandstetter, J., and Hochreiter, S · 2021
Closest in time.
Learning associative inference using fast weight memory
Schlag, I., Munkhdalai, T., and Schmidhuber, J · 2021
Closest in time.
Long range arena: A benchmark for efficient transformers
Tay, Y., Dehghani, M., Abnar, S., Shen, Y., Bahri, D., Pham, P., Rao, J., Yang, L., Ruder, S., and Metzler, D · 2021
Closest in time.