Fetching the paper…
Reading the bibliography…
The weight matrix (WM) of a neural network (NN) is its program.
Adaptive switching circuits
Widrow, B. and Hoff, M. E · 1960
Earlier work this paper cites.
A formal theory of inductive inference. Part I
Solomonoff, R. J · 1964
Earlier work this paper cites.
Speculations concerning the first ultraintelligent machine
Good, I · 1965
Earlier work this paper cites.
Evolutionary principles in self-referential learning, or on learning how to learn: the meta-meta-… hook. Institut für Informatik, Technische Universität München, 1987
Schmidhuber, J · 1987
Earlier work this paper cites.
Connectionism and cognitive architecture: A critical analysis
Fodor, J. A., Pylyshyn, Z. W., et al · 1988
Earlier work this paper cites.
Phoneme recognition using time-delay neural networks
Waibel, A., Hanazawa, T., Hinton, G., Shikano, K., and Lang, K. J · 1989
Earlier work this paper cites.
Fixed-weight networks can learn
Cotter, N. E. and Conwell, P. R · 1990
Earlier work this paper cites.
A stochastic version of the delta rule
Hanson, S. J · 1990
Earlier work this paper cites.
Making the world differentiable: On using fully recurrent self-supervised neural networks for dynamic reinforcement learning and planning in non-stationary environments
Schmidhuber, J · 1990
Earlier work this paper cites.
An efficient gradient-based algorithm for on-line training of recurrent network trajectories
Williams, R. J. and Peng, J · 1990
Earlier work this paper cites.
Learning algorithms and fixed dynamics
Cotter, N. E. and Conwell, P. R · 1991
Earlier work this paper cites.
Learning to control fast-weight memories: An alternative to recurrent nets
Schmidhuber, J · 1991
Earlier work this paper cites.
Long short-term memory
Hochreiter, S. and Schmidhuber, J · 1997
Earlier work this paper cites.
Learning to forget: Continual prediction with LSTM
Gers, F. A., Schmidhuber, J., and Cummins, F · 2000
Earlier work this paper cites.
Learning to learn using gradient descent
Hochreiter, S., Younger, A. S., and Conwell, P. R · 2001
Earlier work this paper cites.
Gödel machines: Fully self-referential optimal universal self-improvers
Schmidhuber, J · 2006
Earlier work this paper cites.
The logic of intelligence
Wang, P · 2007
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A · 2009
Earlier work this paper cites.
Self-programming: Operationalizing autonomy
Nivel, E. and Thórisson, K. R · 2009
Earlier work this paper cites.
Recurrent neural network based language model
Mikolov, T., Karafiát, M., Burget, L., Cernocký, J., and Khudanpur, S · 2010
Earlier work this paper cites.
Singularity Hypotheses: A Scientific and Philosophical Assessment
Eden, A. H., Moor, J. H., Soraker, J. H., and Steinhart, E · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2014
Earlier work this paper cites.
Dropout: a simple way to prevent neural networks from overfitting
Srivastava, N., Hinton, G., Krizhevsky, A., Sutskever, I., and Salakhutdinov, R · 2014
Earlier work this paper cites.
Human-level concept learning through probabilistic program induction
Lake, B. M., Salakhutdinov, R., and Tenenbaum, J. B · 2015
Cited alongside, same era.
Efficient object localization using convolutional networks
Tompson, J., Goroshin, R., Jain, A., LeCun, Y., and Bregler, C · 2015
Cited alongside, same era.
Using fast weights to attend to the recent past
Ba, J., Hinton, G. E., Mnih, V., Leibo, J. Z., and Ionescu, C · 2016
Cited alongside, same era.
Increasing the action gap: New operators for reinforcement learning
Bellemare, M. G., Ostrovski, G., Guez, A., Thomas, P. S., and Munos, R · 2016
Cited alongside, same era.
RL 2 : Fast reinforcement learning via slow reinforcement learning
Duan, Y., Schulman, J., Chen, X., Bartlett, P. L., Sutskever, I., and Abbeel, P · 2016
Cited alongside, same era.
Meta-learning with memory-augmented neural networks
Risks from learned optimization in advanced machine learning systems
Hubinger, E., van Merwijk, C., Mikulik, V., Skalse, J., and Garrabrant, S · 2019
Later among the works it cites.
Torchbeast: A PyTorch platform for distributed RL
Küttler, H., Nardelli, N., Lavril, T., Selvatici, M., Sivakumar, V., Rocktäschel, T., and Grefenstette, E · 2019
Later among the works it cites.
Backpropamine: training self-modifying neural networks with differentiable neuromodulated plasticity
Miconi, T., Rawal, A., Clune, J., and Stanley, K. O · 2019
Later among the works it cites.
Metalearned neural memory
Munkhdalai, T., Sordoni, A., Wang, T., and Trischler, A · 2019
Later among the works it cites.
Leveraging procedural generation to benchmark reinforcement learning
Cobbe, K., Hesse, C., Hilton, J., and Schulman, J · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Santoro, A., Bartunov, S., Botvinick, M., Wierstra, D., and Lillicrap, T. P · 2016
Cited alongside, same era.
Growing recursive self-improvers
Steunebrink, B. R., Thórisson, K. R., and Schmidhuber, J · 2016
Cited alongside, same era.
Matching networks for one shot learning
Vinyals, O., Blundell, C., Lillicrap, T., Kavukcuoglu, K., and Wierstra, D · 2016
Cited alongside, same era.
Model-agnostic meta-learning for fast adaptation of deep networks
Finn, C., Abbeel, P., and Levine, S · 2017
Cited alongside, same era.
Meta networks
Munkhdalai, T. and Yu, H · 2017
Cited alongside, same era.
Optimization as a model for few-shot learning
Ravi, S. and Larochelle, H · 2017
Cited alongside, same era.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Cited alongside, same era.
Livewired: The inside story of the ever-changing brain
Eagleman, D · 2020
Later among the works it cites.
Transformers are RNNs: Fast autoregressive transformers with linear attention
Katharopoulos, A., Vyas, A., Pappas, N., and Fleuret, F · 2020
Later among the works it cites.
Meta-learning through hebbian plasticity in random networks
Najarro, E. and Risi, S · 2020
Later among the works it cites.
Adaptive reinforcement learning through evolving self-modifying neural networks
Schmidgall, S · 2020
Later among the works it cites.
Meta-baseline: exploring simple meta-learning for few-shot learning
Chen, Y., Liu, Z., Xu, H., Darrell, T., and Wang, X · 2021
Later among the works it cites.
Rethinking attention with performers
Choromanski, K., Likhosherstov, V., Dohan, D., Song, X., Gane, A., Sarlos, T., Hawkins, P., Davis, J., Mohiuddin, A., Kaiser, L., et al · 2021
Later among the works it cites.
Going beyond linear transformers with recurrent fast weight programmers
Irie, K., Schlag, I., Csordás, R., and Schmidhuber, J · 2021
Later among the works it cites.
Meta-learning backpropagation and improving it
Kirsch, L. and Schmidhuber, J · 2021
Later among the works it cites.
Pitfalls of static language modelling
Lazaridou, A., Kuncoro, A., Gribovskaya, E., Agrawal, D., Liska, A., Terzi, T., Gimenez, M., d’Autume, C. d. M., Ruder, S., Yogatama, D., et al · 2021
Later among the works it cites.
The CLEAR Benchmark: Continual LEArning on Real-World Imagery
Lin, Z., Shi, J., Pathak, D., and Ramanan, D · 2021
Later among the works it cites.
Random feature attention
Peng, H., Pappas, N., Yogatama, D., Schwartz, R., Smith, N. A., and Kong, L · 2021
Later among the works it cites.
Meta-learning bidirectional update rules
Sandler, M., Vladymyrov, M., Zhmoginov, A., Miller, N., Madams, T., Jackson, A., and y Arcas, B. A · 2021
Later among the works it cites.
Linear Transformers are secretly fast weight programmers
Schlag, I., Irie, K., and Schmidhuber, J · 2021
Later among the works it cites.
MLP-mixer: An all-MLP architecture for vision
Tolstikhin, I. O., Houlsby, N., Kolesnikov, A., Beyer, L., Zhai, X., Unterthiner, T., Yung, J., Steiner, A., Keysers, D., Uszkoreit, J., et al · 2021
Later among the works it cites.
Addressing catastrophic forgetting in few-shot problems
Yap, P. C., Ritter, H., and Barber, D · 2021
Later among the works it cites.
The neural data router: Adaptive control flow in transformers improves systematic generalization
Csordás, R., Irie, K., and Schmidhuber, J · 2022
Closest in time.
Hypertransformer: Model generation for supervised and semi-supervised few-shot learning
Zhmoginov, A., Sandler, M., and Vladymyrov, M · 2022
Closest in time.