Fetching the paper…
Reading the bibliography…
General-purpose learning systems should improve themselves in open-ended fashion in ever-changing environments.
Adaptive switching circuits
Bernard Widrow and Marcian E Hoff · 1960
Earlier work this paper cites.
Theoretical foundations of potential function method in pattern recognition
Mark A. Aizerman, Emmanuil M. Braverman, and Lev I. Rozonoer · 1964
Earlier work this paper cites.
Gödel, Escher, Bach: an Etemal Golden Braid,
Douglas R Hofstadter · 1979
Earlier work this paper cites.
Studies of mind and brain: Neural principles of learning, perception, development, cognition, and motor control
Stephen T Grossberg · 1982
Earlier work this paper cites.
Evolutionary principles in self-referential learning, or on learning how to learn: the meta-meta-… hook
Jürgen Schmidhuber · 1987
Earlier work this paper cites.
Connectionism and cognitive architecture: A critical analysis
Jerry A Fodor and Zenon W Pylyshyn · 1988
Earlier work this paper cites.
Catastrophic interference in connectionist networks: The sequential learning problem
Michael McCloskey and Neal J Cohen · 1989
Earlier work this paper cites.
Fixed-weight networks can learn
Neil E Cotter and Peter R Conwell · 1990
Earlier work this paper cites.
Episodic memory in connectionist networks
Chris A Kortge · 1990
Earlier work this paper cites.
Connectionist models of recognition memory: constraints imposed by learning and forgetting functions
Roger Ratcliff · 1990
Earlier work this paper cites.
Making the world differentiable: On using fully recurrent self-supervised neural networks for dynamic reinforcement learning and planning in non-stationary environments
Jürgen Schmidhuber · 1990
Earlier work this paper cites.
Learning algorithms and fixed dynamics
Neil E Cotter and Peter R Conwell · 1991
Earlier work this paper cites.
Using semi-distributed representations to overcome catastrophic forgetting in connectionist networks
Robert M French · 1991
Earlier work this paper cites.
Learning to control fast-weight memories: An alternative to recurrent nets
Jürgen Schmidhuber · 1991
Earlier work this paper cites.
A self-referential weight matrix
Jürgen Schmidhuber · 1993
Earlier work this paper cites.
Continual Learning in Reinforcement Environments
Mark B. Ring · 1994
Earlier work this paper cites.
On learning how to learn learning strategies
Jürgen Schmidhuber · 1994
Earlier work this paper cites.
Why there are complementary learning systems in the hippocampus and neocortex: insights from the successes and failures of connectionist models of learning and memory
James L McClelland, Bruce L McNaughton, and Randall C O’Reilly · 1995
Earlier work this paper cites.
Catastrophic forgetting, rehearsal and pseudorehearsal
Anthony Robins · 1995
Earlier work this paper cites.
Beyond “genetic programming": Incremental self-improvement
Jürgen Schmidhuber · 1995
Earlier work this paper cites.
Multitask learning
Rich Caruana · 1997
Earlier work this paper cites.
Shifting inductive bias with success-story algorithm, adaptive Levin search, and incremental self-improvement
Jürgen Schmidhuber, Jieyu Zhao, and Marco Wiering · 1997
Earlier work this paper cites.
The MNIST database of handwritten digits
Yann LeCun, Corinna Cortes, and Christopher JC Burges · 1998
Earlier work this paper cites.
Lifelong learning algorithms
Sebastian Thrun · 1998
Earlier work this paper cites.
Catastrophic forgetting in connectionist networks
Robert M French · 1999
Earlier work this paper cites.
Fixed-weight on-line learning
A Steven Younger, Peter R Conwell, and Neil E Cotter · 1999
Earlier work this paper cites.
Learning to learn using gradient descent
Sepp Hochreiter, A. Steven Younger, and Peter R. Conwell · 2001
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky · 2009
Earlier work this paper cites.
Notmnist dataset
Yaroslav Bulatov · 2011
Earlier work this paper cites.
Reading digits in natural images with unsupervised feature learning
Yuval Netzer, Tao Wang, Adam Coates, Alessandro Bissacco, Baolin Wu, Andrew Y Ng, et al · 2011
Earlier work this paper cites.
Compete to compute
Rupesh Kumar Srivastava, Jonathan Masci, Sohrob Kazerounian, Faustino J. Gomez, and Jürgen Schmidhuber · 2013
Earlier work this paper cites.
Learning to learn neural networks
Tom Bosc · 2015
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Sergey Ioffe and Christian Szegedy · 2015
Earlier work this paper cites.
Siamese neural networks for one-shot image recognition
Gregory Koch, Richard Zemel, Ruslan Salakhutdinov, et al · 2015
Cited alongside, same era.
Human-level concept learning through probabilistic program induction
Brenden M Lake, Ruslan Salakhutdinov, and Joshua B Tenenbaum · 2015
Cited alongside, same era.
RL 2 : Fast reinforcement learning via slow reinforcement learning
Yan Duan, John Schulman, Xi Chen, Peter L Bartlett, Ilya Sutskever, and Pieter Abbeel · 2016
Cited alongside, same era.
Learning without forgetting
Zhizhong Li and Derek Hoiem · 2016
Cited alongside, same era.
Andrei A Rusu, Neil C Rabinowitz, Guillaume Desjardins, Hubert Soyer, James Kirkpatrick, Koray Kavukcuoglu, Razvan Pascanu, and Raia Hadsell · 2016
Cited alongside, same era.
Metalearned neural memory
Tsendsuren Munkhdalai, Alessandro Sordoni, Tong Wang, and Adam Trischler · 2019
Later among the works it cites.
Pytorch: An imperative style, high-performance deep learning library
Adam Paszke et al · 2019
Later among the works it cites.
Fast and flexible multi-task classification using conditional neural adaptive processes
James Requeima, Jonathan Gordon, John Bronskill, Sebastian Nowozin, and Richard E. Turner · 2019
Later among the works it cites.
Learning to learn without forgetting by maximizing transfer and minimizing interference
Matthew Riemer, Ignacio Cases, Robert Ajemian, Miao Liu, Irina Rish, Yuhai Tu, and Gerald Tesauro · 2019
Later among the works it cites.
Experience replay for continual learning
David Rolnick, Arun Ahuja, Jonathan Schwarz, Timothy P. Lillicrap, and Gregory Wayne · 2019
Later among the works it cites.
Learning to continually learn
Shawn Beaulieu, Lapo Frati, Thomas Miconi, Joel Lehman, Kenneth O. Stanley, Jeff Clune, and Nick Cheney · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Meta-learning with memory-augmented neural networks
Adam Santoro, Sergey Bartunov, Matthew Botvinick, Daan Wierstra, and Timothy P. Lillicrap · 2016
Cited alongside, same era.
Instance normalization: The missing ingredient for fast stylization
Dmitry Ulyanov, Andrea Vedaldi, and Victor Lempitsky · 2016
Cited alongside, same era.
Matching networks for one shot learning
Oriol Vinyals, Charles Blundell, Tim Lillicrap, Koray Kavukcuoglu, and Daan Wierstra · 2016
Cited alongside, same era.
Model-agnostic meta-learning for fast adaptation of deep networks
Chelsea Finn, Pieter Abbeel, and Sergey Levine · 2017
Cited alongside, same era.
Automated curriculum learning for neural networks
Alex Graves, Marc G. Bellemare, Jacob Menick, Rémi Munos, and Koray Kavukcuoglu · 2017
Cited alongside, same era.
Overcoming catastrophic forgetting in neural networks
James Kirkpatrick, Razvan Pascanu, Neil Rabinowitz, Joel Veness, Guillaume Desjardins, Andrei A Rusu, Kieran Milan, John Quan, Tiago Ramalho, Agnieszka Grabska-Barwinska, et al · 2017
Cited alongside, same era.
Gradient episodic memory for continual learning
David Lopez-Paz and Marc’Aurelio Ranzato · 2017
Cited alongside, same era.
Later among the works it cites.
TaskNorm: Rethinking batch normalization for meta-learning
John Bronskill, Jonathan Gordon, James Requeima, Sebastian Nowozin, and Richard E. Turner · 2020
Later among the works it cites.
Language models are few-shot learners
Tom B Brown et al · 2020
Later among the works it cites.
Online fast adaptation and knowledge accumulation (OSAKA): a new approach to continual learning
Massimo Caccia, Pau Rodríguez, Oleksiy Ostapenko, Fabrice Normandin, Min Lin, Lucas Page-Caccia, Issam Hadj Laradji, Irina Rish, Alexandre Lacoste, David Vázquez, and Laurent Charlin · 2020
Later among the works it cites.
Livewired: The inside story of the ever-changing brain
David Eagleman · 2020
Later among the works it cites.
Adversarial continual learning
Sayna Ebrahimi, Franziska Meier, Roberto Calandra, Trevor Darrell, and Marcus Rohrbach · 2020
Later among the works it cites.
Transformers are RNNs: Fast autoregressive transformers with linear attention
Angelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, and François Fleuret · 2020
Later among the works it cites.
Meta-dataset: A dataset of datasets for learning to learn from few examples
Eleni Triantafillou, Tyler Zhu, Vincent Dumoulin, Pascal Lamblin, Utku Evci, Kelvin Xu, Ross Goroshin, Carles Gelada, Kevin Swersky, Pierre-Antoine Manzagol, and Hugo Larochelle · 2020
Later among the works it cites.
Generative vs. discriminative: Rethinking the meta-continual learning
Mohammadamin Banayeeanzade, Rasoul Mirzaiezadeh, Hosein Hasani, and Mahdieh Soleymani · 2021
Later among the works it cites.
Rethinking attention with performers
Krzysztof Choromanski, Valerii Likhosherstov, David Dohan, Xingyou Song, Andreea Gane, Tamas Sarlos, Peter Hawkins, Jared Davis, Afroz Mohiuddin, Lukasz Kaiser, et al · 2021
Later among the works it cites.
The devil is in the detail: Simple tricks improve systematic generalization of transformers
Róbert Csordás, Kazuki Irie, and Jürgen Schmidhuber · 2021
Later among the works it cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby · 2021
Later among the works it cites.
Meta learning backpropagation and improving it
Louis Kirsch and Jürgen Schmidhuber · 2021
Later among the works it cites.
Random feature attention
Hao Peng, Nikolaos Pappas, Dani Yogatama, Roy Schwartz, Noah A Smith, and Lingpeng Kong · 2021
Later among the works it cites.
Meta-learning bidirectional update rules
Mark Sandler, Max Vladymyrov, Andrey Zhmoginov, Nolan Miller, Tom Madams, Andrew Jackson, and Blaise Agüera y Arcas · 2021
Later among the works it cites.
Linear Transformers are secretly fast weight programmers
Imanol Schlag, Kazuki Irie, and Jürgen Schmidhuber · 2021
Later among the works it cites.
MLP-Mixer: An all-MLP architecture for vision
Ilya O. Tolstikhin, Neil Houlsby, Alexander Kolesnikov, Lucas Beyer, Xiaohua Zhai, Thomas Unterthiner, Jessica Yung, Andreas Steiner, Daniel Keysers, Jakob Uszkoreit, Mario Lucic, and Alexey Dosovitskiy · 2021
Later among the works it cites.
Efficient continual learning with modular networks and task-driven priors
Tom Veniat, Ludovic Denoyer, and Marc’Aurelio Ranzato · 2021
Later among the works it cites.
Addressing catastrophic forgetting in few-shot problems
Pau Ching Yap, Hippolyt Ritter, and David Barber · 2021
Later among the works it cites.
A simple but strong baseline for online continual learning: Repeated augmented rehearsal
Yaqian Zhang, Bernhard Pfahringer, Eibe Frank, Albert Bifet, Nick Jin Sean Lim, and Yunzhe Jia · 2022
Later among the works it cites.
Meta-in-context learning in large language models
Julian Coda-Forno, Marcel Binz, Zeynep Akata, Matthew Botvinick, Jane X Wang, and Eric Schulz · 2023
Closest in time.
Why can GPT learn in-context? language models secretly perform gradient descent as meta-optimizers
Damai Dai, Yutao Sun, Li Dong, Yaru Hao, Shuming Ma, Zhifang Sui, and Furu Wei · 2023
Closest in time.
Are LSTMs good few-shot learners?
Mike Huisman, Thomas M Moerland, Aske Plaat, and Jan N van Rijn · 2023
Closest in time.
Practical computational power of linear transformers and their recurrent and self-referential extensions
Kazuki Irie, Róbert Csordás, and Jürgen Schmidhuber · 2023
Closest in time.
Recasting continual learning as sequence modeling
Soochan Lee, Jaehyeon Son, and Gunhee Kim · 2023
Closest in time.
Addressing loss of plasticity and catastrophic forgetting in continual learning
Mohamed Elsayed and A. Rupam Mahmood · 2024
Closest in time.
Learning to learn without forgetting using attention
Anna Vettoruzzo, Joaquin Vanschoren, Mohamed-Rafik Bouguelia, and Thorsteinn Rögnvaldsson · 2024
Closest in time.