Fetching the paper…
Reading the bibliography…
Temporal difference (TD) learning is often used to update the estimate of the value function which is used by RL agents to extract useful policies.
A massively parallel architecture for a self-organizing neural pattern recognition machine
Gail A Carpenter and Stephen Grossberg · 1987
Earlier work this paper cites.
Evolutionary principles in self-referential learning, or on learning how to learn: the meta-meta-… hook
Jürgen Schmidhuber · 1987
Earlier work this paper cites.
Learning to predict by the methods of temporal differences
Richard S Sutton · 1988
Earlier work this paper cites.
Learning from delayed rewards
Christopher John Cornish Hellaby Watkins · 1989
Earlier work this paper cites.
Adapting bias by gradient descent: An incremental version of delta-bar-delta
Richard S Sutton · 1992
Earlier work this paper cites.
Neuro-dynamic programming
Dimitri Bertsekas and John N Tsitsiklis · 1996
Earlier work this paper cites.
Multitask learning
Rich Caruana · 1997
Earlier work this paper cites.
Child: A first step towards continual learning
Mark B Ring · 1997
Earlier work this paper cites.
Analytical mean squared error curves for temporal difference learning
Satinder Singh and Peter Dayan · 1998
Earlier work this paper cites.
Reinforcement learning of local shape in the game of go
David Silver, Richard S Sutton, and Martin Müller · 2007
Earlier work this paper cites.
On the role of tracking in stationary environments
Richard S Sutton, Anna Koop, and David Silver · 2007
Earlier work this paper cites.
Sample-based learning and search with permanent and transient memories
David Silver, Richard S Sutton, and Martin Müller · 2008
Earlier work this paper cites.
Transfer learning for reinforcement learning domains: A survey
Matthew E Taylor and Peter Stone · 2009
Earlier work this paper cites.
Horde: A scalable real-time architecture for learning knowledge from unsupervised sensorimotor interaction
Richard S Sutton, Joseph Modayil, Michael Delp, Thomas Degris, Patrick M Pilarski, Adam White, and Doina Precup · 2011
Earlier work this paper cites.
A survey on concept drift adaptation
João Gama, Indrė Žliobaitė, Albert Bifet, Mykola Pechenizkiy, and Abdelhamid Bouchachia · 2014
Earlier work this paper cites.
What learning systems do intelligent agents need? complementary learning systems theory updated
Dharshan Kumaran, Demis Hassabis, and James L McClelland · 2016
Earlier work this paper cites.
Mastering the game of go with deep neural networks and tree search
David Silver, Aja Huang, Chris J Maddison, Arthur Guez, Laurent Sifre, George Van Den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Veda Panneershelvam, Marc Lanctot, et al · 2016
Earlier work this paper cites.
Model-agnostic meta-learning for fast adaptation of deep networks
Chelsea Finn, Pieter Abbeel, and Sergey Levine · 2017
Cited alongside, same era.
Overcoming catastrophic forgetting in neural networks
James Kirkpatrick, Razvan Pascanu, Neil Rabinowitz, Joel Veness, Guillaume Desjardins, Andrei A Rusu, Kieran Milan, John Quan, Tiago Ramalho, Agnieszka Grabska-Barwinska, et al · 2017
Cited alongside, same era.
The persistence and transience of memory
Blake A Richards and Paul W Frankland · 2017
Cited alongside, same era.
Mastering the game of go without human knowledge
David Silver, Julian Schrittwieser, Karen Simonyan, Ioannis Antonoglou, Aja Huang, Arthur Guez, Thomas Hubert, Lucas Baker, Matthew Lai, Adrian Bolton, et al · 2017
Cited alongside, same era.
Distral: Robust multitask reinforcement learning
Yee Teh, Victor Bapst, Wojciech M Czarnecki, John Quan, James Kirkpatrick, Raia Hadsell, Nicolas Heess, and Razvan Pascanu · 2017
Cited alongside, same era.
Policy and value transfer in lifelong reinforcement learning
Minatar: An atari-inspired testbed for thorough and reproducible reinforcement learning experiments
Kenny Young and Tian Tian · 2019
Later among the works it cites.
Continual reinforcement learning with multi-timescale replay
Christos Kaplanis, Claudia Clopath, and Murray Shanahan · 2020
Later among the works it cites.
Jelly bean world: A testbed for never-ending learning
Emmanouil Antonios Platanios, Abulhair Saparov, and Tom Mitchell · 2020
Later among the works it cites.
Continual backprop: Stochastic gradient descent with persistent randomness
Shibhansh Dohare, Richard S Sutton, and A Rupam Mahmood · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
David Abel, Yuu Jinnai, Sophie Yue Guo, George Konidaris, and Michael Littman · 2018
Cited alongside, same era.
A finite time analysis of temporal difference learning with linear function approximation
Jalaj Bhandari, Daniel Russo, and Raghav Singal · 2018
Cited alongside, same era.
Minimalistic gridworld environment for openai gym, 2018
Maxime Chevalier-Boisvert, Lucas Willems, and Suman Pal · 2018
Cited alongside, same era.
Continual reinforcement learning with complex synapses
Christos Kaplanis, Murray Shanahan, and Claudia Clopath · 2018
Cited alongside, same era.
Learning to learn without forgetting by maximizing transfer and minimizing interference
Matthew Riemer, Ignacio Cases, Robert Ajemian, Miao Liu, Irina Rish, Yuhai Tu, and Gerald Tesauro · 2018
Cited alongside, same era.
Reinforcement learning: An introduction
Richard S Sutton and Andrew G Barto · 2018
Cited alongside, same era.
Reinforcement learning, fast and slow
Matthew Botvinick, Sam Ritter, Jane X Wang, Zeb Kurth-Nelson, Charles Blundell, and Demis Hassabis · 2019
Cited alongside, same era.
Sebastian Flennerhag, Yannick Schroecker, Tom Zahavy, Hado van Hasselt, David Silver, and Satinder Singh · 2021
Later among the works it cites.
Elahe Arani, Fahad Sarfraz, and Bahram Zonooz · 2022
Later among the works it cites.
Lifelong reinforcement learning with modulating masks
Eseoghene Ben-Iwhiwhu, Saptarshi Nath, Praveen K Pilly, Soheil Kolouri, and Andrea Soltoggio · 2022
Later among the works it cites.
Task-agnostic continual reinforcement learning: In praise of a simple baseline
Massimo Caccia, Jonas Mueller, Taesup Kim, Laurent Charlin, and Rasool Fakoor · 2022
Later among the works it cites.
Neural distillation as a state representation bottleneck in reinforcement learning
Valentin Guillet, Dennis George Wilson, and Emmanuel Rachelson · 2022
Later among the works it cites.
Same state, different task: Continual reinforcement learning without interference
Samuel Kessler, Jack Parker-Holder, Philip Ball, Stefan Zohren, and Stephen J Roberts · 2022
Later among the works it cites.
Towards continual reinforcement learning: A review and perspectives
Khimya Khetarpal, Matthew Riemer, Irina Rish, and Doina Precup · 2022
Later among the works it cites.
Understanding and preventing capacity loss in reinforcement learning
Clare Lyle, Mark Rowland, and Will Dabney · 2022
Later among the works it cites.
The primacy bias in deep reinforcement learning
Evgenii Nikishin, Max Schwarzer, Pierluca D’Oro, Pierre-Luc Bacon, and Aaron Courville · 2022
Later among the works it cites.
Self-activating neural ensembles for continual reinforcement learning
Sam Powers, Eliot Xing, and Abhinav Gupta · 2022
Later among the works it cites.
A definition of continual reinforcement learning
David Abel, André Barreto, Benjamin Van Roy, Doina Precup, Hado van Hasselt, and Satinder Singh · 2023
Closest in time.
Utility-based perturbed gradient descent: An optimizer for continual learning
Mohamed Elsayed and A Rupam Mahmood · 2023
Closest in time.