Fetching the paper…
Reading the bibliography…
In a standard view of the reinforcement learning problem, an agent's goal is to efficiently identify a policy that maximizes long-term reward.
Reinforcement learning, bit by bit
Xiuyuan Lu, Benjamin Van Roy, Vikranth Dwaracherla, Morteza Ibrahimi, Ian Osband, and Zheng Wen · 1935
Earlier work this paper cites.
Essays in positive economics
Milton Friedman · 1953
Earlier work this paper cites.
Catastrophic interference in connectionist networks: The sequential learning problem
Michael McCloskey and Neal J Cohen · 1989
Earlier work this paper cites.
Learning with a slowly changing distribution
Peter L Bartlett · 1992
Earlier work this paper cites.
Introduction: The challenge of reinforcement learning
Richard S Sutton · 1992
Earlier work this paper cites.
Acting optimally in partially observable stochastic domains
Anthony R. Cassandra, Leslie Pack Kaelbling, and Michael L. Littman · 1994
Earlier work this paper cites.
Continual learning in reinforcement environments
Mark B Ring · 1994
Earlier work this paper cites.
Provably bounded-optimal agents
Stuart J Russell and Devika Subramanian · 1994
Earlier work this paper cites.
Is learning the n-th thing any easier than learning the first?
Sebastian Thrun · 1995
Earlier work this paper cites.
Lifelong robot learning
Sebastian Thrun and Tom M Mitchell · 1995
Earlier work this paper cites.
Child: A first step towards continual learning
Mark B Ring · 1997
Earlier work this paper cites.
Reinforcement learning with self-modifying policies
Jürgen Schmidhuber, Jieyu Zhao, and Nicol N Schraudolph · 1998
Earlier work this paper cites.
Lifelong learning algorithms
Sebastian Thrun · 1998
Earlier work this paper cites.
Catastrophic forgetting in connectionist networks
Robert M French · 1999
Earlier work this paper cites.
A theory of universal artificial intelligence based on algorithmic complexity
Marcus Hutter · 2000
Earlier work this paper cites.
Online learning of non-stationary sequences
Claire Monteleoni and Tommi Jaakkola · 2003
Earlier work this paper cites.
Universal artificial intelligence: Sequential decisions based on algorithmic probability
Marcus Hutter · 2004
Earlier work this paper cites.
The reward hypothesis, 2004
Richard S Sutton · 2004
Earlier work this paper cites.
Toward a formal framework for continual learning
Mark B Ring · 2005
Earlier work this paper cites.
Multi-task reinforcement learning: a hierarchical Bayesian approach
Aaron Wilson, Alan Fern, Soumya Ray, and Prasad Tadepalli · 2007
Earlier work this paper cites.
Transfer learning for reinforcement learning domains: A survey
Matthew E. Taylor and Peter Stone · 2009
Earlier work this paper cites.
Toward an architecture for never-ending language learning
Andrew Carlson, Justin Betteridge, Bryan Kisiel, Burr Settles, Estevam R Hruschka, and Tom Mitchell · 2010
Earlier work this paper cites.
Metalearning
Tom Schaul and Jürgen Schmidhuber · 2010
Cited alongside, same era.
Machine lifelong learning: Challenges and benefits for artificial general intelligence
Daniel L Silver · 2011
Cited alongside, same era.
An empirical investigation of catastrophic forgetting in gradient-based neural networks
Ian J Goodfellow, Mehdi Mirza, Da Xiao, Aaron Courville, and Yoshua Bengio · 2013
Cited alongside, same era.
ELLA: An efficient lifelong learning algorithm
Paul Ruvolo and Eric Eaton · 2013
Cited alongside, same era.
Stochastic multi-armed-bandit problem with non-stationary rewards
Omar Besbes, Yonatan Gur, and Assaf Zeevi · 2014
Cited alongside, same era.
PAC-inspired option discovery in lifelong reinforcement learning
Emma Brunskill and Lihong Li · 2014
Experience replay for continual learning
David Rolnick, Arun Ahuja, Jonathan Schwarz, Timothy Lillicrap, and Gregory Wayne · 2019
Later among the works it cites.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Later among the works it cites.
Embracing change: Continual learning in deep neural networks
Raia Hadsell, Dushyant Rao, Andrei A Rusu, and Razvan Pascanu · 2020
Later among the works it cites.
Continual learning for robotics: Definition, framework, learning strategies, opportunities and challenges
Timothée Lesort, Vincenzo Lomonaco, Andrei Stoian, Davide Maltoni, David Filliat, and Natalia Díaz-Rodríguez · 2020
Later among the works it cites.
Jelly bean world: A testbed for never-ending learning
Emmanouil Antonios Platanios, Abulhair Saparov, and Tom Mitchell · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Online learning in Markov decision processes with changing cost sequences
Travis Dick, András György, and Csaba Szepesvari · 2014
Cited alongside, same era.
Theory of general reinforcement learning
Tor Lattimore · 2014
Cited alongside, same era.
Markov Decision Processes: Discrete Stochastic Dynamic Programming
Martin L Puterman · 2014
Cited alongside, same era.
Safe policy search for lifelong reinforcement learning with sublinear regret
Haitham Bou Ammar, Rasul Tutunov, and Eric Eaton · 2015
Cited alongside, same era.
Nonparametric general reinforcement learning
Jan Leike · 2016
Cited alongside, same era.
Model-agnostic meta-learning for fast adaptation of deep networks
Chelsea Finn, Pieter Abbeel, and Sergey Levine · 2017
Cited alongside, same era.
Shibhansh Dohare, Richard S Sutton, and A Rupam Mahmood · 2021
Later among the works it cites.
Abstractions of general reinforcement Learning
Sultan J Majeed · 2021
Later among the works it cites.
Simple agent, complex environment: Efficient reinforcement learning with agent states
Shi Dong, Benjamin Van Roy, and Zhengyuan Zhou · 2022
Later among the works it cites.
Model-based lifelong reinforcement learning with Bayesian exploration
Haotian Fu, Shangqun Yu, Michael Littman, and George Konidaris · 2022
Later among the works it cites.
Towards continual reinforcement learning: A review and perspectives
Khimya Khetarpal, Matthew Riemer, Irina Rish, and Doina Precup · 2022
Later among the works it cites.
Meta-gradients in non-stationary environments
Jelena Luketina, Sebastian Flennerhag, Yannick Schroecker, David Abel, Tom Zahavy, and Satinder Singh · 2022
Later among the works it cites.
Online continual learning in image classification: An empirical survey
Zheda Mai, Ruiwen Li, Jihwan Jeong, David Quispe, Hyunwoo Kim, and Scott Sanner · 2022
Later among the works it cites.
Continual learning in environments with polynomial mixing times
Matthew Riemer, Sharath Chandra Raparthy, Ignacio Cases, Gopeshh Subbaraj, Maximilian Puelma Touzel, and Irina Rish · 2022
Later among the works it cites.
The quest for a common model of the intelligent decision maker
Richard S Sutton · 2022
Later among the works it cites.
Loss of plasticity in continual deep reinforcement learning
Zaheer Abbas, Rosie Zhao, Joseph Modayil, Adam White, and Marlos C Machado · 2023
Closest in time.
A domain-agnostic approach for characterization of lifelong learning systems
Megan M Baker, Alexander New, Mario Aguilar-Simon, Ziad Al-Halah, Sébastien MR Arnold, Ese Ben-Iwhiwhu, Andrew P Brna, Ethan Brooks, Ryan C Brown, Zachary Daniels, et al · 2023
Closest in time.
Settling the reward hypothesis
Michael Bowling, John D. Martin, David Abel, and Will Dabney · 2023
Closest in time.
Loss of plasticity in deep continual learning
Shibhansh Dohare, Juan Hernandez-Garcia, Parash Rahman, Richard Sutton, and A Rupam Mahmood · 2023
Closest in time.
Continual learning as computationally constrained reinforcement learning
Saurabh Kumar, Henrik Marklund, Ashish Rao, Yifan Zhu, Hong Jun Jeon, Yueyang Liu, and Benjamin Van Roy · 2023
Closest in time.
A definition of non-stationary bandits
Yueyang Liu, Benjamin Van Roy, and Kuang Xu · 2023
Closest in time.
Understanding plasticity in neural networks
Clare Lyle, Zeyu Zheng, Evgenii Nikishin, Bernardo Avila Pires, Razvan Pascanu, and Will Dabney · 2023
Closest in time.