Fetching the paper…
Reading the bibliography…
In lifelong learning, an agent learns throughout its entire life without resets, in a constantly changing environment, as we humans do.
Individual Comparisons by Ranking Methods
Frank Wilcoxon · 1945
Earlier work this paper cites.
The organization of behavior: A neuropsychological theory
Donald O. Hebb · 1949
Earlier work this paper cites.
Counterintuitive behavior of social systems
Jay W. Forrester · 1971
Earlier work this paper cites.
Temporal credit assignment in reinforcement learning
Richard Stuart Sutton · 1984
Earlier work this paper cites.
Art 2: self-organization of stable category recognition codes for analog input patterns
Gail A. Carpenter and Stephen Grossberg · 1987
Earlier work this paper cites.
Evolutionary principles in self-referential learning, or on learning how to learn: the meta-meta-… hook
Jürgen Schmidhuber · 1987
Earlier work this paper cites.
Catastrophic interference in connectionist networks: The sequential learning problem
Michael McCloskey and Neal J. Cohen · 1989
Earlier work this paper cites.
Making the world differentiable: On using fully recurrent self-supervised neural networks for dynamic reinforcement learning and planning in non-stationary environments
Jürgen Schmidhuber · 1990
Earlier work this paper cites.
A possibility for implementing curiosity and boredom in model-building neural controllers
Jürgen Schmidhuber · 1991
Earlier work this paper cites.
On the optimization of a synaptic learning rule
Samy Bengio, Yoshua Bengio, Jocelyn Cloutier, and Jan Gecsei · 1992
Earlier work this paper cites.
Self-Improving Reactive Agents Based On Reinforcement Learning, Planning and Teaching
Long Ji Lin · 1992
Earlier work this paper cites.
Q-learning
Christopher J. C. H. Watkins and Peter Dayan · 1992
Earlier work this paper cites.
Neural network exploration using optimal experiment design
David A. Cohn · 1994
Earlier work this paper cites.
Continual learning in reinforcement environments
Mark Bishop Ring et al · 1994
Earlier work this paper cites.
Reinforcement driven information acquisition in non-deterministic environments
Jan Storck, Sepp Hochreiter, and Jürgen Schmidhuber · 1995
Earlier work this paper cites.
Lifelong robot learning
Sebastian Thrun and Tom M Mitchell · 1995
Earlier work this paper cites.
Detection of abrupt changes: Theory and application
Michèle Basseville and Igor V. Nikiforov · 1996
Earlier work this paper cites.
Multitask learning
Rich Caruana · 1997
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Planning and Acting in Partially Observable Stochastic Domains
Leslie P. Kaelbling, Michael L. Littman, and Anthony R. Cassandra · 1998
Earlier work this paper cites.
Reinforcement learning: An introduction , volume 1
Richard S. Sutton and Andrew G. Barto · 1998
Earlier work this paper cites.
Lifelong learning algorithms
Sebastian Thrun · 1998
Earlier work this paper cites.
Catastrophic forgetting in connectionist networks
Robert M French · 1999
Earlier work this paper cites.
Automatic discovery of subgoals in reinforcement learning using diverse density
Amy McGovern and Andrew G. Barto · 2001
Earlier work this paper cites.
Subgoal discovery for hierarchical reinforcement learning using learned policies
Sandeep Goel and Manfred Huber · 2003
Earlier work this paper cites.
Detecting change in data streams
Daniel Kifer, Shai Ben-David, and Johannes Gehrke · 2004
Earlier work this paper cites.
Learning in the presence of concept drift and hidden contexts
Gerhard Widmer and Miroslav Kubát · 2004
Earlier work this paper cites.
Bayesian online changepoint detection, 2007
Ryan Prescott Adams and David J. C. MacKay · 2007
Earlier work this paper cites.
Change Point Detection and Meta-Bandits for Online Learning in Dynamic Environments
Cédric Hartland, Nicolas Baskiotis, Sylvain Gelly, Michèle Sebag, and Olivier Teytaud · 2007
Earlier work this paper cites.
What is intrinsic motivation? A typology of computational approaches
Pierre-Yves Oudeyer and Frédéric Kaplan · 2007
Earlier work this paper cites.
An analysis of model-based interval estimation for markov decision processes
Alexander L. Strehl and Michael L. Littman · 2007
Earlier work this paper cites.
On upper-confidence bound policies for non-stationary bandit problems, 2008
Aurélien Garivier and Eric Moulines · 2008
Earlier work this paper cites.
A theory of learning from different domains
Shai Ben-David, John Blitzer, Koby Crammer, Alex Kulesza, Fernando Pereira, and Jennifer Wortman Vaughan · 2010
Earlier work this paper cites.
Multimodal deep learning
Jiquan Ngiam, Aditya Khosla, Mingyu Kim, Juhan Nam, Honglak Lee, and Andrew Y Ng · 2011
Earlier work this paper cites.
Convergence Proof for Actor-Critic Methods Applied to PPO and RUDDER
Markus Holzleitner, Lukas Gruber, Jose A. Arjona-Medina, Johannes Brandstetter, and Sepp Hochreiter · 2012
Earlier work this paper cites.
Motion-dependent representation of space in area mt+
Gerrit W. Maus, Jason Fischer, and David Whitney · 2013
Earlier work this paper cites.
The stability-plasticity dilemma: Investigating the continuum from catastrophic forgetting to age-limited learning effects
Martial Mermillod, Aurélia Bugaiska, and Patrick Bonin · 2013
Earlier work this paper cites.
Playing atari with deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, and Martin Riedmiller · 2013
Cited alongside, same era.
Stochastic multi-armed-bandit problem with non-stationary rewards
Omar Besbes, Yonatan Gur, and Assaf Zeevi · 2014
Cited alongside, same era.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Cited alongside, same era.
Incentivizing exploration in reinforcement learning with deep predictive models
Bradly C. Stadie, Sergey Levine, and Pieter Abbeel · 2015
Cited alongside, same era.
Unifying count-based exploration and intrinsic motivation
Marc Bellemare, Sriram Srinivasan, Georg Ostrovski, Tom Schaul, David Saxton, and Remi Munos · 2016
Cited alongside, same era.
Continuous learning in single-incremental-task scenarios
Davide Maltoni and Vincenzo Lomonaco · 2019
Later among the works it cites.
Reinforcement learning in non-stationary environments
Sindhu Padakandla, Prabuchandran K. J., and Shalabh Bhatnagar · 2019
Later among the works it cites.
Stable baselines3, 2019
Antonin Raffin, Ashley Hill, Maximilian Ernestus, Adam Gleave, Anssi Kanervisto, and Noah Dormann · 2019
Later among the works it cites.
Weighted linear bandits for non-stationary environments
Yoan Russac, Claire Vernade, and Olivier Cappé · 2019
Later among the works it cites.
Forward and backward knowledge transfer for sentiment classification
Hao Wang, Bing Liu, Shuai Wang, Nianzu Ma, and Yan Yang · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Openai gym, 2016
Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wojciech Zaremba · 2016
Cited alongside, same era.
Andrei A. Rusu, Neil C. Rabinowitz, Guillaume Desjardins, Hubert Soyer, James Kirkpatrick, Koray Kavukcuoglu, Razvan Pascanu, and Raia Hadsell · 2016
Cited alongside, same era.
Prioritized experience replay
Tom Schaul, John Quan, Ioannis Antonoglou, and David Silver · 2016
Cited alongside, same era.
Surprise-based intrinsic motivation for deep reinforcement learning
Joshua Achiam and Shankar Sastry · 2017
Cited alongside, same era.
Pathnet: Evolution channels gradient descent in super neural networks
Chrisantha Fernando, Dylan S. Banarse, Charles Blundell, Yori Zwols, David R. Ha, Andrei A. Rusu, Alexander Pritzel, and Daan Wierstra · 2017
Cited alongside, same era.
Reinforcement learning with deep energy-based policies
Tuomas Haarnoja, Haoran Tang, Pieter Abbeel, and Sergey Levine · 2017
Cited alongside, same era.
Overcoming catastrophic forgetting in neural networks
James Kirkpatrick, Razvan Pascanu, Neil Rabinowitz, Joel Veness, Guillaume Desjardins, Andrei A Rusu, Kieran Milan, John Quan, Tiago Ramalho, Agnieszka Grabska-Barwinska, et al · 2017
Cited alongside, same era.
Cross-domain few-shot learning by representation fusion
Thomas Adler, Johannes Brandstetter, Michael Widrich, Andreas Mayr, David Kreil, Michael Kopp, Günter Klambauer, and Sepp Hochreiter · 2020
Later among the works it cites.
Never give up: Learning directed exploration strategies
Adrià Puigdomènech Badia, Pablo Sprechmann, Alex Vitvitskyi, Daniel Guo, Bilal Piot, Steven Kapturowski, Olivier Tieleman, Martín Arjovsky, Alexander Pritzel, Andew Bolt, et al · 2020
Later among the works it cites.
Model-free non-stationarity detection and adaptation in reinforcement learning
Giuseppe Canonaco, Marcello Restelli, and Manuel Roveri · 2020
Later among the works it cites.
Optimizing for the future in non-stationary mdps
Yash Chandak, Georgios Theocharous, Shiv Shankar, Martha White, Sridhar Mahadevan, and Philip S. Thomas · 2020
Later among the works it cites.
Reinforcement learning for non-stationary markov decision processes: The blessing of (more) optimism
Wang Chi Cheung, David Simchi-Levi, and Ruihao Zhu · 2020
Later among the works it cites.
Continual learning in recurrent neural networks with hypernetworks
Benjamin Ehret, Christian Henning, Maria R Cervera, Alexander Meulemans, Johannes von Oswald, and Benjamin F Grewe · 2020
Later among the works it cites.
La-maml: Look-ahead meta learning for continual learning
Gunshi Gupta, Karmesh Yadav, and Liam Paull · 2020
Later among the works it cites.
Neural replicator dynamics: Multiagent learning via hedging policy gradients
Daniel Hennes, Dustin Morrill, Shayegan Omidshafiei, Rémi Munos, Julien Perolat, Marc Lanctot, Audrunas Gruslys, Jean-Baptiste Lespiau, Paavo Parmas, Edgar Duèñez Guzmán, and Karl Tuyls · 2020
Later among the works it cites.
The impact of non-stationarity on generalisation in deep reinforcement learning
Maximilian Igl, Gregory Farquhar, Jelena Luketina, Wendelin Boehmer, and Shimon Whiteson · 2020
Later among the works it cites.
Towards continual reinforcement learning: A review and perspectives
Khimya Khetarpal, Matthew Riemer, Irina Rish, and Doina Precup · 2020
Later among the works it cites.
Bandit algorithms
Tor Lattimore and Csaba Szepesvári · 2020
Later among the works it cites.
Count-based exploration with the successor representation
Marlos C. Machado, Marc G. Bellemare, and Michael Bowling · 2020
Later among the works it cites.
Align-rudder: Learning from few demonstrations by reward redistribution
Vihang P Patil, Markus Hofmarcher, Marius-Constantin Dinu, Matthias Dorfer, Patrick M Blies, Johannes Brandstetter, Jose A Arjona-Medina, and Sepp Hochreiter · 2020
Later among the works it cites.
Jelly bean world: A testbed for never-ending learning
Emmanouil Antonios Platanios, Abulhair Saparov, and Tom Mitchell · 2020
Later among the works it cites.
RIDE: rewarding impact-driven exploration for procedurally-generated environments
Roberta Raileanu and Tim Rocktäschel · 2020
Later among the works it cites.
Supermasks in superposition
Mitchell Wortsman, Vivek Ramanujan, Rosanne Liu, Aniruddha Kembhavi, Mohammad Rastegari, Jason Yosinski, and Ali Farhadi · 2020
Later among the works it cites.
A simple approach for non-stationary linear bandits
Peng Zhao, Lijun Zhang, Yuan Jiang, and Zhi-Hua Zhou · 2020
Later among the works it cites.
Minimum-delay adaptation in non-stationary reinforcement learning via online high-confidence change-point detection
Lucas Nunes Alegre, Ana L. C. Bazzan, and Bruno C. da Silva · 2021
Later among the works it cites.
Non-stationary off-policy optimization
Joey Hong, Branislav Kveton, Manzil Zaheer, Yinlam Chow, and Amr Ahmed · 2021
Later among the works it cites.
Modern Hopfield Networks for Return Decomposition for Delayed Rewards
Michael Widrich, Markus Hofmarcher, Vihang Prakash Patil, Angela Bitto-Nemling, and Sepp Hochreiter · 2021
Later among the works it cites.
Continual world: A robotic benchmark for continual reinforcement learning
Maciej Wołczyk, Michał Zając, Razvan Pascanu, Lukasz Kucinski, and Piotr Miłoś · 2021
Later among the works it cites.
Deep reinforcement learning amidst continual structured non-stationarity
Annie Xie, James Harrison, and Chelsea Finn · 2021
Later among the works it cites.
Task-agnostic continual learning using online variational bayes with fixed-point updates
Chen Zeno, Itay Golan, Elad Hoffer, and Daniel Soudry · 2021
Later among the works it cites.
Noveld: A simple yet effective exploration criterion
Tianjun Zhang, Huazhe Xu, Xiaolong Wang, Yi Wu, Kurt Keutzer, Joseph E. Gonzalez, and Yuandong Tian · 2021
Later among the works it cites.
XAI and Strategy Extraction via Reward Redistribution , pp. 177–205
Marius-Constantin Dinu, Markus Hofmarcher, Vihang P. Patil, Matthias Dorfer, Patrick M. Blies, Johannes Brandstetter, Jose A. Arjona-Medina, and Sepp Hochreiter · 2022
Closest in time.
Bootstrapped meta-learning
Sebastian Flennerhag, Yannick Schroecker, Tom Zahavy, Hado van Hasselt, David Silver, and Satinder Singh · 2022
Closest in time.
Same state, different task: Continual reinforcement learning without interference
Samuel Kessler, Jack Parker-Holder, Philip Ball, Stefan Zohren, and Stephen J Roberts · 2022
Closest in time.
Lifelong hyper-policy optimization with multiple importance sampling regularization
Pierre Liotet, Francesco Vidaich, Alberto Maria Metelli, and Marcello Restelli · 2022
Closest in time.
History compression via language models in reinforcement learning
Fabian Paischer, Thomas Adler, Vihang Patil, Angela Bitto-Nemling, Markus Holzleitner, Sebastian Lehner, Hamid Eghbal-zadeh, and Sepp Hochreiter · 2022
Closest in time.
Change point detection for compositional multivariate data
K. J. Prabuchandran, Nitin Singh, Pankaj Dayama, Ashutosh Agarwal, and Vinayaka Pandit · 2022
Closest in time.
Infinite steps cartpole problem with variable reward, 2020
Suraj Regmi · 2022
Closest in time.
Abstraction for deep reinforcement learning
Murray Shanahan and Melanie Mitchell · 2022
Closest in time.