Fetching the paper…
Reading the bibliography…
In this article, we aim to provide a literature review of different formulations and approaches to continual reinforcement learning (RL), also known as lifelong or non-stationary RL.
Provably efficient rl with rich observations via latent state decoding
Simon S Du, Akshay Krishnamurthy, Nan Jiang, Alekh Agarwal, Miroslav Dudík, and John Langford · 1901
Earlier work this paper cites.
Meta-learnt priors slow down catastrophic forgetting in neural networks
Giacomo Spigler · 1909
Earlier work this paper cites.
Is a good representation sufficient for sample efficient reinforcement learning?
Simon S Du, Sham M Kakade, Ruosong Wang, and Lin F Yang · 1910
Earlier work this paper cites.
Leveraging procedural generation to benchmark reinforcement learning
Karl Cobbe, Christopher Hesse, Jacob Hilton, and John Schulman · 1912
Earlier work this paper cites.
A note on measurement of utility
Paul A Samuelson · 1937
Earlier work this paper cites.
Bayesian decision problems and Markov chains
James John Martin · 1967
Earlier work this paper cites.
The theory of affordances
James J Gibson · 1977
Earlier work this paper cites.
A massively parallel architecture for a self-organizing neural pattern recognition machine
Gail A Carpenter and Stephen Grossberg · 1987
Earlier work this paper cites.
Evolutionary principles in self-referential learning, or on learning how to learn: the meta-meta-… hook
Jürgen Schmidhuber · 1987
Earlier work this paper cites.
Catastrophic interference in connectionist networks: The sequential learning problem
Michael McCloskey and Neal J Cohen · 1989
Earlier work this paper cites.
Affordances and the body: An intentional analysis of gibson’s ecological approach to visual perception
Harry Heft · 1989
Earlier work this paper cites.
Learning a synaptic learning rule
Yoshua Bengio, Samy Bengio, and Jocelyn Cloutier · 1990
Earlier work this paper cites.
Using semi-distributed representations to overcome catastrophic forgetting in connectionist networks
Robert M French · 1991
Earlier work this paper cites.
Curious model-building control systems
Jürgen Schmidhuber · 1991
Earlier work this paper cites.
Dyna, an integrated architecture for learning, planning, and reacting
Richard S Sutton · 1991
Earlier work this paper cites.
On the optimization of a synaptic learning rule
Samy Bengio, Yoshua Bengio, Jocelyn Cloutier, and Jan Gecsei · 1992
Earlier work this paper cites.
Self-improving reactive agents based on reinforcement learning, planning and teaching
Long-Ji Lin · 1992
Earlier work this paper cites.
Transfer of learning by composing solutions of elemental sequential tasks
Satinder Pal Singh · 1992
Earlier work this paper cites.
Learning to control fast-weight memories: An alternative to dynamic recurrent networks
Jürgen Schmidhuber · 1992
Earlier work this paper cites.
Adapting bias by gradient descent: an incremental version of the delta-bar-delta
Rich Sutton · 1992
Earlier work this paper cites.
Learning to achieve goals
Leslie Pack Kaelbling · 1993
Earlier work this paper cites.
Improving generalization for temporal difference learning: The successor representation
Peter Dayan · 1993
Earlier work this paper cites.
Markov decision processes. 1994
ML Puterman · 1994
Earlier work this paper cites.
Markov games as a framework for multi-agent reinforcement learning
Michael L. Littman · 1994
Earlier work this paper cites.
Catastrophic forgetting, rehearsal and pseudorehearsal
Anthony Robins · 1995
Earlier work this paper cites.
Finding structure in reinforcement learning
Sebastian Thrun and Anton Schwartz · 1995
Earlier work this paper cites.
Why there are complementary learning systems in the hippocampus and neocortex: insights from the successes and failures of connectionist models of learning and memory
James L McClelland, Bruce L McNaughton, and Randall C O’Reilly · 1995
Earlier work this paper cites.
Discovering structure in multiple learning tasks: The tc algorithm
Sebastian Thrun and Joseph O’Sullivan · 1996
Earlier work this paper cites.
Consolidation in neural networks and in the sleeping brain
Anthony Robins · 1996
Earlier work this paper cites.
Child: A first step towards continual learning
Mark B Ring · 1997
Earlier work this paper cites.
Multitask learning
Rich Caruana · 1997
Earlier work this paper cites.
Shifting inductive bias with success-story algorithm, adaptive levin search, and incremental self-improvement
Jürgen Schmidhuber, Jieyu Zhao, and Marco Wiering · 1997
Earlier work this paper cites.
Separated modules for visuomotor control and learning in the cerebellum: a functional mri study
Hiroshi Imamizu · 1997
Earlier work this paper cites.
Introduction to Reinforcement Learning
Richard S. Sutton and Andrew G. Barto · 1998
Earlier work this paper cites.
Multi-criteria reinforcement learning
Zoltán Gábor, Zsolt Kalmár, and Csaba Szepesvári · 1998
Earlier work this paper cites.
Reinforcement learning with hierarchies of machines
Ronald Parr and Stuart J Russell · 1998
Earlier work this paper cites.
Hierarchical solution of markov decision processes using macro-actions
Milos Hauskrecht, Nicolas Meuleau, Leslie Pack Kaelbling, Thomas Dean, and Craig Boutilier · 1998
Earlier work this paper cites.
Reinforcement learning with self-modifying policies
Jürgen Schmidhuber, Jieyu Zhao, and Nicol N Schraudolph · 1998
Earlier work this paper cites.
Internal models in the cerebellum
Daniel M Wolpert, R Chris Miall, and Mitsuo Kawato · 1998
Earlier work this paper cites.
Efficient reinforcement learning in factored mdps
Michael Kearns and Daphne Koller · 1999
Earlier work this paper cites.
Between mdps and semi-mdps: A framework for temporal abstraction in reinforcement learning
Richard S Sutton, Doina Precup, and Satinder Singh · 1999
Earlier work this paper cites.
A general method for incremental self-improvement and multi-agent learning
Juergen Schmidhuber · 1999
Earlier work this paper cites.
Catastrophic forgetting in simple networks: an analysis of the pseudorehearsal solution
Marcus Frean and Anthony Robins · 1999
Earlier work this paper cites.
Hidden-mode markov decision processes for nonstationary sequential decision making
Samuel PM Choi, Dit-Yan Yeung, and Nevin L Zhang · 2000
Earlier work this paper cites.
Stochastic dynamic programming with factored representations
Craig Boutilier, Richard Dearden, and Moisés Goldszmidt · 2000
Earlier work this paper cites.
Hierarchical reinforcement learning with the maxq value function decomposition
Thomas G Dietterich · 2000
Earlier work this paper cites.
Behavioral considerations suggest an average reward td model of the dopamine system
Nathaniel D Daw and David S Touretzky · 2000
Earlier work this paper cites.
Human cerebellar activity reflecting an acquired internal model of a new tool
Hiroshi Imamizu, Satoru Miyauchi, Tomoe Tamada, Yuka Sasaki, Ryousuke Takino, Benno PuÈtz, Toshinori Yoshioka, and Mitsuo Kawato · 2000
Earlier work this paper cites.
Symbolic dynamic programming for first-order mdps
Craig Boutilier, Raymond Reiter, and Bob Price · 2001
Earlier work this paper cites.
Gradient surgery for multi-task learning
Tianhe Yu, Saurabh Kumar, Abhishek Gupta, Sergey Levine, Karol Hausman, and Chelsea Finn · 2001
Earlier work this paper cites.
Options of interest: Temporal abstraction with interest functions
Khimya Khetarpal, Martin Klissarov, Maxime Chevalier-Boisvert, Pierre-Luc Bacon, and Doina Precup · 2001
Earlier work this paper cites.
Friend-or-foe q-learning in general-sum games
Michael L. Littman · 2001
Earlier work this paper cites.
A perspective view and survey of meta-learning
Ricardo Vilalta and Youssef Drissi · 2002
Earlier work this paper cites.
Multiple model-based reinforcement learning
Kenji Doya, Kazuyuki Samejima, Ken-ichi Katagiri, and Mitsuo Kawato · 2002
Earlier work this paper cites.
Optimal Learning: Computational procedures for Bayes-adaptive Markov decision processes
Michael O’Gordon Duff · 2002
Earlier work this paper cites.
Shawn Beaulieu, Lapo Frati, Thomas Miconi, Joel Lehman, Kenneth O Stanley, Jeff Clune, and Nick Cheney · 2002
Earlier work this paper cites.
Reinforcement learning to play an optimal nash equilibrium in team markov games
Xiaofeng Wang and Tuomas Sandholm · 2002
Earlier work this paper cites.
Generalizing plans to new environments in relational mdps
Carlos Guestrin, Daphne Koller, Chris Gearhart, and Neal Kanodia · 2003
Earlier work this paper cites.
Rui Wang, Joel Lehman, Aditya Rawal, Jiale Zhi, Yulun Li, Jeff Clune, and Kenneth O Stanley · 2003
Earlier work this paper cites.
An outline of a theory of affordances
Anthony Chemero · 2003
Earlier work this paper cites.
Correlated-Q learning
Amy Greenwald and Keith Hall · 2003
Earlier work this paper cites.
Reinforcement learning and its relationship to supervised learning
Andrew G Barto and Thomas G Dietterich · 2004
Earlier work this paper cites.
Convergence and no-regret in multiagent learning
Michael Bowling · 2004
Earlier work this paper cites.
Intrinsically motivated reinforcement learning
Nuttapong Chentanez, Andrew G Barto, and Satinder P Singh · 2005
Earlier work this paper cites.
Optimizing for the future in non-stationary mdps
Yash Chandak, Georgios Theocharous, Shiv Shankar, Sridhar Mahadevan, Martha White, and Philip S Thomas · 2005
Earlier work this paper cites.
Empowerment: A universal agent-centric measure of control
Alexander S Klyubin, Daniel Polani, and Chrystopher L Nehaniv · 2005
Earlier work this paper cites.
Uncertainty-based competition between prefrontal and dorsolateral striatal systems for behavioral control
Nathaniel D Daw, Yael Niv, and Peter Dayan · 2005
Earlier work this paper cites.
Dopamine, uncertainty and td learning
Yael Niv, Michael O Duff, and Peter Dayan · 2005
Earlier work this paper cites.
Cyclic equilibria in markov games
Martin Zinkevich, Amy Greenwald, and Michael Littman · 2005
Earlier work this paper cites.
Model compression
Cristian Buciluǎ, Rich Caruana, and Alexandru Niculescu-Mizil · 2006
Earlier work this paper cites.
Towards a unified theory of state abstraction for mdps
Lihong Li, Thomas J Walsh, and Michael L Littman · 2006
Earlier work this paper cites.
Dealing with non-stationary environments using context detection
Bruno C Da Silva, Eduardo W Basso, Ana LC Bazzan, and Paulo M Engel · 2006
Earlier work this paper cites.
Curiosity-driven development
Frederic Kaplan and Pierre-Yves Oudeyer · 2006
Earlier work this paper cites.
Multi-task reinforcement learning: a hierarchical bayesian approach
Aaron Wilson, Alan Fern, Soumya Ray, and Prasad Tadepalli · 2007
Earlier work this paper cites.
An object-oriented representation for efficient reinforcement learning
Carlos Diuk, Andre Cohen, and Michael L Littman · 2008
Earlier work this paper cites.
Driven by compression progress: A simple principle explains essential aspects of subjective beauty, novelty, surprise, interestingness, attention, curiosity, creativity, art, science, music, jokes
Jürgen Schmidhuber · 2008
Earlier work this paper cites.
Transfer learning for reinforcement learning domains: A survey
Matthew E Taylor and Peter Stone · 2009
Earlier work this paper cites.
Multi-task reinforcement learning in partially observable stochastic environments
Hui Li, Xuejun Liao, and Lawrence Carin · 2009
Earlier work this paper cites.
Where do rewards come from
Satinder Singh, Richard L Lewis, and Andrew G Barto · 2009
Earlier work this paper cites.
R-iac: Robust intrinsically motivated exploration and active learning
Adrien Baranes and Pierre-Yves Oudeyer · 2009
Earlier work this paper cites.
Reinforcement learning in the brain
Yael Niv · 2009
Earlier work this paper cites.
Hierarchically organized behavior and its neural foundations: a reinforcement learning perspective
Matthew M Botvinick, Yael Niv, and Andew G Barto · 2009
Earlier work this paper cites.
Human reinforcement learning subdivides structured action spaces by learning effector-specific values
Samuel J Gershman, Bijan Pesaran, and Nathaniel D Daw · 2009
Earlier work this paper cites.
Planning under uncertainty for robotic tasks with mixed observability
Sylvie CW Ong, Shao Wei Png, David Hsu, and Wee Sun Lee · 2010
Earlier work this paper cites.
Near-optimal regret bounds for reinforcement learning
Thomas Jaksch, Ronald Ortner, and Peter Auer · 2010
Earlier work this paper cites.
Computational models of reinforcement learning: the role of dopamine as a reward signal
RD Samson, MJ Frank, and Jean-Marc Fellous · 2010
Earlier work this paper cites.
States versus rewards: dissociable neural prediction error signals underlying model-based and model-free reinforcement learning
Jan Gläscher, Nathaniel Daw, Peter Dayan, and John P O’Doherty · 2010
Earlier work this paper cites.
Multi-agent learning with policy prediction
Chongjie Zhang and Victor R. Lesser · 2010
Earlier work this paper cites.
Toward an architecture for never-ending language learning
Andrew Carlson, Justin Betteridge, Bryan Kisiel, Burr Settles, Estevam R Hruschka, and Tom M Mitchell · 2010
Earlier work this paper cites.
Minecraft, beyond construction and survival
Sean C Duncan · 2011
Earlier work this paper cites.
On the value of information and other rewards
Yael Niv and Stephanie Chan · 2011
Earlier work this paper cites.
Model-based influences on humans’ choices and striatal prediction errors
Nathaniel D Daw, Samuel J Gershman, Ben Seymour, Peter Dayan, and Raymond J Dolan · 2011
Earlier work this paper cites.
Environmental statistics and the trade-off between model-based and td learning in humans
Dylan A Simon and Nathaniel D Daw · 2011
Earlier work this paper cites.
A multitask representation using reusable local policy templates
Benjamin Saul Rosman and Subramanian Ramamoorthy · 2012
Earlier work this paper cites.
Embodiment theory and education: The foundations of cognition in perception and action
Markus Kiefer and Natalie M Trumpp · 2012
Earlier work this paper cites.
How much of reinforcement learning is working memory, not reinforcement learning? a behavioral, computational, and neurogenetic analysis
Anne GE Collins and Michael J Frank · 2012
Earlier work this paper cites.
Generalization of value in reinforcement learning by humans
G Elliott Wimmer, Nathaniel D Daw, and Daphna Shohamy · 2012
Earlier work this paper cites.
The ubiquity of model-based reinforcement learning
Bradley B Doll, Dylan A Simon, and Nathaniel D Daw · 2012
Earlier work this paper cites.
Mechanisms of hierarchical reinforcement learning in corticostriatal circuits 1: computational analysis
Michael J Frank and David Badre · 2012
Earlier work this paper cites.
Mechanisms of hierarchical reinforcement learning in cortico–striatal circuits 2: Evidence from fmri
David Badre and Michael J Frank · 2012
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton · 2012
Earlier work this paper cites.
Finale Doshi-Velez and George Konidaris · 2013
Earlier work this paper cites.
Intrinsic motivation and reinforcement learning
Andrew G Barto · 2013
Earlier work this paper cites.
Powerplay: Training an increasingly general problem solver by continually searching for the simplest still unsolvable problem
Jürgen Schmidhuber · 2013
Earlier work this paper cites.
Scalable bayesian reinforcement learning for multiagent pomdps
Christopher Amato, Frans A Oliehoek, and Eric Shyu · 2013
Earlier work this paper cites.
The arcade learning environment: An evaluation platform for general agents
Marc G Bellemare, Yavar Naddaf, Joel Veness, and Michael Bowling · 2013
Earlier work this paper cites.
Moderate levels of activation lead to forgetting in the think/no-think paradigm
Greg J Detre, Annamalai Natarajan, Samuel J Gershman, and Kenneth A Norman · 2013
Earlier work this paper cites.
A theory for how sensorimotor skills are learned and retained in noisy and nonstationary neural circuits
Robert Ajemian, Alessandro D’Ausilio, Helene Moorman, and Emilio Bizzi · 2013
Earlier work this paper cites.
Acute stress selectively reduces reward sensitivity
Lisa H Berghorst, Ryan Bogdan, Michael J Frank, and Diego A Pizzagalli · 2013
Earlier work this paper cites.
Divide and conquer: hierarchical reinforcement learning and task decomposition in humans
Carlos Diuk, Anna Schapiro, Natalia Córdova, José Ribas-Fernandes, Yael Niv, and Matthew Botvinick · 2013
Earlier work this paper cites.
Intriguing properties of neural networks
Christian Szegedy, Wojciech Zaremba, Ilya Sutskever, Joan Bruna, Dumitru Erhan, Ian Goodfellow, and Rob Fergus · 2013
Earlier work this paper cites.
Teaching on a budget: Agents advising agents in reinforcement learning
Lisa Torrey and Matthew Taylor · 2013
Earlier work this paper cites.
Online multi-task learning for policy gradient methods
Haitham Bou Ammar, Eric Eaton, Paul Ruvolo, and Matthew Taylor · 2014
Earlier work this paper cites.
Sparse multi-task reinforcement learning
Daniele Calandriello, Alessandro Lazaric, and Marcello Restelli · 2014
Earlier work this paper cites.
Pac-inspired option discovery in lifelong reinforcement learning
Emma Brunskill and Lihong Li · 2014
Earlier work this paper cites.
Curiosity driven reinforcement learning for motion planning on humanoids
Mikhail Frank, Jürgen Leitner, Marijn Stollenga, Alexander Förster, and Jürgen Schmidhuber · 2014
Earlier work this paper cites.
Model-based hierarchical reinforcement learning and human action control
Matthew Botvinick and Ari Weinstein · 2014
Earlier work this paper cites.
Optimal behavioral hierarchy
Alec Solway, Carlos Diuk, Natalia Córdova, Debbie Yee, Andrew G Barto, Yael Niv, and Matthew M Botvinick · 2014
Earlier work this paper cites.
Visualizing and understanding convolutional networks
Matthew D Zeiler and Rob Fergus · 2014
Earlier work this paper cites.
Distilling the knowledge in a neural network
Geoffrey Hinton, Oriol Vinyals, and Jeff Dean · 2015
Earlier work this paper cites.
Andrei A Rusu, Sergio Gomez Colmenarejo, Caglar Gulcehre, Guillaume Desjardins, James Kirkpatrick, Razvan Pascanu, Volodymyr Mnih, Koray Kavukcuoglu, and Raia Hadsell · 2015
Earlier work this paper cites.
Recurrent reinforcement learning: a hybrid approach
Xiujun Li, Lihong Li, Jianfeng Gao, Xiaodong He, Jianshu Chen, Li Deng, and Ji He · 2015
Earlier work this paper cites.
Variational information maximisation for intrinsically motivated reinforcement learning
Shakir Mohamed and Danilo Jimenez Rezende · 2015
Earlier work this paper cites.
Model-based learning protects against forming habits
Claire M Gillan, A Ross Otto, Elizabeth A Phelps, and Nathaniel D Daw · 2015
Earlier work this paper cites.
Reinforcement learning in multidimensional environments relies on attention mechanisms
Yael Niv, Reka Daniel, Andra Geana, Samuel J Gershman, Yuan Chang Leong, Angela Radulescu, and Robert C Wilson · 2015
Earlier work this paper cites.
Discovering latent causes in reinforcement learning
Samuel J Gershman, Kenneth A Norman, and Yael Niv · 2015
Earlier work this paper cites.
Imagenet large scale visual recognition challenge
Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, et al · 2015
Earlier work this paper cites.
Learning to reinforcement learn
Jane X Wang, Zeb Kurth-Nelson, Dhruva Tirumala, Hubert Soyer, Joel Z Leibo, Remi Munos, Charles Blundell, Dharshan Kumaran, and Matt Botvinick · 2016
Earlier work this paper cites.
Rl2: Fast reinforcement learning via slow reinforcement learning
Yan Duan, John Schulman, Xi Chen, Peter L Bartlett, Ilya Sutskever, and Pieter Abbeel · 2016
Earlier work this paper cites.
Learning shared representations in multi-task reinforcement learning
Diana Borsa, Thore Graepel, and John Shawe-Taylor · 2016
Earlier work this paper cites.
Andrei A Rusu, Neil C Rabinowitz, Guillaume Desjardins, Hubert Soyer, James Kirkpatrick, Koray Kavukcuoglu, Razvan Pascanu, and Raia Hadsell · 2016
Earlier work this paper cites.
The Benefit of Multitask Representation Learning
Andreas Maurer, Massimiliano Pontil, and Bernardino Romera-Paredes · 2016
Earlier work this paper cites.
Learning without forgetting
Zhizhong Li and Derek Hoiem · 2016
Earlier work this paper cites.
Representation stability as a regularizer for improved text analytics transfer learning
Matthew Riemer, Elham Khabiri, and Richard Goodwin · 2016
Earlier work this paper cites.
Neural architecture search with reinforcement learning
Barret Zoph and Quoc V Le · 2016
Earlier work this paper cites.
Adaptive skills adaptive partitions (asap)
Daniel J Mankowitz, Timothy A Mann, and Shie Mannor · 2016
Earlier work this paper cites.
Karol Gregor, Danilo Jimenez Rezende, and Daan Wierstra · 2016
Earlier work this paper cites.
Learning to navigate in complex environments
Piotr Mirowski, Razvan Pascanu, Fabio Viola, Hubert Soyer, Andrew J Ballard, Andrea Banino, Misha Denil, Ross Goroshin, Laurent Sifre, Koray Kavukcuoglu, et al · 2016
Earlier work this paper cites.
Reinforcement learning with unsupervised auxiliary tasks
Max Jaderberg, Volodymyr Mnih, Wojciech Marian Czarnecki, Tom Schaul, Joel Z Leibo, David Silver, and Koray Kavukcuoglu · 2016
Earlier work this paper cites.
Loss is its own reward: Self-supervision for reinforcement learning
Evan Shelhamer, Parsa Mahmoudieh, Max Argus, and Trevor Darrell · 2016
Cited alongside, same era.
Unifying count-based exploration and intrinsic motivation
Marc Bellemare, Sriram Srinivasan, Georg Ostrovski, Tom Schaul, David Saxton, and Remi Munos · 2016
Cited alongside, same era.
Vime: Variational information maximizing exploration
Rein Houthooft, Xi Chen, Yan Duan, John Schulman, Filip De Turck, and Pieter Abbeel · 2016
Cited alongside, same era.
Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wojciech Zaremba · 2016
Cited alongside, same era.
Charles Beattie, Joel Z Leibo, Denis Teplyashin, Tom Ward, Marcus Wainwright, Heinrich Küttler, Andrew Lefrancq, Simon Green, Víctor Valdés, Amir Sadik, et al · 2016
Norml: No-reward meta learning
Yuxiang Yang, Ken Caluwaerts, Atil Iscen, Jie Tan, and Chelsea Finn · 2019
Later among the works it cites.
Reward shaping via meta-learning
Haosheng Zou, Tongzheng Ren, Dong Yan, Hang Su, and Jun Zhu · 2019
Later among the works it cites.
What can learned intrinsic rewards capture?
Zeyu Zheng, Junhyuk Oh, Matteo Hessel, Zhongwen Xu, Manuel Kroiss, Hado van Hasselt, David Silver, and Satinder Singh · 2019
Later among the works it cites.
Model-based active exploration
Pranav Shyam, Wojciech Jaśkowski, and Faustino Gomez · 2019
Later among the works it cites.
Hao Liu, Alexander Trott, Richard Socher, and Caiming Xiong · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Vizdoom: A doom-based ai research platform for visual reinforcement learning
Michał Kempka, Marek Wydmuch, Grzegorz Runc, Jakub Toczek, and Wojciech Jaśkowski · 2016
Cited alongside, same era.
Virtual embodiment: A scalable long-term strategy for artificial intelligence research
Douwe Kiela, Luana Bulat, Anita L. Vero, and Stephen Clark · 2016
Cited alongside, same era.
Navigating the affordance landscape: feedback control as a process model of behavior and cognition
Giovanni Pezzulo and Paul Cisek · 2016
Cited alongside, same era.
Overcoming catastrophic forgetting in neural networks
James Kirkpatrick, Razvan Pascanu, Neil Rabinowitz, Joel Veness, Guillaume Desjardins, Andrei A Rusu, Kieran Milan, John Quan, Tiago Ramalho, Agnieszka Grabska-Barwinska, et al · 2017
Cited alongside, same era.
Gradient episodic memory for continuum learning
David Lopez-Paz and Marc’Aurelio Ranzato · 2017
Cited alongside, same era.
Minimax regret bounds for reinforcement learning
Mohammad Gheshlaghi Azar, Ian Osband, and Rémi Munos · 2017
Cited alongside, same era.
Model-agnostic meta-learning for fast adaptation of deep networks
Chelsea Finn, Pieter Abbeel, and Sergey Levine · 2017
Cited alongside, same era.
Adrien Ecoffet, Joost Huizinga, Joel Lehman, Kenneth O Stanley, and Jeff Clune · 2019
Later among the works it cites.
Behaviour suite for reinforcement learning
Ian Osband, Yotam Doron, Matteo Hessel, John Aslanides, Eren Sezener, Andre Saraiva, Katrina McKinney, Tor Lattimore, Csaba Szepezvari, Satinder Singh, et al · 2019
Later among the works it cites.
Optimizing agent behavior over long time scales by transporting value
Chia-Chun Hung, Timothy Lillicrap, Josh Abramson, Yan Wu, Mehdi Mirza, Federico Carnevale, Arun Ahuja, and Greg Wayne · 2019
Later among the works it cites.
Hippocampal contributions to model-based planning and spatial memory
Oliver M Vikbladh, Michael R Meager, John King, Karen Blackmon, Orrin Devinsky, Daphna Shohamy, Neil Burgess, and Nathaniel D Daw · 2019
Later among the works it cites.
Search on the replay buffer: Bridging planning and reinforcement learning
Ben Eysenbach, Russ R Salakhutdinov, and Sergey Levine · 2019
Later among the works it cites.
Sequential replay of nonspatial task states in the human hippocampus
Nicolas W Schuck and Yael Niv · 2019
Later among the works it cites.
Positive reward prediction errors during decision-making strengthen memory encoding
Anthony I Jang, Matthew R Nassar, Daniel G Dillon, and Michael J Frank · 2019
Later among the works it cites.
Subgoal-and goal-related reward prediction errors in medial prefrontal cortex
José JF Ribas-Fernandes, Danesh Shahnazian, Clay B Holroyd, and Matthew M Botvinick · 2019
Later among the works it cites.
Learning task-state representations
Yael Niv · 2019
Later among the works it cites.
The bitter lesson
Richard Sutton · 2019
Later among the works it cites.
Improving generalization in meta reinforcement learning using learned objectives
Louis Kirsch, Sjoerd van Steenkiste, and Juergen Schmidhuber · 2019
Later among the works it cites.
On value functions and the agent-environment boundary
Nan Jiang · 2019
Later among the works it cites.
A survey and critique of multiagent deep reinforcement learning
Pablo Hernandez-Leal, Bilal Kartal, and Matthew E Taylor · 2019
Later among the works it cites.
Stable opponent shaping in differentiable games
Alistair Letcher, Jakob Foerster, David Balduzzi, Tim Rocktäschel, and Shimon Whiteson · 2019
Later among the works it cites.
Learning to teach in cooperative multiagent reinforcement learning
Shayegan Omidshafiei, Dong Ki Kim, Miao Liu, Gerald Tesauro, Matthew Riemer, Christopher Amato, Murray Campbell, and Jonathan P How · 2019
Later among the works it cites.
Interference and generalization in temporal difference learning
Emmanuel Bengio, Joelle Pineau, and Doina Precup · 2020
Closest in time.
Martin Mundt, Yong Won Hong, Iuliia Pliushch, and Visvanathan Ramesh · 2020
Closest in time.
Embracing change: Continual learning in deep neural networks
Raia Hadsell, Dushyant Rao, Andrei A. Rusu, and Razvan Pascanu · 2020
Closest in time.
Offline reinforcement learning: Tutorial, review, and perspectives on open problems
Sergey Levine, Aviral Kumar, George Tucker, and Justin Fu · 2020
Closest in time.
Flambe: Structural complexity and representation learning of low rank mdps
Alekh Agarwal, Sham Kakade, Akshay Krishnamurthy, and Wen Sun · 2020
Closest in time.
Invariant causal prediction for block mdps
Amy Zhang, Clare Lyle, Shagun Sodhani, Angelos Filos, Marta Kwiatkowska, Joelle Pineau, Yarin Gal, and Doina Precup · 2020
Closest in time.
Meta-learning in neural networks: A survey
Timothy Hospedales, Antreas Antoniou, Paul Micaelli, and Amos Storkey · 2020
Closest in time.
Reinforcement learning for non-stationary markov decision processes: The blessing of (more) optimism
Wang Chi Cheung, David Simchi-Levi, and Ruihao Zhu · 2020
Closest in time.
A kernel-based approach to non-stationary reinforcement learning in metric spaces
Omar Darwiche Domingues, Pierre Ménard, Matteo Pirotta, Emilie Kaufmann, and Michal Valko · 2020
Closest in time.
Efficient learning in non-stationary linear markov decision processes, 2020
Ahmed Touati and Pascal Vincent · 2020
Closest in time.
Automatic curriculum learning for deep rl: A short survey
Rémy Portelas, Cédric Colas, Lilian Weng, Katja Hofmann, and Pierre-Yves Oudeyer · 2020
Closest in time.
Dream architecture: a developmental approach to open-ended learning in robotics
Stephane Doncieux, Nicolas Bredeche, Léni Le Goff, Benoît Girard, Alexandre Coninx, Olivier Sigaud, Mehdi Khamassi, Natalia Díaz-Rodríguez, David Filliat, Timothy Hospedales, et al · 2020
Closest in time.
Sharing knowledge in multi-task deep reinforcement learning
Carlo D’Eramo, Davide Tateo, Andrea Bonarini, Marcello Restelli, and Jan Peters · 2020
Closest in time.
Mitchell Wortsman, Vivek Ramanujan, Rosanne Liu, Aniruddha Kembhavi, Mohammad Rastegari, Jason Yosinski, and Ali Farhadi · 2020
Closest in time.
Multi-task reinforcement learning with soft modularization
Ruihan Yang, Huazhe Xu, Yi Wu, and Xiaolong Wang · 2020
Closest in time.
The value equivalence principle for model-based reinforcement learning
Christopher Grimm, André Barreto, Satinder Singh, and David Silver · 2020
Closest in time.
On the role of weight sharing during deep option learning
Matthew Riemer, Ignacio Cases, Clemens Rosenbaum, Miao Liu, and Gerald Tesauro · 2020
Closest in time.
Explore, discover and learn: Unsupervised discovery of state-covering skills
Víctor Campos, Alexander Trott, Caiming Xiong, Richard Socher, Xavier Giro-i Nieto, and Jordi Torres · 2020
Closest in time.
Reset-free lifelong learning with skill-space planning
Kevin Lu, Aditya Grover, Pieter Abbeel, and Igor Mordatch · 2020
Closest in time.
Value preserving state-action abstractions
David Abel, Nate Umbanhowar, Khimya Khetarpal, Dilip Arumugam, Doina Precup, and Michael Littman · 2020
Closest in time.
Reward-free exploration for reinforcement learning
Chi Jin, Akshay Krishnamurthy, Max Simchowitz, and Tiancheng Yu · 2020
Closest in time.
Generalized hindsight for reinforcement learning
Alexander C Li, Lerrel Pinto, and Pieter Abbeel · 2020
Closest in time.
Bootstrap latent-predictive representations for multitask reinforcement learning
Daniel Guo, Bernardo Avila Pires, Bilal Piot, Jean-bastien Grill, Florent Altché, Rémi Munos, and Mohammad Gheshlaghi Azar · 2020
Closest in time.
Generalized hidden parameter mdps transferable model-based rl in a handful of trials
Christian F Perez, Felipe Petroski Such, and Theofanis Karaletsos · 2020
Closest in time.
Jean Harb, Tom Schaul, Doina Precup, and Pierre-Luc Bacon · 2020
Closest in time.
Fast adaptation via policy-dynamics value functions
Roberta Raileanu, Max Goldstein, Arthur Szlam, and Rob Fergus · 2020
Closest in time.
Online fast adaptation and knowledge accumulation: a new approach to continual learning
Massimo Caccia, Pau Rodriguez, Oleksiy Ostapenko, Fabrice Normandin, Min Lin, Lucas Caccia, Issam Laradji, Irina Rish, Alexande Lacoste, David Vazquez, et al · 2020
Closest in time.
Model-based adversarial meta-reinforcement learning
Zichuan Lin, Garrett Thomas, Guangwen Yang, and Tengyu Ma · 2020
Closest in time.
Online meta-critic learning for off-policy actor-critic methods
Wei Zhou, Yiying Li, Yongxin Yang, Huaimin Wang, and Timothy M Hospedales · 2020
Closest in time.
Self-tuning deep reinforcement learning
Tom Zahavy, Zhongwen Xu, Vivek Veeriah, Matteo Hessel, Junhyuk Oh, Hado van Hasselt, David Silver, and Satinder Singh · 2020
Closest in time.
Meta-gradient reinforcement learning with an objective discovered online
Zhongwen Xu, Hado P van Hasselt, Matteo Hessel, Junhyuk Oh, Satinder Singh, and David Silver · 2020
Closest in time.
Planning to explore via self-supervised world models
Ramanan Sekar, Oleh Rybkin, Kostas Daniilidis, Pieter Abbeel, Danijar Hafner, and Deepak Pathak · 2020
Closest in time.
Smirl: Surprise minimizing reinforcement learning in unstable environments
Glen Berseth, Daniel Geng, Coline Manon Devin, Nicholas Rhinehart, Chelsea Finn, Dinesh Jayaraman, and Sergey Levine · 2020
Closest in time.
Causalworld: A robotic manipulation benchmark for causal structure and transfer learning
Ossama Ahmed, Frederik Träuble, Anirudh Goyal, Alexander Neitz, Manuel Wüthrich, Yoshua Bengio, Bernhard Schölkopf, and Stefan Bauer · 2020
Closest in time.
Yonatan Bisk, Ari Holtzman, Jesse Thomason, Jacob Andreas, Yoshua Bengio, Joyce Chai, Mirella Lapata, Angeliki Lazaridou, Jonathan May, Aleksandr Nisnevich, et al · 2020
Closest in time.
Jelly bean world: A testbed for never-ending learning
Emmanouil Antonios Platanios, Abulhair Saparov, and Tom Mitchell · 2020
Closest in time.
Efficiency of learning vs. processing: Towards a normative theory of multitasking
Yotam Sagiv, Sebastian Musslick, Yael Niv, and Jonathan D Cohen · 2020
Closest in time.
The role of executive function in shaping reinforcement learning
Milena Rmus, Samuel McDougle, and Anne Collins · 2020
Closest in time.
Learning structures: Predictive representations, replay, and generalization
Ida Momennejad · 2020
Closest in time.
Distributional reinforcement learning in the brain
Adam S Lowet, Qiao Zheng, Sara Matias, Jan Drugowitsch, and Naoshige Uchida · 2020
Closest in time.
Discovering reinforcement learning algorithms
Junhyuk Oh, Matteo Hessel, Wojciech M Czarnecki, Zhongwen Xu, Hado P van Hasselt, Satinder Singh, and David Silver · 2020
Closest in time.
What is an agent?
Anna Harutyunyan · 2020
Closest in time.
Sara Hooker · 2020
Closest in time.
Language models are few-shot learners
Tom B Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Closest in time.
Scaling laws for neural language models
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei · 2020
Closest in time.
A neural scaling law from the dimension of the data manifold
Utkarsh Sharma and Jared Kaplan · 2020
Closest in time.
Scaling laws for autoregressive generative modeling
Tom Henighan, Jared Kaplan, Mor Katz, Mark Chen, Christopher Hesse, Jacob Jackson, Heewoo Jun, Tom B Brown, Prafulla Dhariwal, Scott Gray, et al · 2020
Closest in time.
Learning hierarchical teaching policies for cooperative agents
Dong Ki Kim, Miao Liu, Shayegan Omidshafiei, Sebastian Lopez-Cot, Matthew Riemer, Golnaz Habibi, Gerald Tesauro, Sami Mourad, Murray Campbell, and Jonathan P How · 2020
Closest in time.
Deep reinforcement learning amidst continual structured non-stationarity
Annie Xie, James Harrison, and Chelsea Finn · 2021
Closest in time.
Near-optimal model-free reinforcement learning in non-stationary episodic mdps
Weichao Mao, Kaiqing Zhang, Ruihao Zhu, David Simchi-Levi, and Tamer Basar · 2021
Closest in time.
Sparsely changing latent states for prediction and planning in partially observable domains
Christian Gumbsch, Martin V Butz, and Georg Martius · 2021
Closest in time.
Representation learning for online and offline rl in low-rank mdps
Masatoshi Uehara, Xuezhou Zhang, and Wen Sun · 2021
Closest in time.
Bellman eluder dimension: New rich classes of rl problems, and sample-efficient algorithms
Chi Jin, Qinghua Liu, and Sobhan Miryoosefi · 2021
Closest in time.
Adarl: What, where, and how to adapt in transfer reinforcement learning
Biwei Huang, Fan Feng, Chaochao Lu, Sara Magliacane, and Kun Zhang · 2021
Closest in time.
Provable rich observation reinforcement learning with combinatorial latent states
Dipendra Misra, Qinghua Liu, Chi Jin, and John Langford · 2021
Closest in time.
Learning domain invariant representations in goal-conditioned block mdps
Beining Han, Chongyi Zheng, Harris Chan, Keiran Paster, Michael Zhang, and Jimmy Ba · 2021
Closest in time.
Ltl2action: Generalizing ltl instructions for multi-task rl
Pashootan Vaezipoor, Andrew C Li, Rodrigo A Toro Icarte, and Sheila A Mcilraith · 2021
Closest in time.
How rl agents behave when their actions are modified
Eric D Langlois and Tom Everitt · 2021
Closest in time.
Know your action set: Learning action relations for reinforcement learning
Ayush Jain, Norio Kosaka, Kyung-Min Kim, and Joseph J Lim · 2021
Closest in time.
Dealing with non-stationarity in marl via trust-region decomposition
Wenhao Li, Xiangfeng Wang, Bo Jin, Junjie Sheng, and Hongyuan Zha · 2021
Closest in time.
Automatic data augmentation for generalization in reinforcement learning
Roberta Raileanu, Maxwell Goldstein, Denis Yarats, Ilya Kostrikov, and Rob Fergus · 2021
Closest in time.
Variational automatic curriculum learning for sparse-reward cooperative multi-agent problems
Jiayu Chen, Yuanxin Zhang, Yuanfan Xu, Huimin Ma, Huazhong Yang, Jiaming Song, Yu Wang, and Yi Wu · 2021
Closest in time.
Improving generalization in meta-rl with imaginary tasks from latent dynamics mixture
Suyoung Lee and Sae-Young Chung · 2021
Closest in time.
Autonomous reinforcement learning via subgoal curricula
Archit Sharma, Abhishek Gupta, Sergey Levine, Karol Hausman, and Chelsea Finn · 2021
Closest in time.
Boosted curriculum reinforcement learning
Pascal Klink, Carlo D’Eramo, Jan Peters, and Joni Pajarinen · 2021
Closest in time.
Meta-adaptive nonlinear control: Theory and algorithms
Guanya Shi, Kamyar Azizzadenesheli, Michael O’Connell, Soon-Jo Chung, and Yisong Yue · 2021
Closest in time.
Transient non-stationarity and generalisation in deep reinforcement learning
Maximilian Igl, Gregory Farquhar, Jelena Luketina, JW Böhmer, and Shimon Whiteson · 2021
Closest in time.
Model-augmented prioritized experience replay
Youngmin Oh, Jinwoo Shin, Eunho Yang, and Sung Ju Hwang · 2021
Closest in time.
Posterior meta-replay for continual learning
Christian Henning, Maria Cervera, Francesco D’Angelo, Johannes Von Oswald, Regina Traber, Benjamin Ehret, Seijin Kobayashi, Benjamin F Grewe, and João Sacramento · 2021
Closest in time.
Generalized proximal policy optimization with sample reuse
James Queeney, Yannis Paschalidis, and Christos G Cassandras · 2021
Closest in time.
Universal off-policy evaluation
Yash Chandak, Scott Niekum, Bruno da Silva, Erik Learned-Miller, Emma Brunskill, and Philip S Thomas · 2021
Closest in time.
Towards mental time travel: a hierarchical memory for reinforcement learning agents
Andrew Lampinen, Stephanie Chan, Andrea Banino, and Felix Hill · 2021
Closest in time.
Policy gradients incorporating the future
David Venuto, Elaine Lau, Doina Precup, and Ofir Nachum · 2021
Closest in time.
Sharing less is more: Lifelong learning in deep networks with selective layer transfer
Seungwon Lee, Sima Behpour, and Eric Eaton · 2021
Closest in time.
Toward robust long range policy transfer
Wei-Cheng Tseng, Jin-Siang Lin, Yao-Min Feng, and Min Sun · 2021
Closest in time.
Learning markov state abstractions for deep reinforcement learning
Cameron Allen, Neev Parikh, Omer Gottesman, and George Konidaris · 2021
Closest in time.
Monte carlo tree search with iteratively refining state abstractions
Samuel Sokota, Caleb Y Ho, Zaheen Ahmad, and J Zico Kolter · 2021
Closest in time.
Control-aware representations for model-based reinforcement learning
Brandon Cui, Yinlam Chow, and Mohammad Ghavamzadeh · 2021
Closest in time.
Flexible option learning
Martin Klissarov and Doina Precup · 2021
Closest in time.
Reset-free lifelong learning with skill-space planning
Kevin Lu, Aditya Grover, Pieter Abbeel, and Igor Mordatch · 2021
Closest in time.
Learning one representation to optimize all rewards
Ahmed Touati and Yann Ollivier · 2021
Closest in time.
Offline meta reinforcement learning–identifiability challenges and effective data collection strategies
Ron Dorfman, Idan Shenfeld, and Aviv Tamar · 2021
Closest in time.
Risk-averse bayes-adaptive reinforcement learning
Marc Rigter, Bruno Lacerda, and Nick Hawes · 2021
Closest in time.
Multi-task reinforcement learning with context-based representations
Shagun Sodhani, Amy Zhang, and Joelle Pineau · 2021
Closest in time.
Towards effective context for meta-reinforcement learning: an approach based on contrastive learning
Haotian Fu, Hongyao Tang, Jianye Hao, Chen Chen, Xidong Feng, Dong Li, and Wulong Liu · 2021
Closest in time.
Accelerating online reinforcement learning via model-based meta-learning
John D Co-Reyes, Sarah Feng, Glen Berseth, Jie Qui, and Sergey Levine · 2021
Closest in time.
Comps: Continual meta policy search
Glen Berseth, Zhiwei Zhang, Grace Zhang, Chelsea Finn, and Sergey Levine · 2021
Closest in time.
Sebastian Flennerhag, Yannick Schroecker, Tom Zahavy, Hado van Hasselt, David Silver, and Satinder Singh · 2021
Closest in time.
Continuous coordination as a realistic scenario for lifelong learning
Hadi Nekoei, Akilesh Badrinaaraayanan, Aaron Courville, and Sarath Chandar · 2021
Closest in time.
Cora: Benchmarks, baselines, and metrics as a platform for continual reinforcement learning agents
Sam Powers, Eliot Xing, Eric Kolve, Roozbeh Mottaghi, and Abhinav Gupta · 2021
Closest in time.
Novelgridworlds: A benchmark environment for detecting and adapting to novelties in open worlds
Shivam Goel, Gyan Tatiya, Matthias Scheutz, and Jivko Sinapov · 2021
Closest in time.
Continual backprop: Stochastic gradient descent with persistent randomness
Shibhansh Dohare, A Rupam Mahmood, and Richard S Sutton · 2021
Closest in time.
Continual learning in environments with polynomial mixing times
Matthew Riemer, Sharath Chandra Raparthy, Ignacio Cases, Gopeshh Subbaraj, Maximilian Puelma Touzel, and Irina Rish · 2022
Closest in time.
Influencing long-term behavior in multiagent reinforcement learning
Dong Ki Kim, Matthew Riemer, Miao Liu, Jakob N Foerster, Michael Everett, Chuangchuang Sun, Gerald Tesauro, and Jonathan P How · 2022
Closest in time.
Sample-efficient reinforcement learning for pomdps with linear function approximations
Qi Cai, Zhuoran Yang, and Zhaoran Wang · 2022
Closest in time.
Block contextual mdps for continual learning
Shagun Sodhani, Franziska Meier, Joelle Pineau, and Amy Zhang · 2022
Closest in time.
Anymorph: Learning transferable polices by inferring agent morphology
Brandon Trabucco, Mariano Phielipp, and Glen Berseth · 2022
Closest in time.
Memory-efficient reinforcement learning with knowledge consolidation
Qingfeng Lan, Yangchen Pan, Jun Luo, and A Rupam Mahmood · 2022
Closest in time.
Catastrophic interference in reinforcement learning: A solution based on context division and knowledge distillation
Tiantian Zhang, Xueqian Wang, Bin Liang, and Bo Yuan · 2022
Closest in time.
Lifelong hyper-policy optimization with multiple importance sampling regularization
Pierre Liotet, Francesco Vidaich, Alberto Maria Metelli, and Marcello Restelli · 2022
Closest in time.
Model-free generative replay for lifelong reinforcement learning: Application to starcraft-2
Zachary Daniels, Aswin Raghavan, Jesse Hostetler, Abrar Rahman, Indranil Sur, Michael Piacentino, and Ajay Divakaran · 2022
Closest in time.
Jorge A Mendez and Eric Eaton · 2022
Closest in time.
Modular lifelong reinforcement learning via neural composition
Jorge A Mendez, Harm van Seijen, and Eric Eaton · 2022
Closest in time.
Building a subspace of policies for scalable continual learning
Jean-Baptiste Gaya, Thang Doan, Lucas Caccia, Laure Soulier, Ludovic Denoyer, and Roberta Raileanu · 2022
Closest in time.
Bisimulation makes analogies in goal-conditioned reinforcement learning
Philippe Hansen-Estruch, Amy Zhang, Ashvin Nair, Patrick Yin, and Sergey Levine · 2022
Closest in time.
Structural similarity for improved transfer in reinforcement learning
C Chace Ashcraft, Benjamin Stoler, Chigozie Ewulum, and Susama Agarwala · 2022
Closest in time.
Discovering state and action abstractions for generalized task and motion planning
Aidan Curtis, Tom Silver, Joshua B Tenenbaum, Tomás Lozano-Pérez, and Leslie Kaelbling · 2022
Closest in time.
Context-specific representation abstraction for deep option learning
Marwa Abdulhai, Dong Ki Kim, Matthew Riemer, Miao Liu, Gerald Tesauro, and Jonathan P How · 2022
Closest in time.
Dynamic dialogue policy transformer for continual reinforcement learning
Christian Geishauser, Carel van Niekerk, Nurul Lubis, Michael Heck, Hsien-Chin Lin, Shutong Feng, and Milica Gašić · 2022
Closest in time.
Goal-directed planning via hindsight experience replay
Lorenzo Moro, Amarildo Likmeta, Enrico Prati, Marcello Restelli, et al · 2022
Closest in time.
Same state, different task: Continual reinforcement learning without interference
Samuel Kessler, Jack Parker-Holder, Philip Ball, Stefan Zohren, and Stephen J Roberts · 2022
Closest in time.
Adapt to environment sudden changes by learning a context sensitive policy
Fan-Ming Luo, Shengyi Jiang, Yang Yu, Zongzhang Zhang, and Yi-Feng Zhang · 2022
Closest in time.
Introducing symmetries to black box meta reinforcement learning
Louis Kirsch, Sebastian Flennerhag, Hado van Hasselt, Abram Friesen, Junhyuk Oh, and Yutian Chen · 2022
Closest in time.
Hindsight foresight relabeling for meta-reinforcement learning
Michael Wan, Jian Peng, and Tanmay Gangwani · 2022
Closest in time.
Transformers are meta-reinforcement learners
Luckeciano C Melo · 2022
Closest in time.
Skill-based meta-reinforcement learning
Taewook Nam, Shao-Hua Sun, Karl Pertsch, Sung Ju Hwang, and Joseph J Lim · 2022
Closest in time.
Reactive exploration to cope with non-stationarity in lifelong reinforcement learning
Christian Steinparz, Thomas Schmied, Fabian Paischer, Marius-Constantin Dinu, Vihang Patil, Angela Bitto-Nemling, Hamid Eghbal-zadeh, and Sepp Hochreiter · 2022
Closest in time.
L2explorer: A lifelong reinforcement learning assessment environment
Erik C Johnson, Eric Q Nguyen, Blake Schreurs, Chigozie S Ewulum, Chace Ashcraft, Neil M Fendley, Megan M Baker, Alexander New, and Gautam K Vallabha · 2022
Closest in time.
Blenderbot 3: a deployed conversational agent that continually learns to responsibly engage
Kurt Shuster, Jing Xu, Mojtaba Komeili, Da Ju, Eric Michael Smith, Stephen Roller, Megan Ung, Moya Chen, Kushal Arora, Joshua Lane, et al · 2022
Closest in time.