Fetching the paper…
Reading the bibliography…
In this paper we consider the problem of how a reinforcement learning agent that is tasked with solving a sequence of reinforcement learning problems (a sequence of Markov decision processes) can use knowledge acquired early in its lifetime to improve its ability to solve new problems.
Complexity and cooperation in q-learning
Steven D. Whitehead · 1991
Earlier work this paper cites.
Q-learning
Christopher J. C. H. Watkins and Peter Dayan · 1992
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J. Williams · 1992
Earlier work this paper cites.
Local and global optimization algorithms for generalized learning automata
V. V. Phansalkar and M. A. L. Thathachar · 1995
Earlier work this paper cites.
Introduction to Reinforcement Learning
Richard S. Sutton and Andrew G. Barto · 1998
Earlier work this paper cites.
Automatic discovery of subgoals in reinforcement learning using diverse density
Amy Mcgovern and Andrew G. Barto · 2001
Earlier work this paper cites.
Subgoal discovery for hierarchical reinforcement learning using learned policies
Sandeep Goel and Manfred Huber · 2003
Earlier work this paper cites.
Proto-value functions: Developmental reinforcement learning
Sridhar Mahadevan · 2005
Earlier work this paper cites.
Probabilistic policy reuse in a reinforcement learning agent
Fernando Fernandez and Manuela Veloso · 2006
Earlier work this paper cites.
Probably approximately correct (pac) exploration in reinforcement learning
Alexander L. Strehl · 2008
Earlier work this paper cites.
On upper-confidence bound policies for switching bandit problems
Aurélien Garivier and Eric Moulines · 2011
Earlier work this paper cites.
Conjugate markov decision processes
Philip S. Thomas and Andrew G. Barto · 2011
Cited alongside, same era.
Optimal rewards in multiagent teams
Bingyao Liu, Satinder P. Singh, Richard L. Lewis, and Shiyin Qin · 2012
Cited alongside, same era.
Reinforcement learning controller design for affine nonlinear discrete-time systems using online approximators
Q. Yang and S. Jagannathan · 2012
Cited alongside, same era.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A. Rusu, Joel Veness, Marc G. Bellemare, Alex Graves, Martin Riedmiller, Andreas K. Fidjeland, Georg Ostrovski, Stig Petersen, Charles Beattie, Amir Sadik, Ioannis Antonoglou, Helen King, Dharshan Kumaran, Daan Wierstra, Shane Legg, and Demis Hassabis · 2015
Cited alongside, same era.
Learning deep control policies for autonomous aerial vehicles with mpc-guided policy search
Tianhao Zhang, Gregory Kahn, Sergey Levine, and Pieter Abbeel · 2015
Cited alongside, same era.
Minimax regret bounds for reinforcement learning
Mohammad Gheshlaghi Azar, Ian Osband, and Rémi Munos · 2017
Later among the works it cites.
Model-agnostic meta-learning for fast adaptation of deep networks
Chelsea Finn, Pieter Abbeel, and Sergey Levine · 2017
Later among the works it cites.
Multi-advisor reinforcement learning
Romain Laroche, Mehdi Fatemi, Harm van Seijen, and Joshua Romoff · 2017
Later among the works it cites.
A laplacian framework for option discovery in reinforcement learning
Marlos C. Machado, Marc G. Bellemare, and Michael H. Bowling · 2017
Later among the works it cites.
Count-based exploration in feature space for reinforcement learning
Jarryd Martin, Suraj Narayanan Sasikumar, Tom Everitt, and Marcus Hutter · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Marcin Andrychowicz, Misha Denil, Sergio Gomez Colmenarejo, Matthew W. Hoffman, David Pfau, Tom Schaul, and Nando de Freitas · 2016
Cited alongside, same era.
Openai gym, 2016
Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wojciech Zaremba · 2016
Cited alongside, same era.
Curiosity-driven exploration in deep reinforcement learning via bayesian neural networks
Rein Houthooft, Xi Chen, Yan Duan, John Schulman, Filip De Turck, and Pieter Abbeel · 2016
Cited alongside, same era.
High-dimensional continuous control using generalized advantage estimation
John Schulman, Philipp Moritz, Sergey Levine, Michael Jordan, and Pieter Abbeel · 2016
Cited alongside, same era.
#exploration: A study of count-based exploration for deep reinforcement learning
Haoran Tang, Rein Houthooft, Davis Foote, Adam Stooke, Xi Chen, Yan Duan, John Schulman, Filip De Turck, and Pieter Abbeel · 2016
Cited alongside, same era.
Deepak Pathak, Pulkit Agrawal, Alexei A. Efros, and Trevor Darrell · 2017
Later among the works it cites.
Routing networks: Adaptive selection of non-linear functions for multi-task learning
Clemens Rosenbaum, Tim Klinger, and Matthew Riemer · 2017
Later among the works it cites.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Later among the works it cites.
Hybrid reward architecture for reinforcement learning
Harm van Seijen, Mehdi Fatemi, Joshua Romoff, Romain Laroche, Tavian Barnes, and Jeffrey Tsang · 2017
Later among the works it cites.
Reinforcement learning for UAV attitude control
William Koch, Renato Mancuso, Richard West, and Azer Bestavros · 2018
Later among the works it cites.