Fetching the paper…
Reading the bibliography…
The ability of a reinforcement learning (RL) agent to learn about many reward functions at the same time has many potential benefits, such as the decomposition of complex tasks into simpler ones, the exchange of information between tasks, and the reuse of skills.
Learning from Delayed Rewards
C. Watkins · 1989
Earlier work this paper cites.
Learning to achieve goals
L. P. Kaelbling · 1993
Earlier work this paper cites.
Hierarchical learning in stochastic domains
R. Ashar · 1994
Earlier work this paper cites.
Markov Decision Processes—Discrete Stochastic Dynamic Programming
M. L. Puterman · 1994
Earlier work this paper cites.
A general method for incremental self-improvement and multi-agent learning in unrestricted environments
J. Schmidhuber · 1996
Earlier work this paper cites.
Reinforcement Learning: An Introduction
R. S. Sutton and A. G. Barto · 1998
Earlier work this paper cites.
Transfer learning for reinforcement learning domains: A survey
M. E. Taylor and P. Stone · 2009
Earlier work this paper cites.
Algorithms for Reinforcement Learning
C. Szepesvári · 2010
Earlier work this paper cites.
Horde: A scalable real-time architecture for learning knowledge from unsupervised sensorimotor interaction
R. S. Sutton, J. Modayil, M. Delp, T. Degris, P. M. Pilarski, A. White, and D. Precup · 2011
Earlier work this paper cites.
Transfer in Reinforcement Learning: A Framework and a Survey , pages 143–173
A. Lazaric · 2012
Earlier work this paper cites.
Human-level control through deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski, S. Petersen, C. Beattie, A. Sadik, I. Antonoglou, H. King, D. Kumaran, D. Wierstra, S. Legg, and D. Hassabis · 2015
Earlier work this paper cites.
Universal Value Function Approximators
T. Schaul, D. Horgan, K. Gregor, and D. Silver · 2015
Cited alongside, same era.
Modular multitask reinforcement learning with policy sketches
J. Andreas, D. Klein, and S. Levine · 2016
Cited alongside, same era.
C. Beattie, J. Z. Leibo, D. Teplyashin, T. Ward, M. Wainwright, H. Küttler, A. Lefrancq, S. Green, V. Valdés, A. Sadik, et al · 2016
Cited alongside, same era.
Learning and transfer of modulated locomotor controllers
N. Heess, G. Wayne, Y. Tassa, T. Lillicrap, M. Riedmiller, and D. Silver · 2016
Cited alongside, same era.
Overcoming catastrophic forgetting in neural networks
J. Kirkpatrick, R. Pascanu, N. C. Rabinowitz, J. Veness, G. Desjardins, A. A. Rusu, K. Milan, J. Quan, T. Ramalho, A. Grabska-Barwinska, D. Hassabis, C. Clopath, D. Kumaran, and R. Hadsell · 2016
Learning modular neural network policies for multi-task and multi-robot transfer
C. Devin, A. Gupta, T. Darrell, P. Abbeel, and S. Levine · 2017
Later among the works it cites.
Model-agnostic meta-learning for fast adaptation of deep networks
C. Finn, P. Abbeel, and S. Levine · 2017
Later among the works it cites.
Grounded language learning in a simulated 3d world
K. M. Hermann, F. Hill, S. Green, F. Wang, R. Faulkner, H. Soyer, D. Szepesvari, W. Czarnecki, M. Jaderberg, D. Teplyashin, et al · 2017
Later among the works it cites.
Zero-shot task generalization with multi-task deep reinforcement learning
J. Oh, S. P. Singh, H. Lee, and P. Kohli · 2017
Later among the works it cites.
Distral: Robust multitask reinforcement learning
Y. W. Teh, V. Bapst, W. M. Czarnecki, J. Quan, J. Kirkpatrick, R. Hadsell, N. Heess, and R. Pascanu · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Deep successor reinforcement learning
T. D. Kulkarni, A. Saeedi, S. Gautam, and S. J. Gershman · 2016
Cited alongside, same era.
A. A. Rusu, N. C. Rabinowitz, G. Desjardins, H. Soyer, J. Kirkpatrick, K. Kavukcuoglu, R. Pascanu, and R. Hadsell · 2016
Cited alongside, same era.
Deep reinforcement learning with successor features for navigation across similar environments
J. Zhang, J. T. Springenberg, J. Boedecker, and W. Burgard · 2016
Cited alongside, same era.
Hindsight experience replay
M. Andrychowicz, F. Wolski, A. Ray, J. Schneider, R. Fong, P. Welinder, B. McGrew, J. Tobin, O. P. Abbeel, and W. Zaremba · 2017
Cited alongside, same era.
Successor features for transfer in reinforcement learning
A. Barreto, W. Dabney, R. Munos, J. Hunt, T. Schaul, H. van Hasselt, and D. Silver · 2017
Cited alongside, same era.
Later among the works it cites.
FeUdal networks for hierarchical reinforcement learning
A. S. Vezhnevets, S. Osindero, T. Schaul, N. Heess, M. Jaderberg, D. Silver, and K. Kavukcuoglu · 2017
Later among the works it cites.
Transfer in deep reinforcement learning using successor features and generalised policy improvement
A. Barreto, D. Borsa, J. Quan, T. Schaul, D. Silver, M. Hessel, D. Mankowitz, A. Zidek, and R. Munos · 2018
Closest in time.
Impala: Scalable distributed deep-rl with importance weighted actor-learner architectures
L. Espeholt, H. Soyer, R. Munos, K. Simonyan, V. Mnih, T. Ward, Y. Doron, V. Firoiu, T. Harley, I. Dunning, et al · 2018
Closest in time.
Universal successor representations for transfer reinforcement learning
C. Ma, J. Wen, and Y. Bengio · 2018
Closest in time.
Unicorn: Continual learning with a universal, off-policy agent
D. J. Mankowitz, A. Žídek, A. Barreto, D. Horgan, M. Hessel, J. Quan, J. Oh, H. van Hasselt, D. Silver, and T. Schaul · 2018
Closest in time.