Fetching the paper…
Reading the bibliography…
In this paper we consider the problem of how a reinforcement learning agent tasked with solving a set of related Markov decision processes can use knowledge acquired early in its lifetime to improve its ability to more rapidly solve novel, but related, tasks.
A technique for high-performance data compression
T. A. Welch · 1984
Earlier work this paper cites.
Complexity and cooperation in q-learning
Steven D. Whitehead · 1991
Earlier work this paper cites.
Q-learning
Christopher J. C. H. Watkins and Peter Dayan · 1992
Earlier work this paper cites.
Macro actions in reinforcement learning: An empirical analysis
A. McGovern and R. Sutton · 1998
Earlier work this paper cites.
Learning macro-actions in reinforcement learning
Jette Randlov · 1998
Earlier work this paper cites.
Between mdps and semi-mdps: A framework for temporal abstraction in reinforcement learning
Richard S. Sutton, Doina Precup, and Satinder P. Singh · 1999
Earlier work this paper cites.
Automatic discovery of subgoals in reinforcement learning using diverse density
Amy McGovern and Andrew G. Barto · 2001
Cited alongside, same era.
Q-cut - dynamic discovery of sub-goals in reinforcement learning
Ishai Menache, Shie Mannor, and Nahum Shimkin · 2002
Cited alongside, same era.
Proto-value functions: Developmental reinforcement learning
Sridhar Mahadevan · 2005
Cited alongside, same era.
Conjugate markov decision processes
Philip S. Thomas and Andrew G. Barto · 2011
Cited alongside, same era.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A. Rusu, Joel Veness, Marc G. Bellemare, Alex Graves, Martin Riedmiller, Andreas K. Fidjeland, Georg Ostrovski, Stig Petersen, Charles Beattie, Amir Sadik, Ioannis Antonoglou, Helen King, Dharshan Kumaran, Daan Wierstra, Shane Legg, and Demis Hassabis · 2015
Cited alongside, same era.
Hierarchical reinforcement learning using spatio-temporal abstractions and deep neural networks
Ramnandan Krishnamurthy, Aravind S. Lakshminarayanan, Peeyush Kumar, and Balaraman Ravindran · 2016
Later among the works it cites.
Resource management with deep reinforcement learning
Hongzi Mao, Mohammad Alizadeh, Ishai Menache, and Srikanth Kandula · 2016
Later among the works it cites.
The option-critic architecture
Pierre-Luc Bacon, Jean Harb, and Doina Precup · 2017
Closest in time.
A Laplacian Framework for Option Discovery in Reinforcement Learning
Michael Bowling Marlos C. Machado, Marc G. Bellemare · 2017
Closest in time.
Introduction to Reinforcement Learning
Richard S. Sutton and Andrew G. Barto · 2018
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wojciech Zaremba · 2016
Cited alongside, same era.