Fetching the paper…
Reading the bibliography…
Some real-world domains are best characterized as a single task, but for others this perspective is limiting.
Learning from delayed rewards
C. J. C. H. Watkins · 1989
Earlier work this paper cites.
Q-learning
C. J. Watkins and P. Dayan · 1992
Earlier work this paper cites.
Learning to achieve goals
L. P. Kaelbling · 1993
Earlier work this paper cites.
Hierarchical learning in stochastic domains
R. Ashar · 1994
Earlier work this paper cites.
Incremental multi-step q-learning
J. Peng and R. J. Williams · 1994
Earlier work this paper cites.
Continual learning in reinforcement environments
M. B. Ring · 1994
Earlier work this paper cites.
Long short-term memory
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner · 1998
Earlier work this paper cites.
Reinforcement learning: An introduction , volume 1
R. S. Sutton and A. G. Barto · 1998
Earlier work this paper cites.
Intra-option learning about temporally abstract actions
R. S. Sutton, D. Precup, and S. P. Singh · 1998
Earlier work this paper cites.
Hierarchical reinforcement learning with the maxq value function decomposition
T. G. Dietterich · 2000
Earlier work this paper cites.
Off-policy temporal-difference learning with function approximation
D. Precup, R. S. Sutton, and S. Dasgupta · 2001
Earlier work this paper cites.
Curriculum learning
Y. Bengio, J. Louradour, R. Collobert, and J. Weston · 2009
Earlier work this paper cites.
A convergent o ( n ) o(n) temporal-difference algorithm for off-policy learning with linear function approximation
R. S. Sutton, H. R. Maei, and C. Szepesvári · 2009
Earlier work this paper cites.
Toward off-policy learning control with function approximation
H. R. Maei, C. Szepesvári, S. Bhatnagar, and R. S. Sutton · 2010
Earlier work this paper cites.
Metalearning
T. Schaul and J. Schmidhuber · 2010
Earlier work this paper cites.
Horde: A scalable real-time architecture for learning knowledge from unsupervised sensorimotor interaction
R. S. Sutton, J. Modayil, M. Delp, T. Degris, P. M. Pilarski, A. White, and D. Precup · 2011
Earlier work this paper cites.
Curriculum learning for motor skills
A. Karpathy and M. Van De Panne · 2012
Cited alongside, same era.
Pac-inspired option discovery in lifelong reinforcement learning
E. Brunskill and L. Li · 2014
Cited alongside, same era.
Scaling up robust mdps using function approximation
A. Tamar, S. Mannor, and H. Xu · 2014
Cited alongside, same era.
Safe policy search for lifelong reinforcement learning with sublinear regret
H. B. Ammar, R. Tutunov, and E. Eaton · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski, et al · 2015
Cited alongside, same era.
Universal value function approximators
T. Schaul, D. Horgan, K. Gregor, and D. Silver · 2015
Cited alongside, same era.
Multi-step reinforcement learning: A unifying algorithm
K. De Asis, J. F. Hernandez-Garcia, G. Z. Holland, and R. S. Sutton · 2017
Later among the works it cites.
Model-agnostic meta-learning for fast adaptation of deep networks
C. Finn, P. Abbeel, and S. Levine · 2017
Later among the works it cites.
Reverse curriculum generation for reinforcement learning
C. Florensa, D. Held, M. Wulfmeier, and P. Abbeel · 2017
Later among the works it cites.
Automated curriculum learning for neural networks
A. Graves, M. G. Bellemare, J. Menick, R. Munos, and K. Kavukcuoglu · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Optimizing the cvar via sampling
A. Tamar, Y. Glassner, and S. Mannor · 2015
Cited alongside, same era.
C. Beattie, J. Z. Leibo, D. Teplyashin, T. Ward, M. Wainwright, H. Küttler, A. Lefrancq, S. Green, V. Valdés, A. Sadik, et al · 2016
Cited alongside, same era.
beta-vae: Learning basic visual concepts with a constrained variational framework
I. Higgins, L. Matthey, A. Pal, C. Burgess, X. Glorot, M. Botvinick, S. Mohamed, and A. Lerchner · 2016
Cited alongside, same era.
Reinforcement learning with unsupervised auxiliary tasks
M. Jaderberg, V. Mnih, W. M. Czarnecki, T. Schaul, J. Z. Leibo, D. Silver, and K. Kavukcuoglu · 2016
Cited alongside, same era.
Hierarchical deep reinforcement learning: Integrating temporal abstraction and intrinsic motivation
T. D. Kulkarni, K. Narasimhan, A. Saeedi, and J. Tenenbaum · 2016
Cited alongside, same era.
Adaptive Skills, Adaptive Partitions (ASAP)
D. J. Mankowitz, T. A. Mann, and S. Mannor · 2016
Cited alongside, same era.
N. Heess, D. TB, S. Sriram, J. Lemmon, J. Merel, G. Wayne, Y. Tassa, T. Erez, Z. Wang, S. M. A. Eslami, M. A. Riedmiller, and D. Silver · 2017
Later among the works it cites.
Grounded language learning in a simulated 3d world
K. M. Hermann, F. Hill, S. Green, F. Wang, R. Faulkner, H. Soyer, D. Szepesvari, W. Czarnecki, M. Jaderberg, D. Teplyashin, et al · 2017
Later among the works it cites.
Rainbow: Combining improvements in deep reinforcement learning
M. Hessel, J. Modayil, H. Van Hasselt, T. Schaul, G. Ostrovski, W. Dabney, D. Horgan, B. Piot, M. Azar, and D. Silver · 2017
Later among the works it cites.
The uncertainty bellman equation and exploration
B. O’Donoghue, I. Osband, R. Munos, and V. Mnih · 2017
Later among the works it cites.
Zero-shot task generalization with multi-task deep reinforcement learning
J. Oh, S. Singh, H. Lee, and P. Kohli · 2017
Later among the works it cites.
Intrinsic motivation and automatic curricula via asymmetric self-play
S. Sukhbaatar, I. Kostrikov, A. Szlam, and R. Fergus · 2017
Later among the works it cites.
A deep hierarchical approach to lifelong learning in minecraft
C. Tessler, S. Givony, T. Zahavy, D. J. Mankowitz, and S. Mannor · 2017
Later among the works it cites.
Feudal networks for hierarchical reinforcement learning
A. S. Vezhnevets, S. Osindero, T. Schaul, N. Heess, M. Jaderberg, D. Silver, and K. Kavukcuoglu · 2017
Later among the works it cites.
Scalable distributed deep-rl with importance weighted actor-learner architectures
L. Espeholt, H. Soyer, R. Munos, K. Simonyan, V. Mnih, T. Ward, Y. Doron, V. Firoiu, T. Harley, I. Dunning, S. Legg, and K. Kavukcuoglu · 2018
Closest in time.
Learning robust options
D. J. Mankowitz, T. A. Mann, P.-L. Bacon, D. Precup, and S. Mannor · 2018
Closest in time.
Continual lifelong learning with neural networks: A review
G. I. Parisi, R. Kemker, J. L. Part, C. Kanan, and S. Wermter · 2018
Closest in time.
Learning by playing-solving sparse reward tasks from scratch
M. Riedmiller, R. Hafner, T. Lampe, M. Neunert, J. Degrave, T. Van de Wiele, V. Mnih, N. Heess, and J. T. Springenberg · 2018
Closest in time.