Fetching the paper…
Reading the bibliography…
Data efficiency and robustness to task-irrelevant perturbations are long-standing challenges for deep reinforcement learning algorithms.
Mega-reward: Achieving human-level play without extrinsic rewards
Y. Song, J. Wang, T. Lukasiewicz, Z. Xu, S. Zhang, and M. Xu · 1905
Earlier work this paper cites.
Learning powerful policies by using consistent dynamics model
S. Sodhani, A. Goyal, T. Deleu, Y. Bengio, S. Levine, and J. Tang · 1906
Earlier work this paper cites.
A cluster separation measure
D. L. Davies and D. W. Bouldin · 1979
Earlier work this paper cites.
An on-line algorithm for dynamic reinforcement learning and planning in reactive environments
J. Schmidhuber · 1990
Earlier work this paper cites.
A possibility for implementing curiosity and boredom in model-building neural controllers
J. Schmidhuber · 1991
Earlier work this paper cites.
The scientist in the crib: Minds, brains, and how children learn
A. Gopnik, A. N. Meltzoff, and P. K. Kuhl · 1999
Earlier work this paper cites.
An object-oriented representation for efficient reinforcement learning
C. Diuk, A. Cohen, and M. L. Littman · 2008
Earlier work this paper cites.
Formal theory of creativity, fun, and intrinsic motivation (1990–2010)
J. Schmidhuber · 2010
Earlier work this paper cites.
Pilco: A model-based and data-efficient approach to policy search
M. Deisenroth and C. E. Rasmussen · 2011
Earlier work this paper cites.
Active Learning
B. Settles · 2011
Earlier work this paper cites.
Object focused q-learning for autonomous agents
L. C. Cobo, C. L. Isbell, and A. L. Thomaz · 2013
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
S. Ioffe and C. Szegedy · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski, S. Petersen, C. Beattie, A. Sadik, I. Antonoglou, H. King, D. Kumaran, D. Wierstra, S. Legg, and D. Hassabis · 2015
Earlier work this paper cites.
Prioritized experience replay, 2015
T. Schaul, J. Quan, I. Antonoglou, and D. Silver · 2015
Cited alongside, same era.
Deep spatial autoencoders for visuomotor learning
C. Finn, X. Y. Tan, Y. Duan, T. Darrell, S. Levine, and P. Abbeel · 2016
Cited alongside, same era.
Towards deep symbolic reinforcement learning
M. Garnelo, K. Arulkumaran, and M. Shanahan · 2016
Cited alongside, same era.
Hierarchical deep reinforcement learning: Integrating temporal abstraction and intrinsic motivation
T. D. Kulkarni, K. Narasimhan, A. Saeedi, and J. Tenenbaum · 2016
Cited alongside, same era.
Mastering the game of Go with deep neural networks and tree search
D. Silver, A. Huang, C. J. Maddison, A. Guez, L. Sifre, G. van den Driessche, J. Schrittwieser, I. Antonoglou, V. Panneershelvam, M. Lanctot, S. Dieleman, D. Grewe, J. Nham, N. Kalchbrenner, I. Sutskever, T. Lillicrap, M. Leach, K. Kavukcuoglu, T. Graepel, and D. Hassabis · 2016
Cited alongside, same era.
Implicit quantile networks for distributional reinforcement learning
W. Dabney, G. Ostrovski, D. Silver, and R. Munos · 2018
Later among the works it cites.
Diversity is all you need: Learning skills without a reward function
B. Eysenbach, A. Gupta, J. Ibarz, and S. Levine · 2018
Later among the works it cites.
World models
D. Ha and J. Schmidhuber · 2018
Later among the works it cites.
Learning to play with intrinsically-motivated, self-aware agents
N. Haber, D. Mrowca, S. Wang, L. F. Fei-Fei, and D. L. Yamins · 2018
Later among the works it cites.
Learning latent dynamics for planning from pixels
D. Hafner, T. Lillicrap, I. Fischer, R. Villegas, D. Ha, H. Lee, and J. Davidson · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
M. G. Bellemare, W. Dabney, and R. Munos · 2017
Cited alongside, same era.
Distributional reinforcement learning with quantile regression
W. Dabney, M. Rowland, M. G. Bellemare, and R. Munos · 2017
Cited alongside, same era.
Curiosity-driven exploration by self-supervised prediction
D. Pathak, P. Agrawal, A. A. Efros, and T. Darrell · 2017
Cited alongside, same era.
M. Roderick, C. Grimm, and S. Tellex · 2017
Cited alongside, same era.
A simple neural network module for relational reasoning
A. Santoro, D. Raposo, D. G. T. Barrett, M. Malinowski, R. Pascanu, P. Battaglia, and T. P. Lillicrap · 2017
Cited alongside, same era.
Feudal networks for hierarchical reinforcement learning
A. S. Vezhnevets, S. Osindero, T. Schaul, N. Heess, M. Jaderberg, D. Silver, and K. Kavukcuoglu · 2017
Cited alongside, same era.
Maximum a posteriori policy optimisation
A. Abdolmaleki, J. T. Springenberg, Y. Tassa, R. Munos, N. Heess, and M. Riedmiller · 2018
Cited alongside, same era.
R. Keramati, J. Whang, P. Cho, and E. Brunskill · 2018
Later among the works it cites.
Learning plannable representations with causal infogan
T. Kurutach, A. Tamar, G. Yang, S. J. Russell, and P. Abbeel · 2018
Later among the works it cites.
Curiosity driven exploration of learned disentangled goal spaces
A. Laversanne-Finot, A. Pere, and P.-Y. Oudeyer · 2018
Later among the works it cites.
Reinforcement learning: An introduction
R. S. Sutton and A. G. Barto · 2018
Later among the works it cites.
SOLAR: deep structured latent representations for model-based reinforcement learning
M. Zhang, S. Vikram, L. Smith, P. Abbeel, M. J. Johnson, and S. Levine · 2018
Later among the works it cites.
MONet: Unsupervised scene decomposition and representation
C. P. Burgess, L. Matthey, N. Watters, R. Kabra, I. Higgins, M. Botvinick, and A. Lerchner · 2019
Closest in time.
Model-based reinforcement learning for atari
L. Kaiser, M. Babaeizadeh, P. Milos, B. Osinski, R. H. Campbell, K. Czechowski, D. Erhan, C. Finn, P. Kozakowski, S. Levine, et al · 2019
Closest in time.
Beyond imitation: Zero-shot task transfer on robots by learning concepts as cognitive programs
M. Lázaro-Gredilla, D. Lin, J. S. Guntupalli, and D. George · 2019
Closest in time.
Spriteworld: A flexible, configurable reinforcement learning environment
N. Watters, L. Matthey, S. Borgeaud, R. Kabra, and A. Lerchner · 2019
Closest in time.