Fetching the paper…
Reading the bibliography…
One of the main challenges in reinforcement learning (RL) is generalisation.
Reinforcement Learning: An Introduction
Sutton, R. S. and Barto, A. G · 1998
Earlier work this paper cites.
Policy invariance under reward transformations: theory and application to reward shaping
Ng, A. Y., Harada, D., and Russell, S · 1999
Earlier work this paper cites.
Between mdps and semi-mdps: A framework for temporal abstraction in reinforcement learning
Sutton, R.S., Precup, D., and Singh, S.P · 1999
Earlier work this paper cites.
Finite-time analysis of the multiarmed bandit problem
Auer, P., Cesa-Bianchi, N., and Fischer, P · 2002
Earlier work this paper cites.
Recent advances in hierarchical reinforcement learning
Barto, A. G. and Mahadevan, S · 2003
Earlier work this paper cites.
Cortex and mind: Unifying cognition
Fuster, J. M · 2003
Earlier work this paper cites.
Q-decomposition for reinforcement learning agents
Russell, S. and Zimdar, A. L · 2003
Earlier work this paper cites.
Multiple-goal reinforcement learning with modular sarsa(0)
Sprague, N. and Ballard, D · 2003
Earlier work this paper cites.
Intrinsically motivated reinforcement learning: A promising framework for developmental robotics
Stout, A., Konidaris, G., and Barto, A. G · 2005
Earlier work this paper cites.
An object-oriented representation for efficient reinforcement learning
Diuk, C., Cohen, A., and Littman, M. L · 2008
Earlier work this paper cites.
Algorithms for reinforcement learning
Szepesvári, C · 2009
Cited alongside, same era.
A theoretical and empirical analysis of expected sarsa
van Seijen, H., van Hasselt, H., Whiteson, S., and Wiering, M · 2009
Cited alongside, same era.
Formal theory of creativity, fun, and intrinsic motivation (1990–2010)
Schmidhuber, J · 2010
Cited alongside, same era.
Horde: A scalable real-time architecture for learning knowledge from unsupervised sensorimotor interaction
Sutton, R. S., Modayil, J., Delp, M., Degris, T., Pilarski, P. M., White, A., and Precup, Doina · 2011
Cited alongside, same era.
The arcade learning environment: An evaluation platform for general agents
Bellemare, M. G., Naddaf, Y., Veness, J., and Bowling, M · 2013
Cited alongside, same era.
Learning and memory: From brain to behavior
Gluck, M. A., Mercado, E., and Myers, C. E · 2013
Universal value function approximators
Schaul, T., Horgan, D., Gregor, K., and Silver, D · 2015
Later among the works it cites.
Hierarchical deep reinforcement learning: Integrating temporal abstraction and intrinsic motivation
Kulkarni, T. D., Narasimhan, K. R., Saeedi, A., and Tenenbaum, J. B · 2016
Later among the works it cites.
Asynchronous methods for deep reinforcement learning
Mnih, V., Badia, A. P., Mirza, M., Graves, A., Harley, T., Lillicrap, T. P., Silver, D., and Kavukcuoglu, K · 2016
Later among the works it cites.
Prioritized experience replay
Schaul, T., Quan, J., Antonoglou, I., and Silver, D · 2016
Later among the works it cites.
Strategic attentive writer for learning macro-actions
Vezhnevets, A., Mnih, V., Osindero, S., Graves, A., Vinyals, O., Agapiou, J., and Kavukcuoglu, K · 2016
Later among the works it cites.
Dueling network architectures for deep reinforcement learning
Wang, Z., Schaul, T., Hessel, M., van Hasselt, H., Lanctot, M., and Freitas, N · 2016
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
A survey of multi-objective sequential decision-making
Roijers, D. M., Vamplew, P., Whiteson, S., and Dazeley, R · 2013
Cited alongside, same era.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification
He, K., Zhang, X., Ren, S., and Sun, J · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., Petersen, S., Beattie, C., Sadik, A., Antonoglou, I., Kumaran, H. King D., Wierstra, D., Legg, S., and Hassabis, D · 2015
Cited alongside, same era.
Massively parallel methods for deep reinforcement learning
Nair, A., Srinivasan, P., Blackwell, S., Alcicek, C., Fearon, R., Maria, A. De, Panneershelvam, V., Suleyman, M., Beattie, C., Petersen, S., Legg, S., Mnih, V., Kavukcuoglu, K., and Silver, D · 2015
Cited alongside, same era.
Learning values across many orders of magnitude
van Hasselt, H., Guez, A., Hessel, M., Mnih, V., and Silver, D
Cited in the paper.
Later among the works it cites.
The option-critic architecture
Bacon, P., Harb, J., and Precup, D · 2017
Closest in time.
Reinforcement learning with unsupervised auxiliary tasks
Jaderberg, M., Mnih, V., Czarnecki, W.M., Schaul, T., Leibo, J.Z., Silver, D., and Kavukcuoglu, K · 2017
Closest in time.
Deep reinforcement learning with double q-learning
van Hasselt, H., Guez, A., and Silver, D · 2094
Closest in time.