Fetching the paper…
Reading the bibliography…
Efficient exploration remains a challenging research problem in reinforcement learning, especially when an environment contains large state spaces, deceptive local optima, or sparse rewards.
Reinforcement learning: An introduction
R. S. Sutton and A. G. Barto · 1998
Earlier work this paper cites.
Relative entropy policy search
Jan Peters, Katharina Mülling, and Yasemin Altun · 2010
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
E. Todorov, T. Erez, and Y. Tassa · 2012
Earlier work this paper cites.
The arcade learning environment: An evaluation platform for general agents
M. G. Bellemare, Y. Naddaf, J. Veness, and M. Bowling · 2013
Earlier work this paper cites.
Deterministic policy gradient algorithms
D. et al · 2014
Earlier work this paper cites.
Human-level control through deep reinforcement learning
V. et al · 2015
Earlier work this paper cites.
Incentivizing exploration in reinforcement learning with deep predictive models
B. C. Stadie, S. Levine, and P. Abbeel · 2015
Earlier work this paper cites.
Trust region policy optimization
J. et al · 2015
Earlier work this paper cites.
Mastering the game of Go with deep neural networks and tree search
D. et al · 2016
Cited alongside, same era.
Learning deep neural network policies with continuous memory states
M. et al · 2016
Cited alongside, same era.
Asynchronous methods for deep reinforcement learning
V. et al · 2016
Cited alongside, same era.
Vime: Variational information maximizing exploration
R. et al · 2016
Cited alongside, same era.
Conditional image generation with pixelcnn decoders
A. et al · 2016
Cited alongside, same era.
Deep exploration via bootstrapped DQN
I. Osband, C. Blundell, A. Pritzel, and B. Van Roy · 2016
Cited alongside, same era.
Deep exploration via randomized value functions
I. Osband, D. Russo, Z. Wen, and B. Van Roy · 2017
Later among the works it cites.
#Exploration: A study of count-based exploration for deep reinforcement learning
H. et al · 2017
Later among the works it cites.
Curiosity-driven exploration by self-supervised prediction
D. Pathak, P. Agrawal, A. A. Efros, and T. Darrell · 2017
Later among the works it cites.
E. et al · 2017
Later among the works it cites.
Scalable trust-region method for deep reinforcement learning using Kronecker-factored approximation
Y. et al · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
T. P. et al · 2016
Cited alongside, same era.
Unifying count-based exploration and intrinsic motivation
M. et al · 2016
Cited alongside, same era.
Abandoning objectives: Evolution through the search for novelty alone
J. Lehman and K. O. Stanley
Cited in the paper.
Evolving a diversity of virtual creatures through novelty search and local competition
J. Lehman and K. O. Stanley
Cited in the paper.
Novelty search and the problem with objectives
J. Lehman and K. O. Stanley
Cited in the paper.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
T. Haarnoja, A. Zhou, P. Abbeel, and S. Levine
Cited in the paper.
J. et al · 2017
Later among the works it cites.
Parameter space noise for exploration
M. et al · 2018
Closest in time.
Noisy networks for exploration
M. et al · 2018
Closest in time.