Fetching the paper…
Reading the bibliography…
Exploration is a key problem in reinforcement learning, since agents can only learn from data they acquire in the environment.
On the likelihood that one unknown probability exceeds another in view of the evidence of two samples
W. R. Thompson · 1933
Earlier work this paper cites.
Introduction to Reinforcement Learning
R. S. Sutton and A. G. Barto · 1998
Earlier work this paper cites.
Exploiting open-endedness to solve problems through the search for novelty
J. Lehman and K. O. Stanley · 2008
Earlier work this paper cites.
Formal theory of creativity, fun, and intrinsic motivation (1990–2010)
J. Schmidhuber · 2010
Earlier work this paper cites.
Gaussian process optimization in the bandit setting: No regret and experimental design
N. Srinivas, A. Krause, S. Kakade, and M. Seeger · 2010
Earlier work this paper cites.
Critical factors in the performance of novelty search
S. Kistemaker and S. Whiteson · 2011
Earlier work this paper cites.
k-dpps: Fixed-size determinantal point processes
A. Kulesza and B. Taskar · 2011
Earlier work this paper cites.
Abandoning objectives: Evolution through the search for novelty alone
J. Lehman and K. O. Stanley · 2011
Earlier work this paper cites.
Hogwild: A lock-free approach to parallelizing stochastic gradient descent
B. Recht, C. Re, S. Wright, and F. Niu · 2011
Earlier work this paper cites.
Enumerative combinatorics volume 1 second edition
R. P. Stanley · 2011
Earlier work this paper cites.
Analysis of Thompson sampling for the multi-armed bandit problem
S. Agrawal and N. Goyal · 2012
Earlier work this paper cites.
Determinantal Point Processes for Machine Learning
A. Kulesza and B. Taskar · 2012
Earlier work this paper cites.
Playing atari with deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. Graves, I. Antonoglou, D. Wierstra, and M. A. Riedmiller · 2013
Earlier work this paper cites.
Robots that can adapt like animals
A. Cully, J. Clune, D. Tarapore, and J.-B. Mouret · 2015
Earlier work this paper cites.
Illuminating search spaces by mapping elites
J.-B. Mouret and J. Clune · 2015
Earlier work this paper cites.
Trust region policy optimization
J. Schulman, S. Levine, P. Abbeel, M. Jordan, and P. Moritz · 2015
Earlier work this paper cites.
OpenAI Gym, 2016
G. Brockman, V. Cheung, L. Pettersson, J. Schneider, J. Schulman, J. Tang, and W. Zaremba · 2016
Earlier work this paper cites.
Deep exploration via bootstrapped DQN
I. Osband, C. Blundell, A. Pritzel, and B. Van Roy · 2016
Earlier work this paper cites.
Quality diversity: A new frontier for evolutionary computation
J. K. Pugh, L. B. Soros, and K. O. Stanley · 2016
Earlier work this paper cites.
An information-theoretic analysis of Thompson sampling
D. Russo and B. Van Roy · 2016
Earlier work this paper cites.
Model-agnostic meta-learning for fast adaptation of deep networks
C. Finn, P. Abbeel, and S. Levine · 2017
Earlier work this paper cites.
Ray: A distributed framework for emerging AI applications
P. Moritz, R. Nishihara, S. Wang, A. Tumanov, R. Liaw, E. Liang, W. Paul, M. I. Jordan, and I. Stoica · 2017
Cited alongside, same era.
Random gradient-free minimization of convex functions
Y. Nesterov and V. Spokoiny · 2017
Cited alongside, same era.
A tutorial on Thompson sampling
D. Russo, B. Roy, A. Kazerouni, and I. Osband · 2017
Cited alongside, same era.
Evolution Strategies as a scalable alternative to reinforcement learning
T. Salimans, J. Ho, X. Chen, S. Sidor, and I. Sutskever · 2017
Cited alongside, same era.
F. P. Such, V. Madhavan, E. Conti, J. Lehman, K. O. Stanley, and J. Clune · 2017
Cited alongside, same era.
Evolvability ES: Scalable and direct optimization of evolvability
A. Gajewski, J. Clune, K. O. Stanley, and J. Lehman · 2019
Later among the works it cites.
Novelty search for deep reinforcement learning policy network weights by action sequence edit metric distance
E. C. Jackson and M. Daley · 2019
Later among the works it cites.
Evolutionary reinforcement learning for sample-efficient multiagent coordination
S. Khadka, S. Majumdar, and K. Tumer · 2019
Later among the works it cites.
A generalized framework for population based training
A. Li, O. Spyra, S. Perel, V. Dalibard, M. Jaderberg, C. Gu, D. Budden, T. Harley, and P. Gupta · 2019
Later among the works it cites.
Emergent coordination through competition
S. Liu, G. Lever, N. Heess, J. Merel, S. Tunyasuvunakool, and T. Graepel · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Structured evolution with compact architectures for scalable policy optimization
K. Choromanski, M. Rowland, V. Sindhwani, R. E. Turner, and A. Weller · 2018
Cited alongside, same era.
Improving exploration in Evolution Strategies for deep reinforcement learning via a population of novelty-seeking agents
E. Conti, V. Madhavan, F. P. Such, J. Lehman, K. O. Stanley, and J. Clune · 2018
Cited alongside, same era.
Noisy networks for exploration
M. Fortunato, M. G. Azar, B. Piot, J. Menick, M. Hessel, I. Osband, A. Graves, V. Mnih, R. Munos, D. Hassabis, O. Pietquin, C. Blundell, and S. Legg · 2018
Cited alongside, same era.
Addressing function approximation error in actor-critic methods
S. Fujimoto, H. van Hoof, and D. Meger · 2018
Cited alongside, same era.
Meta-reinforcement learning of structured exploration strategies
A. Gupta, R. Mendonca, Y. Liu, P. Abbeel, and S. Levine · 2018
Cited alongside, same era.
Soft actor-critic algorithms and applications
T. Haarnoja, A. Zhou, K. Hartikainen, G. Tucker, S. Ha, J. Tan, V. Kumar, H. Zhu, A. Gupta, P. Abbeel, and S. Levine · 2018
Cited alongside, same era.
Diversity-driven exploration strategy for deep reinforcement learning
Z.-W. Hong, T.-Y. Shann, S.-Y. Su, Y.-H. Chang, T.-J. Fu, and C.-Y. Lee · 2018
Cited alongside, same era.
CEM-RL: Combining evolutionary and gradient-based methods for policy search
Pourchot and Sigaud · 2019
Later among the works it cites.
ProMP: Proximal meta-policy search
J. Rothfuss, D. Lee, I. Clavera, T. Asfour, and P. Abbeel · 2019
Later among the works it cites.
Meta-learning with latent embedding optimization
A. A. Rusu, D. Rao, J. Sygnowski, O. Vinyals, R. Pascanu, S. Osindero, and R. Hadsell · 2019
Later among the works it cites.
Adapting behaviour for learning progress
T. Schaul, D. Borsa, D. Ding, D. Szepesvari, G. Ostrovski, W. Dabney, and S. Osindero · 2019
Later among the works it cites.
Introduction to multi-armed bandits
A. Slivkins · 2019
Later among the works it cites.
Designing neural networks through neuroevolution
K. Stanley, J. Clune, J. Lehman, and R. Miikkulainen · 2019
Later among the works it cites.
Nonlinear stein variational gradient descent for learning diversified mixture models
D. Wang and Q. Liu · 2019
Later among the works it cites.
A generalized algorithm for multi-objective reinforcement learning and policy adaptation
R. Yang, X. Sun, and K. Narasimhan · 2019
Later among the works it cites.
Model-based reinforcement learning for biological sequence design
C. Angermueller, D. Dohan, D. Belanger, R. Deshpande, K. Murphy, and L. Colwell · 2020
Closest in time.
Agent57: Outperforming the atari human benchmark
A. P. Badia, B. Piot, S. Kapturowski, P. Sprechmann, A. Vitvitskyi, D. Guo, and C. Blundell · 2020
Closest in time.
Ready policy one: World building through active learning
P. Ball, J. Parker-Holder, A. Pacchiano, K. Choromanski, and S. Roberts · 2020
Closest in time.
Automating representation discovery with MAP-Elites
A. Gaier, A. Asteroth, and J. Mouret · 2020
Closest in time.
Dynamical distance learning for semi-supervised and unsupervised skill discovery
K. Hartikainen, X. Geng, T. Haarnoja, and S. Levine · 2020
Closest in time.
Population-guided parallel policy search for reinforcement learning
W. Jung, G. Park, and Y. Sung · 2020
Closest in time.
Learning to score behaviors for guided policy optimization
A. Pacchiano, J. Parker-Holder, Y. Tang, A. Choromanska, K. Choromanski, and M. I. Jordan · 2020
Closest in time.