Fetching the paper…
Reading the bibliography…
Exploration in high-dimensional, continuous spaces with sparse rewards is an open problem in reinforcement learning.
Efficient exploration via state marginal matching
L. Lee, B. Eysenbach, E. Parisotto, E. P. Xing, S. Levine, and R. Salakhutdinov · 1906
Earlier work this paper cites.
Self-supervised exploration via disagreement
D. Pathak, D. Gandhi, and A. Gupta · 1906
Earlier work this paper cites.
Dota 2 with large scale deep reinforcement learning
C. Berner, G. Brockman, B. Chan, V. Cheung, P. Debiak, C. Dennison, D. Farhi, Q. Fischer, S. Hashme, C. Hesse, R. Józefowicz, S. Gray, C. Olsson, J. Pachocki, M. Petrov, H. P. de Oliveira Pinto, J. Raiman, T. Salimans, J. Schlatter, J. Schneider, S. Sidor, I. Sutskever, J. Tang, F. Wolski, and S. Zhang · 1912
Earlier work this paper cites.
Asymptotic properties of k-means clustering algorithm as a density estimation procedure
M. A. Wong · 1980
Earlier work this paper cites.
A possibility for implementing curiosity and boredom in model-building neural controllers
J. Schmidhuber · 1991
Earlier work this paper cites.
Improving generalisation for temporal difference learning: The successor representation
P. Dayan · 1993
Earlier work this paper cites.
A new class of entropy estimators for multi-dimensional densities
E. Miller · 2003
Earlier work this paper cites.
dm_control: Software and tasks for continuous control
Y. Tassa, S. Tunyasuvunakool, A. Muldal, Y. Doron, S. Liu, S. Bohez, J. Merel, T. Erez, T. P. Lillicrap, and N. Heess · 2006
Earlier work this paper cites.
A policy gradient method for task-agnostic exploration
M. Mutti, L. Pratissoli, and M. Restelli · 2007
Earlier work this paper cites.
An analysis of model-based interval estimation for markov decision processes
A. L. Strehl and M. L. Littman · 2008
Cited alongside, same era.
Near-bayesian exploration in polynomial time
J. Z. Kolter and A. Y. Ng · 2009
Cited alongside, same era.
Human-level control through deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski, S. Petersen, C. Beattie, A. Sadik, I. Antonoglou, H. King, D. Kumaran, D. Wierstra, S. Legg, and D. Hassabis · 2015
Cited alongside, same era.
Unifying count-based exploration and intrinsic motivation
M. G. Bellemare, S. Srinivasan, G. Ostrovski, T. Schaul, D. Saxton, and R. Munos · 2016
Cited alongside, same era.
Curiosity-driven exploration in deep reinforcement learning via bayesian neural networks
R. Houthooft, X. Chen, Y. Duan, J. Schulman, F. D. Turck, and P. Abbeel · 2016
Cited alongside, same era.
Exploration by random network distillation
Y. Burda, H. Edwards, A. J. Storkey, and O. Klimov · 2018
Later among the works it cites.
Provably efficient maximum entropy exploration
E. Hazan, S. M. Kakade, K. Singh, and A. V. Soest · 2018
Later among the works it cites.
Grandmaster level in starcraft ii using multi-agent reinforcement learning
O. Vinyals, I. Babuschkin, W. M. Czarnecki, M. Mathieu, A. Dudzik, J. Chung, D. H. Choi, R. Powell, T. Ewalds, P. Georgiev, J. Oh, D. Horgan, M. Kroiss, I. Danihelka, A. Huang, L. Sifre, T. Cai, J. P. Agapiou, M. Jaderberg, A. S. Vezhnevets, R. Leblond, T. Pohlen, V. Dalibard, D. Budden, Y. Sulsky, J. Molloy, T. L. Paine, C. Gulcehre, Z. Wang, T. Pfaff, Y. Wu, R. Ring, D. Yogatama, D. Wünsch, K. McKinney, O. Smith, T. Schaul, T. Lillicrap, K. Kavukcuoglu, D. Hassabis, C. Apps, and D. Silver · 2019
Later among the works it cites.
Experiment tracking with weights and biases, 2020
L. Biewald · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Mastering the game of go with deep neural networks and tree search
D. Silver, A. Huang, C. J. Maddison, A. Guez, L. Sifre, G. van den Driessche, J. Schrittwieser, I. Antonoglou, V. Panneershelvam, M. Lanctot, S. Dieleman, D. Grewe, J. Nham, N. Kalchbrenner, I. Sutskever, T. Lillicrap, M. Leach, K. Kavukcuoglu, T. Graepel, and D. Hassabis · 2016
Cited alongside, same era.
#exploration: A study of count-based exploration for deep reinforcement learning
H. Tang, R. Houthooft, D. Foote, A. Stooke, X. Chen, Y. Duan, J. Schulman, F. D. Turck, and P. Abbeel · 2016
Cited alongside, same era.
Curiosity-driven exploration by self-supervised prediction
D. Pathak, P. Agrawal, A. A. Efros, and T. Darrell · 2017
Cited alongside, same era.
Proximal policy optimization algorithms
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov · 2017
Cited alongside, same era.
Behavior from the void: Unsupervised active pre-training
H. Liu and P. Abbeel
Cited in the paper.
APS: active pretraining with successor features
H. Liu and P. Abbeel
Cited in the paper.
Z. D. Guo, M. G. Azar, A. Saade, S. Thakoor, B. Piot, B. Á. Pires, M. Valko, T. Mesnard, T. Lattimore, and R. Munos · 2021
Later among the works it cites.
URLB: unsupervised reinforcement learning benchmark
M. Laskin, D. Yarats, H. Liu, K. Lee, A. Zhan, K. Lu, C. Cang, L. Pinto, and P. Abbeel · 2021
Later among the works it cites.
Stable-baselines3: Reliable reinforcement learning implementations
A. Raffin, A. Hill, A. Gleave, A. Kanervisto, M. Ernestus, and N. Dormann · 2021
Later among the works it cites.
State entropy maximization with random encoders for efficient exploration
Y. Seo, L. Chen, J. Shin, H. Lee, P. Abbeel, and K. Lee · 2021
Later among the works it cites.
Reinforcement learning with prototypical representations
D. Yarats, R. Fergus, A. Lazaric, and L. Pinto · 2021
Later among the works it cites.