Fetching the paper…
Reading the bibliography…
In a reward-free environment, what is a suitable intrinsic objective for an agent to pursue so that it can learn an optimal task-agnostic exploration policy? In this paper, we argue that the entropy of the state distribution induced by finite-horizon trajectories is a sensible target.
Go-explore: A new approach for hard-exploration problems
Ecoffet, A.; Huizinga, J.; Lehman, J.; Stanley, K. O.; and Clune, J. 2019 · 1901
Earlier work this paper cites.
Meta-learning via learned loss
Bechtle, S.; Molchanov, A.; Chebotar, Y.; Grefenstette, E.; Righetti, L.; Sukhatme, G.; and Meier, F. 2019 · 1906
Earlier work this paper cites.
Efficient exploration via state marginal matching
Lee, L.; Eysenbach, B.; Parisotto, E.; Xing, E.; Levine, S.; and Salakhutdinov, R. 2019 · 1906
Earlier work this paper cites.
Autonomous exploration for navigating in non-stationary CMPs
Gajane, P.; Ortner, R.; Auer, P.; and Szepesvari, C. 2019 · 1910
Earlier work this paper cites.
Dota 2 with large scale deep reinforcement learning
Berner, C.; Brockman, G.; Chan, B.; Cheung, V.; Debiak, P.; Dennison, C.; Farhi, D.; Fischer, Q.; Hashme, S.; Hesse, C.; et al. 2019 · 1912
Earlier work this paper cites.
A mathematical theory of communication
Shannon, C. E. 1948 · 1948
Earlier work this paper cites.
The Poisson approximation to the Poisson binomial distribution
Hodges, J. L.; and Le Cam, L. 1960 · 1960
Earlier work this paper cites.
Evolutionary principles in self-referential learning, or on learning how to learn: the meta-meta-… hook
Schmidhuber, J. 1987 · 1987
Earlier work this paper cites.
A possibility for implementing curiosity and boredom in model-building neural controllers
Schmidhuber, J. 1991 · 1991
Earlier work this paper cites.
Nonparametric entropy estimation: An overview
Beirlant, J.; Dudewicz, E. J.; Györfi, L.; and Van der Meulen, E. C. 1997 · 1997
Earlier work this paper cites.
Nearest neighbor estimates of entropy
Singh, H.; Misra, N.; Hnizdo, V.; Fedorowicz, A.; and Demchuk, E. 2003 · 2003
Earlier work this paper cites.
Active Model Estimation in Markov Decision Processes
Tarbouriech, J.; Shekhar, S.; Pirotta, M.; Ghavamzadeh, M.; and Lazaric, A. 2020 · 2003
Earlier work this paper cites.
Intrinsically motivated reinforcement learning
Chentanez, N.; Barto, A. G.; and Singh, S. P. 2005 · 2005
Earlier work this paper cites.
Incremental skill acquisition for self-motivated learning animats
Bonarini, A.; Lazaric, A.; and Restelli, M. 2006 · 2006
Earlier work this paper cites.
Self-development framework for reinforcement learning agents
Bonarini, A.; Lazaric, A.; Restelli, M.; and Vitali, P. 2006 · 2006
Earlier work this paper cites.
Adaptive reward-free exploration
Kaufmann, E.; Ménard, P.; Domingues, O. D.; Jonsson, A.; Leurent, E.; and Valko, M. 2020 · 2006
Earlier work this paper cites.
Task-agnostic exploration in reinforcement learning
Zhang, X.; Singla, A.; et al. 2020 · 2006
Earlier work this paper cites.
Intrinsic motivation systems for autonomous mental development
Oudeyer, P.-Y.; Kaplan, F.; and Hafner, V. V. 2007 · 2007
Earlier work this paper cites.
Differential entropy estimation by particles
Ajgl, J.; and Šimandl, M. 2011 · 2011
Earlier work this paper cites.
Scikit-learn: Machine learning in Python
Pedregosa, F.; Varoquaux, G.; Gramfort, A.; Michel, V.; Thirion, B.; Grisel, O.; Blondel, M.; Prettenhofer, P.; Weiss, R.; Dubourg, V.; Vanderplas, J.; Passos, A.; Cournapeau, D.; Brucher, M.; Perrot, M.; and Duchesnay, E. 2011 · 2011
Cited alongside, same era.
Autonomous exploration for navigating in mdps
Lim, S. H.; and Auer, P. 2012 · 2012
Cited alongside, same era.
Exploration in model-based reinforcement learning by empirically estimating learning progress
Lopes, M.; Lang, T.; Toussaint, M.; and Oudeyer, P.-Y. 2012 · 2012
Cited alongside, same era.
Mujoco: A physics engine for model-based control
Todorov, E.; Erez, T.; and Tassa, Y. 2012 · 2012
Cited alongside, same era.
A survey on policy search for robotics
Deisenroth, M. P.; Neumann, G.; Peters, J.; et al. 2013 · 2013
Cited alongside, same era.
Monte Carlo theory, methods and examples
Owen, A. B. 2013 · 2013
Curiosity-driven exploration by self-supervised prediction
Pathak, D.; Agrawal, P.; Efros, A. A.; and Darrell, T. 2017 · 2017
Later among the works it cites.
# Exploration: A study of count-based exploration for deep reinforcement learning
Tang, H.; Houthooft, R.; Foote, D.; Stooke, A.; Chen, O. X.; Duan, Y.; Schulman, J.; DeTurck, F.; and Abbeel, P. 2017 · 2017
Later among the works it cites.
Variational option discovery algorithms
Achiam, J.; Edwards, H.; Amodei, D.; and Abbeel, P. 2018 · 2018
Later among the works it cites.
Unsupervised meta-learning for reinforcement learning
Gupta, A.; Eysenbach, B.; Finn, C.; and Levine, S. 2018 · 2018
Later among the works it cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Haarnoja, T.; Zhou, A.; Abbeel, P.; and Levine, S. 2018 · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Markov decision processes: Discrete stochastic dynamic programming
Puterman, M. L. 2014 · 2014
Cited alongside, same era.
Empowerment–an introduction
Salge, C.; Glackin, C.; and Polani, D. 2014 · 2014
Cited alongside, same era.
Human-level control through deep reinforcement learning
Mnih, V.; Kavukcuoglu, K.; Silver, D.; Rusu, A. A.; Veness, J.; Bellemare, M. G.; Graves, A.; Riedmiller, M.; Fidjeland, A. K.; Ostrovski, G.; et al. 2015 · 2015
Cited alongside, same era.
Variational information maximisation for intrinsically motivated reinforcement learning
Mohamed, S.; and Rezende, D. J. 2015 · 2015
Cited alongside, same era.
Trust region policy optimization
Schulman, J.; Levine, S.; Abbeel, P.; Jordan, M.; and Moritz, P. 2015 · 2015
Cited alongside, same era.
Unifying count-based exploration and intrinsic motivation
Bellemare, M.; Srinivasan, S.; Ostrovski, G.; Schaul, T.; Saxton, D.; and Munos, R. 2016 · 2016
Cited alongside, same era.
Policy optimization via importance sampling
Metelli, A. M.; Papini, M.; Faccio, F.; and Restelli, M. 2018 · 2018
Later among the works it cites.
Reinforcement learning: An introduction
Sutton, R. S.; and Barto, A. G. 2018 · 2018
Later among the works it cites.
Quantifying generalization in reinforcement learning
Cobbe, K.; Klimov, O.; Hesse, C.; Kim, T.; and Schulman, J. 2019 · 2019
Later among the works it cites.
Diversity is all you need: Learning skills without a reward function
Eysenbach, B.; Gupta, A.; Ibarz, J.; and Levine, S. 2019 · 2019
Later among the works it cites.
A divergence minimization perspective on imitation learning methods
Ghasemipour, S. K. S.; Zemel, R. S.; and Gu, S. 2019 · 2019
Later among the works it cites.
Provably efficient maximum entropy exploration
Hazan, E.; Kakade, S.; Singh, K.; and Van Soest, A. 2019 · 2019
Later among the works it cites.
Active Exploration in Markov Decision Processes
Tarbouriech, J.; and Lazaric, A. 2019 · 2019
Later among the works it cites.
Learning dexterous in-hand manipulation
Andrychowicz, O. M.; Baker, B.; Chociej, M.; Jozefowicz, R.; McGrew, B.; Pachocki, J.; Petron, A.; Plappert, M.; Powell, G.; Ray, A.; et al. 2020 · 2020
Closest in time.
Reward-free exploration for reinforcement learning
Jin, C.; Krishnamurthy, A.; Simchowitz, M.; and Yu, T. 2020 · 2020
Closest in time.
Kinematic state abstraction and provably efficient rich-observation reinforcement learning
Misra, D.; Henaff, M.; Krishnamurthy, A.; and Langford, J. 2020 · 2020
Closest in time.
An intrinsically-motivated approach for learning highly exploring and fast mixing policies
Mutti, M.; and Restelli, M. 2020 · 2020
Closest in time.
Skew-fit: State-covering self-supervised reinforcement learning
Pong, V. H.; Dalal, M.; Lin, S.; Nair, A.; Bahl, S.; and Levine, S. 2020 · 2020
Closest in time.
What can learned intrinsic rewards capture?
Zheng, Z.; Oh, J.; Hessel, M.; Xu, Z.; Kroiss, M.; van Hasselt, H.; Silver, D.; and Singh, S. 2020 · 2020
Closest in time.