Fetching the paper…
Reading the bibliography…
Curiosity-based reward schemes can present powerful exploration mechanisms which facilitate the discovery of solutions for complex, sparse or long-horizon tasks.
Catastrophic interference in connectionist networks: The sequential learning problem
M. McCloskey and N. J. Cohen · 1989
Earlier work this paper cites.
A possibility for implementing curiosity and boredom in model-building neural controllers
J. Schmidhuber · 1991
Earlier work this paper cites.
Empowerment: A universal agent-centric measure of control
A. S. Klyubin, D. Polani, and C. L. Nehaniv · 2005
Earlier work this paper cites.
Intrinsically motivated reinforcement learning
S. Singh, A. G. Barto, and N. Chentanez · 2005
Earlier work this paper cites.
Intrinsic motivation systems for autonomous mental development
P.-Y. Oudeyer, F. Kaplan, and V. V. Hafner · 2007
Earlier work this paper cites.
R-iac: Robust intrinsically motivated exploration and active learning
A. Baranes and P.-Y. Oudeyer · 2009
Earlier work this paper cites.
What is intrinsic motivation? a typology of computational approaches
P.-Y. Oudeyer and F. Kaplan · 2009
Earlier work this paper cites.
Formal theory of creativity, fun, and intrinsic motivation (1990–2010)
J. Schmidhuber · 2010
Earlier work this paper cites.
Representation matters: Improving perception and exploration for robotics
M. Wulfmeier, A. Byravan, T. Hertweck, I. Higgins, A. Gupta, T. Kulkarni, M. Reynolds, D. Teplyashin, R. Hafner, T. Lampe, et al · 2011
Earlier work this paper cites.
Exploration in model-based reinforcement learning by empirically estimating learning progress
M. Lopes, T. Lang, M. Toussaint, and P.-Y. Oudeyer · 2012
Earlier work this paper cites.
An information-theoretic approach to curiosity-driven reinforcement learning
S. Still and D. Precup · 2012
Earlier work this paper cites.
Powerplay: Training an increasingly general problem solver by continually searching for the simplest still unsolvable problem
J. Schmidhuber · 2013
Earlier work this paper cites.
First experiments with powerplay
R. K. Srivastava, B. R. Steunebrink, and J. Schmidhuber · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2014
Cited alongside, same era.
Variational information maximisation for intrinsically motivated reinforcement learning
S. Mohamed and D. J. Rezende · 2015
Cited alongside, same era.
Unifying count-based exploration and intrinsic motivation
M. Bellemare, S. Srinivasan, G. Ostrovski, T. Schaul, D. Saxton, and R. Munos · 2016
Cited alongside, same era.
K. Gregor, D. J. Rezende, and D. Wierstra · 2016
Cited alongside, same era.
Vime: Variational information maximizing exploration
R. Houthooft, X. Chen, Y. Duan, J. Schulman, F. De Turck, and P. Abbeel · 2016
Plan online, learn offline: Efficient learning and exploration via model-based control
K. Lowrey, A. Rajeswaran, S. Kakade, E. Todorov, and I. Mordatch · 2018
Later among the works it cites.
Learning by playing solving sparse reward tasks from scratch
M. Riedmiller, R. Hafner, T. Lampe, M. Neunert, J. Degrave, T. Wiele, V. Mnih, N. Heess, and J. T. Springenberg · 2018
Later among the works it cites.
Learning gentle object manipulation with curiosity-driven deep reinforcement learning
S. H. Huang, M. Zambelli, J. Kay, M. F. Martins, Y. Tassa, P. M. Pilarski, and R. Hadsell · 2019
Later among the works it cites.
Self-supervised exploration via disagreement
D. Pathak, D. Gandhi, and A. Gupta · 2019
Later among the works it cites.
Dynamics-aware unsupervised discovery of skills
A. Sharma, S. Gu, S. Levine, V. Kumar, and K. Hausman · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Loss is its own reward: Self-supervision for reinforcement learning
E. Shelhamer, P. Mahmoudieh, M. Argus, and T. Darrell · 2016
Cited alongside, same era.
Surprise-based intrinsic motivation for deep reinforcement learning
J. Achiam and S. Sastry · 2017
Cited alongside, same era.
Curiosity-driven exploration by self-supervised prediction
D. Pathak, P. Agrawal, A. A. Efros, and T. Darrell · 2017
Cited alongside, same era.
Intrinsic motivation and automatic curricula via asymmetric self-play
S. Sukhbaatar, Z. Lin, I. Kostrikov, G. Synnaeve, A. Szlam, and R. Fergus · 2017
Cited alongside, same era.
Maximum a posteriori policy optimisation
A. Abdolmaleki, J. T. Springenberg, Y. Tassa, R. Munos, N. Heess, and M. Riedmiller · 2018
Cited alongside, same era.
Diversity is all you need: Learning skills without a reward function
B. Eysenbach, A. Gupta, J. Ibarz, and S. Levine · 2018
Cited alongside, same era.
Learning to play with intrinsically-motivated, self-aware agents
N. Haber, D. Mrowca, S. Wang, F.-F. Li, and D. L. Yamins · 2018
Cited alongside, same era.
Later among the works it cites.
Embracing change: Continual learning in deep neural networks
R. Hadsell, D. Rao, A. A. Rusu, and R. Pascanu · 2020
Later among the works it cites.
Simple sensor intentions for exploration
T. Hertweck, M. Riedmiller, M. Bloesch, J. T. Springenberg, N. Siegel, M. Wulfmeier, R. Hafner, and N. Heess · 2020
Later among the works it cites.
Learning latent plans from play
C. Lynch, M. Khansari, T. Xiao, V. Kumar, J. Tompson, S. Levine, and P. Sermanet · 2020
Later among the works it cites.
Planning to explore via self-supervised world models
R. Sekar, O. Rybkin, K. Daniilidis, P. Abbeel, D. Hafner, and D. Pathak · 2020
Later among the works it cites.
Emergent real-world robotic skills via unsupervised off-policy reinforcement learning
A. Sharma, M. Ahn, S. Levine, V. Kumar, K. Hausman, and S. Gu · 2020
Later among the works it cites.
Asymmetric self-play for automatic goal discovery in robotic manipulation
O. OpenAI, M. Plappert, R. Sampedro, T. Xu, I. Akkaya, V. Kosaraju, P. Welinder, R. D’Sa, A. Petron, H. P. d. O. Pinto, et al · 2021
Closest in time.
Data-efficient hindsight off-policy option learning
M. Wulfmeier, D. Rao, R. Hafner, T. Lampe, A. Abdolmaleki, T. Hertweck, M. Neunert, D. Tirumala, N. Siegel, N. Heess, et al · 2021
Closest in time.