Fetching the paper…
Reading the bibliography…
We consider the problem of exploration in meta reinforcement learning.
Evolutionary principles in self-referential learning, or on learning how to learn: The meta-meta-… hook
Schmidhuber · 1987
Earlier work this paper cites.
A Possibility for Implementing Curiosity and Boredom in Model-Building Neural Controllers
Schmidhuber, Jürgen · 1991
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Williams, Ronald J · 1992
Earlier work this paper cites.
Reinforcement driven information acquisition in non-deterministic environments
Storck, Jan, Hochreiter, Sepp, and Schmidhuber, Jürgen · 1995
Earlier work this paper cites.
Is learning the n-th thing any easier than learning the first?
Thrun · 1996
Earlier work this paper cites.
Reinforcement learning with self-modifying policies
Schmidhuber, J., Zhao, J., and Schraudolph, N · 1997
Earlier work this paper cites.
Hq-learning
Wiering and Schmidhuber · 1997
Earlier work this paper cites.
Exploration strategies for model-based learning in multi-agent systems
Carmel, D. and Markovitch, S · 1999
Earlier work this paper cites.
R-max - a general polynomial time algorithm for near-optimal reinforcement learning
Brafman, R. I. and Tennenholtz, M · 2002
Earlier work this paper cites.
Near-optimal reinforcement learning in polynomial time
Kearns, M. J. and Singh, S. P · 2002
Earlier work this paper cites.
Exploring the predictable
Kompella, Stollenga, Luciw, and Schmidhuber · 2002
Earlier work this paper cites.
Recent advances in hierarchical reinforcement learning
Barto and Mahadevan · 2003
Earlier work this paper cites.
Gödel machines: Fully self-referential optimal universal self-improvers
Schmidhuber · 2006
Cited alongside, same era.
Near-bayesian exploration in polynomial time
Kolter, J. Z. and Ng, A. Y · 2009
Cited alongside, same era.
Transfer learning for reinforcement learning domains: A survey
Taylor and Stone · 2009
Cited alongside, same era.
Artificial Curiosity for Autonomous Space Exploration
Graziano, Vincent, Glasmachers, Tobias, Schaul, Tom, Pape, Leo, Cuccu, Giuseppe, Leitner, Jurgen, and Schmidhuber, Jürgen · 2011
Cited alongside, same era.
Planning to be surprised: Optimal Bayesian exploration in dynamic environments
Sun, Yi, Gomez, Faustino, and Schmidhuber, Jürgen · 2011
Cited alongside, same era.
Learning skills from play: Artificial curiosity on a Katana robot arm
Ngo, Hung, Luciw, Matthew, Forster, Alexander, and Schmidhuber, Juergen · 2012
Unifying count based exploration and intrinsic motivation
Bellemare, M., Srinivasan, S., Ostrovski, G., Schaul, T., Saxton, D., and Munos, R · 2016
Later among the works it cites.
Vime: Variational information maximizing exploration
Houthooft, R., Chen, X., Duan, Y., Schulman, J., De Turck, F., and Abbeel, P · 2016
Later among the works it cites.
Asynchronous methods for deep reinforcement learning
Mnih, Volodymyr, Badia, Adria Puigdomenech, Mirza, Mehdi, Graves, Alex, Lillicrap, Timothy P, Harley, Tim, Silver, David, and Kavukcuoglu, Koray · 2016
Later among the works it cites.
Deep exploration via bootstrapped dqn
Osband, I., Blundell, C., Pritzel, A., and Van Roy, B · 2016
Later among the works it cites.
Rusu, A. A., Rabinowitz, N. C., Desjardins, G., Soyer, H., Kirkpatrick, J., Kavukcuoglu, K., Pascanu, R., and Hadsell, R · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Evolving large-scale neural networks for vision-based reinforcement learning
Koutník, Jan, Cuccu, Giuseppe, Schmidhuber, Jürgen, and Gomez, Faustino · 2013
Cited alongside, same era.
Lifelong machine learning systems: Beyond learning algorithms
Silver, Yand, and Li · 2013
Cited alongside, same era.
The option-critic architecture
Bacon, P. and Precup, D · 2015
Cited alongside, same era.
Continuous control with deep reinforcement learning
Lillicrap, T., Hunt, J., Pritzel, A., Heess, N., Erez, T., Tassa, Y., Silver, D., and Wierstra, D · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
Mnih, Volodymyr, Kavukcuoglu, Koray, Silver, David, Rusu, Andrei A, Veness, Joel, Bellemare, Marc G, Graves, Alex, Riedmiller, Martin, Fidjeland, Andreas K, Ostrovski, Georg, et al · 2015
Cited alongside, same era.
Incentivizing exploration in reinforcement learning with deep predictive models
Stadie, Bradly C., Levine, S., and Abbeel, P · 2015
Cited alongside, same era.
exploration: A study of count-based exploration for deep reinforcement learning
Tang, H., Houthooft, R., Foote, D., Stooke, A., Chen, X., Duan, Y., Schulman, J., De Turck, F., and Abbeel, P · 2016
Later among the works it cites.
A deep hierarchical approach to lifelong learning in minecraft
Tessler, C. Givony, S., Zahavy, T., Mankowitz, D.J., and Mannor, S · 2016
Later among the works it cites.
Model-agnostic metalearning for fast adaptation of deep networks
Finn, C., Abbeel, P., and Levine, S · 2017
Later among the works it cites.
Count-based exploration with neural density models
Ostrovski, G., Bellemare, M. G., Oord, A. v. d., and Munos, R · 2017
Later among the works it cites.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Later among the works it cites.
Feudal networks for hierarchical reinforcement learning
Vezhnevets, A. S., Osindero, S., Schaul, T., Heess, N., Jaderberg, M., Silver, D., and Kavukcuoglu, K · 2017
Later among the works it cites.
Promp: Proximal meta-policy search
Anonymous · 2019
Closest in time.