Fetching the paper…
Reading the bibliography…
Model-based reinforcement learning (RL) is appealing because (i) it enables planning and thus more strategic exploration, and (ii) by decoupling dynamics from rewards, it enables fast transfer to new reward functions.
Learning from delayed rewards
Watkins, C · 1989
Earlier work this paper cites.
Integrated architectures for learning, planning, and reacting based on approximating dynamic programming
Sutton, R. S · 1990
Earlier work this paper cites.
Planning simple trajectories using neural subgoal generators
Schmidhuber, J · 1993
Earlier work this paper cites.
Gambling in a rigged casino: The adversarial multi-armed bandit problem
Auer, P., Cesa-Bianchi, N., Freund, Y., and Schapire, R. E · 1995
Earlier work this paper cites.
Reinforcement learning with soft state aggregation
Singh, S. P., Jaakkola, T., and Jordan, M · 1995
Earlier work this paper cites.
The MAXQ method for hierarchical reinforcement learning
Dietterich, T. G · 1998
Earlier work this paper cites.
Between mdps and semi-mdps: A framework for temporal abstraction in reinforcement learning
Sutton, R. S., Precup, D., and Singh, S · 1999
Earlier work this paper cites.
State abstraction for programmable reinforcement learning agents
Andre, D. and Russell, S. J · 2002
Earlier work this paper cites.
R-max-a general polynomial time algorithm for near-optimal reinforcement learning
Brafman, R. and Tennenholtz, M · 2002
Earlier work this paper cites.
Near-optimal reinforcement learning in polynomial time
Kearns, M. and Singh, S · 2002
Earlier work this paper cites.
On the sample complexity of reinforcement learning
Kakade, S. M. et al · 2003
Earlier work this paper cites.
Using inaccurate models in reinforcement learning
Abbeel, P., Quigley, M., and Ng, A. Y · 2006
Earlier work this paper cites.
Towards a unified theory of state abstraction for mdps
Li, L., Walsh, T. J., and Littman, M. L · 2006
Earlier work this paper cites.
An analysis of model-based interval estimation for markov decision processes
Strehl, A. L. and Littman, M. L · 2008
Earlier work this paper cites.
Rectified linear units improve restricted boltzmann machines
Nair, V. and Hinton, G. E · 2010
Earlier work this paper cites.
The arcade learning environment: An evaluation platform for general agents
Bellemare, M. G., Naddaf, Y., Veness, J., and Bowling, M · 2013
Earlier work this paper cites.
Deep learning for real-time atari game play using offline monte-carlo tree search planning
Guo, X., Singh, S., Lee, H., Lewis, R. L., and Wang, X · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. and Ba, J · 2014
Earlier work this paper cites.
Model regularization for stable sample rollouts
Talvitie, E · 2014
Cited alongside, same era.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., et al · 2015
Cited alongside, same era.
Action-conditional video prediction using deep networks in atari games
Oh, J., Guo, X., Lee, H., Lewis, R. L., and Singh, S · 2015
Cited alongside, same era.
Agnostic system identification for monte carlo planning
Talvitie, E · 2015
Cited alongside, same era.
Imagination-augmented agents for deep reinforcement learning
Weber, T., Racanière, S., Reichert, D. P., Buesing, L., Guez, A., Rezende, D. J., Badia, A. P., Vinyals, O., Heess, N., Li, Y., et al · 2015
Cited alongside, same era.
Unifying count-based exploration and intrinsic motivation
Overcoming exploration in reinforcement learning with demonstrations
Nair, A., McGrew, B., Andrychowicz, M., Zaremba, W., and Abbeel, P · 2017
Later among the works it cites.
Count-based exploration with neural density models
Ostrovski, G., Bellemare, M. G., Oord, A., and Munos, R · 2017
Later among the works it cites.
Roderick, M., Grimm, C., and Tellex, S · 2017
Later among the works it cites.
World of bits: An open-domain platform for web-based agents
Shi, T., Karpathy, A., Fan, L., Hernandez, J., and Liang, P · 2017
Later among the works it cites.
#exploration: A study of count-based exploration for deep reinforcement learning
Tang, H., Houthooft, R., Foote, D., Stooke, A., Chen, X., Duan, Y., Schulman, J., DeTurck, F., and Abbeel, P · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Bellemare, M., Srinivasan, S., Ostrovski, G., Schaul, T., Saxton, D., and Munos, R · 2016
Cited alongside, same era.
Unsupervised learning for physical interaction through video prediction
Finn, C., Goodfellow, I., and Levine, S · 2016
Cited alongside, same era.
From softmax to sparsemax: A sparse model of attention and multi-label classification
Martins, A. and Astudillo, R · 2016
Cited alongside, same era.
Deep reinforcement learning with double Q-learning
van Hasselt, H., Guez, A., and Silver, D · 2016
Cited alongside, same era.
Dueling network architectures for deep reinforcement learning
Wang, Z., Schaul, T., Hessel, M., Hasselt, H. V., Lanctot, M., and Freitas, N. D · 2016
Cited alongside, same era.
The option-critic architecture
Bacon, P., Harb, J., and Precup, D · 2017
Cited alongside, same era.
Recurrent environment simulators
Chiappa, S., Racaniere, S., Wierstra, D., and Mohamed, S · 2017
Cited alongside, same era.
Vezhnevets, A. S., Osindero, S., Schaul, T., Heess, N., Jaderberg, M., Silver, D., and Kavukcuoglu, K · 2017
Later among the works it cites.
Playing hard exploration games by watching youtube
Aytar, Y., Pfaff, T., Budden, D., Paine, T. L., Wang, Z., and de Freitas, N · 2018
Later among the works it cites.
Exploration by random network distillation
Burda, Y., Edwards, H., Storkey, A., and Klimov, O · 2018
Later among the works it cites.
Contingency-aware exploration in reinforcement learning
Choi, J., Guo, Y., Moczulski, M., Oh, J., Wu, N., Norouzi, M., and Lee, H · 2018
Later among the works it cites.
Deep Q-learning from demonstrations
Hester, T., Vecerik, M., Pietquin, O., Lanctot, M., Schaul, T., Piot, B., Sendonaris, A., Dulac-Arnold, G., Osband, I., Agapiou, J., Leibo, J. Z., and Gruslys, A · 2018
Later among the works it cites.
Strategic object oriented reinforcement learning
Keramati, R., Whang, J., Cho, P., and Brunskill, E · 2018
Later among the works it cites.
Reinforcement learning on web interfaces using workflow-guided exploration
Liu, E. Z., Guu, K., Pasupat, P., Shi, T., and Liang, P · 2018
Later among the works it cites.
Neural network dynamics for model-based deep reinforcement learning with model-free fine-tuning
Nagabandi, A., Kahn, G., Fearing, R. S., and Levine, S · 2018
Later among the works it cites.
Self-imitation learning
Oh, J., Guo, Y., Singh, S., and Lee, H · 2018
Later among the works it cites.
Observe and look further: Achieving consistent performance on ATARI
Pohlen, T., Piot, B., Hester, T., Azar, M. G., Horgan, D., Budden, D., Barth-Maron, G., van Hasselt, H., Quan, J., Večer’ık, M., et al · 2018
Later among the works it cites.
Go-explore: a new approach for hard-exploration problems
Ecoffet, A., Huizinga, J., Lehman, J., Stanley, K. O., and Clune, J · 2019
Later among the works it cites.
Never give up: Learning directed exploration strategies
Badia, A. P., Sprechmann, P., Vitvitskyi, A., Guo, D., Piot, B., Kapturowski, S., Tieleman, O., Arjovsky, M., Pritzel, A., Bolt, A., et al · 2020
Closest in time.