Fetching the paper…
Reading the bibliography…
We introduce Dynamic Planning Networks (DPN), a novel architecture for deep reinforcement learning, that combines model-based and model-free aspects for online planning.
Dyna, an integrated architecture for learning, planning, and reacting
Sutton, R. S · 1991
Earlier work this paper cites.
Efficient learning and planning within the dyna framework
Peng, J. and Williams, R. J · 1993
Earlier work this paper cites.
Intrinsically motivated reinforcement learning
Chentanez, N., Barto, A. G., and Singh, S. P · 2005
Earlier work this paper cites.
Efficient selectivity and backup operators in monte-carlo tree search
Coulom, R · 2006
Earlier work this paper cites.
Pilco: A model-based and data-efficient approach to policy search
Deisenroth, M. and Rasmussen, C. E · 2011
Earlier work this paper cites.
Lecture 6.5—RmsProp: Divide the gradient by a running average of its recent magnitude
Tieleman, T. and Hinton, G · 2012
Earlier work this paper cites.
Learning neural network policies with guided policy search under unknown dynamics
Levine, S. and Abbeel, P · 2014
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., et al · 2015
Earlier work this paper cites.
Action-conditional video prediction using deep networks in atari games
Oh, J., Guo, X., Lee, H., Lewis, R. L., and Singh, S. P · 2015
Earlier work this paper cites.
Schmidhuber, J · 2015
Earlier work this paper cites.
Trust region policy optimization
Schulman, J., Levine, S., Abbeel, P., Jordan, M., and Moritz, P · 2015
Earlier work this paper cites.
Adaptive computation time for recurrent neural networks
Graves, A · 2016
Cited alongside, same era.
Continuous deep q-learning with model-based acceleration
Gu, S., Lillicrap, T., Sutskever, I., and Levine, S · 2016
Cited alongside, same era.
Categorical reparameterization with gumbel-softmax
Jang, E., Gu, S., and Poole, B · 2016
Cited alongside, same era.
The concrete distribution: A continuous relaxation of discrete random variables
Maddison, C. J., Mnih, A., and Teh, Y. W · 2016
Cited alongside, same era.
Asynchronous methods for deep reinforcement learning
Mnih, V., Badia, A. P., Mirza, M., Graves, A., Lillicrap, T., Harley, T., Silver, D., and Kavukcuoglu, K · 2016
Cited alongside, same era.
Oh, J., Singh, S., and Lee, H · 2017
Later among the works it cites.
Learning model-based planning from scratch
Pascanu, R., Li, Y., Vinyals, O., Heess, N., Buesing, L., Racanière, S., Reichert, D. P., Weber, T., Wierstra, D., and Battaglia, P · 2017
Later among the works it cites.
Parameter space noise for exploration
Plappert, M., Houthooft, R., Dhariwal, P., Sidor, S., Chen, R. Y., Chen, X., Asfour, T., Abbeel, P., and Andrychowicz, M · 2017
Later among the works it cites.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Later among the works it cites.
Reinforcement learning: An introduction
Sutton, R. S. and Barto, A. G · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Value iteration networks
Tamar, A., Levine, S., Abbeel, P., WU, Y., and Thomas, G · 2016
Cited alongside, same era.
Strategic attentive writer for learning macro-actions
Vezhnevets, A., Mnih, V., Agapiou, J., Osindero, S., Graves, A., Vinyals, O., Kavukcuoglu, K., et al · 2016
Cited alongside, same era.
Recurrent environment simulators
Chiappa, S., Racaniere, S., Wierstra, D., and Mohamed, S · 2017
Cited alongside, same era.
Deep visual foresight for planning robot motion
Finn, C. and Levine, S · 2017
Cited alongside, same era.
Game engine learning from video
Guzdial, M., Li, B., and Riedl, M. O · 2017
Cited alongside, same era.
Model-based planning with discrete and continuous actions
Henaff, M., Whitney, W. F., and LeCun, Y · 2017
Cited alongside, same era.
Treeqn and atreec: Differentiable tree planning for deep reinforcement learning
Farquhar, G., Rocktäschel, T., Igl, M., and Whiteson, S
Cited in the paper.
Later among the works it cites.
Imagination-augmented agents for deep reinforcement learning
Weber, T., Racanière, S., Reichert, D. P., Buesing, L., Guez, A., Rezende, D. J., Badia, A. P., Vinyals, O., Heess, N., Li, Y., et al · 2017
Later among the works it cites.
Combined reinforcement learning via abstract representations
François-Lavet, V., Bengio, Y., Precup, D., and Pineau, J · 2018
Closest in time.
Learning to search with mctsnets
Guez, A., Weber, T., Antonoglou, I., Simonyan, K., Vinyals, O., Wierstra, D., Munos, R., and Silver, D · 2018
Closest in time.
Ha, D. and Schmidhuber, J · 2018
Closest in time.
Openai five benchmark: Results, Aug 2018
OpenAI · 2018
Closest in time.
Srinivas, A., Jabri, A., Abbeel, P., Levine, S., and Finn, C · 2018
Closest in time.