Fetching the paper…
Reading the bibliography…
This paper proposes a novel deep reinforcement learning (RL) architecture, called Value Prediction Network (VPN), which integrates model-free and model-based RL methods into a single neural network.
Integrated architectures for learning, planning, and reacting based on approximating dynamic programming
R. S. Sutton · 1990
Earlier work this paper cites.
Q-learning
C. J. Watkins and P. Dayan · 1992
Earlier work this paper cites.
Between MDPs and semi-MDPs: A framework for temporal abstraction in reinforcement learning
R. S. Sutton, D. Precup, and S. Singh · 1999
Earlier work this paper cites.
Temporal abstraction in reinforcement learning
D. Precup · 2000
Earlier work this paper cites.
Learning options in reinforcement learning
M. Stolle and D. Precup · 2002
Earlier work this paper cites.
Bandit based monte-carlo planning
L. Kocsis and C. Szepesvári · 2006
Earlier work this paper cites.
Dyna-style planning with linear function approximation and prioritized sweeping
R. S. Sutton, C. Szepesvári, A. Geramifard, and M. H. Bowling · 2008
Earlier work this paper cites.
Variational bayesian learning of nonlinear hidden state-space models for model predictive control
T. Raiko and M. Tornio · 2009
Earlier work this paper cites.
Multi-step Dyna planning for policy evaluation and control
H. Yao, S. Bhatnagar, D. Diao, R. S. Sutton, and C. Szepesvári · 2009
Earlier work this paper cites.
The arcade learning environment: An evaluation platform for general agents
M. G. Bellemare, Y. Naddaf, J. Veness, and M. Bowling · 2012
Earlier work this paper cites.
A survey of monte carlo tree search methods
C. B. Browne, E. Powley, D. Whitehouse, S. M. Lucas, P. I. Cowling, P. Rohlfshagen, S. Tavener, D. Perez, S. Samothrakis, and S. Colton · 2012
Earlier work this paper cites.
Temporal-difference search in computer go
D. Silver, R. S. Sutton, and M. Müller · 2012
Earlier work this paper cites.
Fast and accurate deep network learning by exponential linear units (ELUs)
D.-A. Clevert, T. Unterthiner, and S. Hochreiter · 2015
Earlier work this paper cites.
Deep recurrent q-learning for partially observable MDPs
M. Hausknecht and P. Stone · 2015
Earlier work this paper cites.
Learning continuous control policies by stochastic value gradients
N. Heess, G. Wayne, D. Silver, T. P. Lillicrap, Y. Tassa, and T. Erez · 2015
Cited alongside, same era.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2015
Cited alongside, same era.
DeepMPC: Learning deep latent features for model predictive control
I. Lenz, R. A. Knepper, and A. Saxena · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski, S. Petersen, C. Beattie, A. Sadik, I. Antonoglou, H. King, D. Kumaran, D. Wierstra, S. Legg, and D. Hassabis · 2015
Cited alongside, same era.
Action-conditional video prediction using deep networks in atari games
J. Oh, X. Guo, H. Lee, R. L. Lewis, and S. Singh · 2015
Cited alongside, same era.
Incentivizing exploration in reinforcement learning with deep predictive models
Asynchronous methods for deep reinforcement learning
V. Mnih, A. P. Badia, M. Mirza, A. Graves, T. P. Lillicrap, T. Harley, D. Silver, and K. Kavukcuoglu · 2016
Later among the works it cites.
Control of memory, active perception, and action in minecraft
J. Oh, V. Chockalingam, S. Singh, and H. Lee · 2016
Later among the works it cites.
Mastering the game of go with deep neural networks and tree search
D. Silver, A. Huang, C. J. Maddison, A. Guez, L. Sifre, G. van den Driessche, J. Schrittwieser, I. Antonoglou, V. Panneershelvam, M. Lanctot, S. Dieleman, D. Grewe, J. Nham, N. Kalchbrenner, I. Sutskever, T. Lillicrap, M. Leach, K. Kavukcuoglu, T. Graepel, and D. Hassabis · 2016
Later among the works it cites.
Value iteration networks
A. Tamar, S. Levine, P. Abbeel, Y. Wu, and G. Thomas · 2016
Later among the works it cites.
Strategic attentive writer for learning macro-actions
A. Vezhnevets, V. Mnih, S. Osindero, A. Graves, O. Vinyals, J. Agapiou, and K. Kavukcuoglu · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
B. C. Stadie, S. Levine, and P. Abbeel · 2015
Cited alongside, same era.
Tensorflow: Large-scale machine learning on heterogeneous distributed systems
M. Abadi, A. Agarwal, P. Barham, E. Brevdo, Z. Chen, C. Citro, G. S. Corrado, A. Davis, J. Dean, M. Devin, S. Ghemawat, I. J. Goodfellow, A. Harp, G. Irving, M. Isard, Y. Jia, R. Józefowicz, L. Kaiser, M. Kudlur, J. Levenberg, D. Mané, R. Monga, S. Moore, D. G. Murray, C. Olah, M. Schuster, J. Shlens, B. Steiner, I. Sutskever, K. Talwar, P. A. Tucker, V. Vanhoucke, V. Vasudevan, F. B. Viégas, O. Vinyals, P. Warden, M. Wattenberg, M. Wicke, Y. Yu, and X. Zheng · 2016
Cited alongside, same era.
Unsupervised learning for physical interaction through video prediction
C. Finn, I. J. Goodfellow, and S. Levine · 2016
Cited alongside, same era.
Continuous deep q-learning with model-based acceleration
S. Gu, T. P. Lillicrap, I. Sutskever, and S. Levine · 2016
Cited alongside, same era.
Deep learning for reward design to improve monte carlo tree search in atari games
X. Guo, S. P. Singh, R. L. Lewis, and H. Lee · 2016
Cited alongside, same era.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Cited alongside, same era.
N. Kalchbrenner, A. van den Oord, K. Simonyan, I. Danihelka, O. Vinyals, A. Graves, and K. Kavukcuoglu · 2016
Cited alongside, same era.
Z. Wang, T. Schaul, M. Hessel, H. van Hasselt, M. Lanctot, and N. de Freitas · 2016
Later among the works it cites.
Recurrent environment simulators
S. Chiappa, S. Racaniere, D. Wierstra, and S. Mohamed · 2017
Closest in time.
Deep visual foresight for planning robot motion
C. Finn and S. Levine · 2017
Closest in time.
Reinforcement learning with unsupervised auxiliary tasks
M. Jaderberg, V. Mnih, W. Czarnecki, T. Schaul, J. Z. Leibo, D. Silver, and K. Kavukcuoglu · 2017
Closest in time.
Dynamic action repetition for deep reinforcement learning
A. S. Lakshminarayanan, S. Sharma, and B. Ravindran · 2017
Closest in time.
Prediction and control with temporal segment models
N. Mishra, P. Abbeel, and I. Mordatch · 2017
Closest in time.
Neural map: Structured memory for deep reinforcement learning
E. Parisotto and R. Salakhutdinov · 2017
Closest in time.
The predictron: End-to-end learning and planning
D. Silver, H. van Hasselt, M. Hessel, T. Schaul, A. Guez, T. Harley, G. Dulac-Arnold, D. Reichert, N. Rabinowitz, A. Barreto, and T. Degris · 2017
Closest in time.