Fetching the paper…
Reading the bibliography…
This paper proposes a new optimization objective for value-based deep reinforcement learning.
Learning from delayed rewards
C J C H Watkins · 1989
Earlier work this paper cites.
Integrated architectures for learning, planning, and reacting based on approximating dynamic programming
R S Sutton · 1990
Earlier work this paper cites.
Reinforcement Learning: An Introduction
R S Sutton and A G Barto · 1998
Earlier work this paper cites.
An analysis of linear models, linear value-function approximation, and feature selection for reinforcement learning
R Parr, L Li, G Taylor, C Painter-Wakefield, and M L Littman · 2008
Earlier work this paper cites.
Sample-based learning and search with permanent and transient memories
D Silver, R S Sutton, and M Muller · 2008
Earlier work this paper cites.
Model Predictive Control System Design and Implementation using MATLAB
L Wang · 2009
Earlier work this paper cites.
Reinforcement Learning and Dynamic Programming using Function Approximators
L Busoniu, R Babuska, B De Schutter, and D Ernst · 2010
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
X Glorot and Y Bengio · 2010
Earlier work this paper cites.
PILCO: A model-based and data-efficient approach to policy search
M P Deisenroth and C E Rasmussen · 2011
Earlier work this paper cites.
Deep sparse rectifier neural networks
X Glorot, A Bordes, and Y Bengio · 2011
Earlier work this paper cites.
Dyna-style planning with linear function approximation and prioritized sweeping
R S Sutton, C Szepesvari, A Geramifard, and M P Bowling · 2012
Earlier work this paper cites.
A survey of Monte Carlo tree search methods
C Browne, E Powley, D Whitehouse, S Lucas, P I Cowling, P Rohlfshagen, S Tavener, D Perez, S Samothrakis, and S Colton · 2012
Earlier work this paper cites.
The Arcade Learning Environment: An evaluation platform for general agents
M G Bellemare, Y Naddaf, J Veness, and M Bowling · 2013
Earlier work this paper cites.
Guided policy search
S Levine and V Koltun · 2013
Cited alongside, same era.
Learning neural network policies with guided policy search under unknown dynamics
S Levine and P Abbeel · 2014
Cited alongside, same era.
Human-level control through deep reinforcement learning
V Mnih, K Kavukcuoglu, D Silver, A A Rusu, J Veness, M G Bellemare, A Graves, M Riedmiller, A K Fidjeland, G Ostrovski, S Petersen, C Beattie, A Sadik, I Antonoglou, H King, D Kumaran, D Wierstra, S Legg, and D Hassabis · 2015
Cited alongside, same era.
Learning continuous control policies by stochastic value gradients
N Heess, G Wayne, D Silver, T Lillicrap, Y Tassa, and T Erez · 2015
Cited alongside, same era.
From pixels to torques: Policy learning with deep dynamical models
N Wahlström, T B Schön, and M P Deisenroth · 2015
Cited alongside, same era.
Embed to control: A locally linear latent dynamics model for control from raw images
Model-based reinforcement learning for playing Atari games
J Fu and I Hsu · 2016
Later among the works it cites.
Artificial Intelligence: A Modern Approach
S J Russell and P Norvig · 2016
Later among the works it cites.
Recurrent environment simulators
S Chiappa, S Racaniere, D Wierstra, and S Mohamed · 2017
Later among the works it cites.
A deep learning approach for joint video frame and reward prediction in Atari games
F Leibfried, N Kushman, and K Hofmann · 2017
Later among the works it cites.
Deep action conditional neural network for frame prediction in Atari games
E Wang, A Kosson, and T Mu · 2017
Later among the works it cites.
Imagination-augmented agents for deep reinforcement learning
T Weber, S Racaniere, D Reichert, L Buesing, A Guez, D J Rezende, A Puigdomenech Badia, O Vinyals, N Heess, Y Li, R Pascanu, P Battaglia, D Silver, and D Wierstra · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
M Watter, J T Springenberg, J Boedecker, and M Riedmiller · 2015
Cited alongside, same era.
Action-conditional video prediction using deep networks in Atari games
J Oh, X Guo, H Lee, R Lewis, and S Singh · 2015
Cited alongside, same era.
Incentivizing exploration in reinforcement learning with deep predictive models
B C Stadie, S Levine, and P Abbeel · 2015
Cited alongside, same era.
Learning to generate chairs with convolutional neural networks
A Dosovitskiy, J T Springenberg, and T Brox · 2015
Cited alongside, same era.
Adam: A method for stochastic optimization
D P Kingma and J Ba · 2015
Cited alongside, same era.
Linear feature encoding for reinforcement learning
Z Song, R Parr, X Liao, and L Carin · 2016
Cited alongside, same era.
Deep spatial autoencoders for visuomotor learning
C Finn, X Y Tan, Y Duan, T Darrell, S Levine, and P Abbeel · 2016
Cited alongside, same era.
Later among the works it cites.
Curiosity-driven exploration by self-supervised prediction
D Pathak, P Agrawal, A A Efros, and T Darrell · 2017
Later among the works it cites.
Reinforcement learning with unsupervised auxiliary tasks
M Jaderberg, V Mnih, W M Czarnecki, T Schaul, J Z Leibo, D Silver, and K Kavukcuoglu · 2017
Later among the works it cites.
Deep reinforcement learning with model learning and Monte Carlo tree search in Minecraft
S Alaniz · 2017
Later among the works it cites.
Value prediction network
J Oh, S Singh, and H Lee · 2017
Later among the works it cites.
Temporal difference models: model-free deep rl for model-based control
V Pong, S Gu, M Dalal, and S Levine · 2018
Closest in time.
Learning and querying fast generative models for reinforcement rearning
L Buesing, T Weber, S Racaniere, S M Ali Eslami, D Rezende, D P Reichert, F Viola, F Besse, K Gregor, D Hassabis, and D Wierstra · 2018
Closest in time.
Rainbow: Combining improvements in deep reinforcement learning
M Hessel, J Modayil, H van Hasselt, T Schaul, G Ostrovski, W Dabney, D Horgan, B Piot, M Azar, and D Silver · 2018
Closest in time.