Fetching the paper…
Reading the bibliography…
We examine the question of when and how parametric models are most useful in reinforcement learning.
Model predictive heuristic control
J. Richalet, A. Rault, J. Testud, and J. Papon · 1978
Earlier work this paper cites.
Learning to predict by the methods of temporal differences
R. S. Sutton · 1988
Earlier work this paper cites.
Learning from delayed rewards
C. J. C. H. Watkins · 1989
Earlier work this paper cites.
Integrated architectures for learning, planning, and reacting based on approximating dynamic programming
R. S. Sutton · 1990
Earlier work this paper cites.
Self-improving reactive agents based on reinforcement learning, planning and teaching
L. Lin · 1992
Earlier work this paper cites.
Q-learning
C. J. C. H. Watkins and P. Dayan · 1992
Earlier work this paper cites.
R. J. Williams and L. C. Baird III · 1993
Earlier work this paper cites.
Residual algorithms: Reinforcement learning with function approximation
L. Baird · 1995
Earlier work this paper cites.
On the virtues of linear learning and trajectory distributions
R. S. Sutton · 1995
Earlier work this paper cites.
Linear least-squares algorithms for temporal difference learning
S. J. Bradtke and A. G. Barto · 1996
Earlier work this paper cites.
An analysis of temporal-difference learning with function approximation
J. N. Tsitsiklis and B. Van Roy · 1997
Earlier work this paper cites.
Intra-Option Learning about Temporally Abstract Actions
R. S. Sutton, D. Precup, and S. P. Singh · 1998
Earlier work this paper cites.
Least-squares temporal difference learning
J. A. Boyan · 1999
Earlier work this paper cites.
Model predictive control: past, present and future
M. Morari and J. H. Lee · 1999
Earlier work this paper cites.
Eligibility traces for off-policy policy evaluation
D. Precup, R. S. Sutton, and S. P. Singh · 2000
Earlier work this paper cites.
Neural fitted Q iteration - first experiences with a data efficient neural reinforcement learning method
M. Riedmiller · 2005
Cited alongside, same era.
An analysis of linear models, linear value-function approximation, and feature selection for reinforcement learning
R. Parr, L. Li, G. Taylor, C. Painter-Wakefield, and M. L. Littman · 2008
Cited alongside, same era.
A convergent O(n) algorithm for off-policy temporal-difference learning with linear function approximation
R. S. Sutton, C. Szepesvári, and H. R. Maei · 2008
Cited alongside, same era.
Fast gradient-descent methods for temporal-difference learning with linear function approximation
R. S. Sutton, H. R. Maei, D. Precup, S. Bhatnagar, D. Silver, C. Szepesvári, and E. Wiewiora · 2009
Cited alongside, same era.
Hierarchical planning in the now
L. P. Kaelbling and T. Lozano-Pérez · 2010
Cited alongside, same era.
An emphatic approach to the problem of off-policy temporal-difference learning
R. S. Sutton, A. R. Mahmood, and M. White · 2016
Later among the works it cites.
Deep reinforcement learning with Double Q-learning
H. van Hasselt, A. Guez, and D. Silver · 2016
Later among the works it cites.
Dueling network architectures for deep reinforcement learning
Z. Wang, N. de Freitas, T. Schaul, M. Hessel, H. van Hasselt, and M. Lanctot · 2016
Later among the works it cites.
A distributional perspective on reinforcement learning
M. G. Bellemare, W. Dabney, and R. Munos · 2017
Later among the works it cites.
Noisy networks for exploration
M. Fortunato, M. G. Azar, B. Piot, J. Menick, I. Osband, A. Graves, V. Mnih, R. Munos, D. Hassabis, O. Pietquin, C. Blundell, and S. Legg · 2017
Later among the works it cites.
The predictron: End-to-end learning and planning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
H. van Hasselt · 2010
Cited alongside, same era.
The arcade learning environment: An evaluation platform for general agents
M. G. Bellemare, Y. Naddaf, J. Veness, and M. Bowling · 2013
Cited alongside, same era.
Playing atari with deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. Graves, I. Antonoglou, D. Wierstra, and M. Riedmiller · 2013
Cited alongside, same era.
Model predictive control: Recent developments and future promise
D. Q. Mayne · 2014
Cited alongside, same era.
Off-policy TD(
H. van Hasselt, A. R. Mahmood, and R. S. Sutton · 2014
Cited alongside, same era.
Human-level control through deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski, S. Petersen, C. Beattie, A. Sadik, I. Antonoglou, H. King, D. Kumaran, D. Wierstra, S. Legg, and D. Hassabis · 2015
Cited alongside, same era.
Learning to predict independent of span
H. van Hasselt and R. S. Sutton · 2015
Cited alongside, same era.
D. Silver, H. van Hasselt, M. Hessel, T. Schaul, A. Guez, T. Harley, G. Dulac-Arnold, D. Reichert, N. Rabinowitz, A. Barreto, and T. Degris · 2017
Later among the works it cites.
Imagination-augmented agents for deep reinforcement learning
T. Weber, S. Racanière, D. P. Reichert, L. Buesing, A. Guez, D. J. Rezende, A. P. Badia, O. Vinyals, N. Heess, Y. Li, R. Pascanu, P. Battaglia, D. Silver, and D. Wierstra · 2017
Later among the works it cites.
Multiple-step greedy policies in approximate and online reinforcement learning
Y. Efroni, G. Dalal, B. Scherrer, and S. Mannor · 2018
Later among the works it cites.
Organizing experience: a deeper look at replay mechanisms for sample-based planning in continuous state domains
Y. Pan, M. Zaheer, A. White, A. Patterson, and M. White · 2018
Later among the works it cites.
Reinforcement Learning: An Introduction
R. S. Sutton and A. G. Barto · 2018
Later among the works it cites.
Deep reinforcement learning and the deadly triad
H. van Hasselt, Y. Doron, F. Strub, M. Hessel, N. Sonnerat, and J. Modayil · 2018
Later among the works it cites.
Model-based reinforcement learning for atari
L. Kaiser, M. Babaeizadeh, P. Milos, B. Osinski, R. H. Campbell, K. Czechowski, D. Erhan, C. Finn, P. Kozakowski, S. Levine, R. Sepassi, G. Tucker, and H. Michalewski · 2019
Closest in time.
Plan online, learn offline: Efficient learning and exploration via model-based control
K. Lowrey, A. Rajeswaran, S. Kakade, E. Todorov, and I. Mordatch · 2019
Closest in time.
An online learning approach to model predictive control
N. Wagener, C.-A. Cheng, J. Sacks, and B. Boots · 2019
Closest in time.
Planning with expectation models
Y. Wan, M. Zaheer, A. White, M. White, and R. S. Sutton · 2019
Closest in time.