Fetching the paper…
Reading the bibliography…
We revisit residual algorithms in both model-free and model-based reinforcement learning settings.
Dynamic programming
Richard E. Bellman. 1957 · 1957
Earlier work this paper cites.
Robust estimation of a location parameter
Peter J Huber et al · 1964
Earlier work this paper cites.
Learning to predict by the methods of temporal differences
Richard S Sutton. 1988 · 1988
Earlier work this paper cites.
Model predictive control: theory and practice—a survey
Carlos E Garcia, David M Prett, and Manfred Morari. 1989 · 1989
Earlier work this paper cites.
Integrated architectures for learning, planning, and reacting based on approximating dynamic programming. In Proceedings of the 7th International Conference on Machine Learning
Richard S Sutton. 1990 · 1990
Earlier work this paper cites.
Self-improving reactive agents based on reinforcement learning, planning and teaching
Long-Ji Lin. 1992 · 1992
Earlier work this paper cites.
Q-learning
Christopher JCH Watkins and Peter Dayan. 1992 · 1992
Earlier work this paper cites.
Tight performance bounds on greedy policies based on imperfect value functions
Ronald J Williams and Leemon C Baird. 1993 · 1993
Earlier work this paper cites.
Residual algorithms: Reinforcement learning with function approximation
Leemon Baird. 1995 · 1995
Earlier work this paper cites.
Stable function approximation in dynamic programming
Geoffrey J Gordon. 1995 · 1995
Earlier work this paper cites.
Neuro-Dynamic Programming
Dimitri P Bertsekas and John N Tsitsiklis. 1996 · 1996
Earlier work this paper cites.
Analysis of temporal-diffference learning with function approximation. In Advances in Neural Information Pprocessing Systems
John N Tsitsiklis and Benjamin Van Roy. 1997 · 1997
Earlier work this paper cites.
Approximate solutions to Markov decision processes
Geoffrey J Gordon. 1999 · 1999
Earlier work this paper cites.
Between MDPs and semi-MDPs: A framework for temporal abstraction in reinforcement learning
Richard S Sutton, Doina Precup, and Satinder Singh. 1999 · 1999
Earlier work this paper cites.
Least-squares policy iteration
Michail G Lagoudakis and Ronald Parr. 2003 · 2003
Earlier work this paper cites.
Error bounds for approximate policy iteration. In Proceedings of the 20th International Conference on Machine Learning
Rémi Munos. 2003 · 2003
Earlier work this paper cites.
Learning near-optimal policies with Bellman-residual minimization based fitted policy iteration and a single sample path
András Antos, Csaba Szepesvári, and Rémi Munos. 2008 · 2008
Earlier work this paper cites.
A worst-case comparison between temporal difference and residual gradient with linear function approximation. In Proceedings of the 25th International Conference on Machine Learning
Lihong Li. 2008 · 2008
Earlier work this paper cites.
Fast gradient-descent methods for temporal-difference learning with linear function approximation. In Proceedings of the 26th International Conference on Machine Learning
Richard S Sutton, Hamid Reza Maei, Doina Precup, Shalabh Bhatnagar, David Silver, Csaba Szepesvári, and Eric Wiewiora. 2009 · 2009
Earlier work this paper cites.
Toward off-policy learning control with function approximation.. In Proceedings of the 27th International Conference on Machine Learning
Hamid Reza Maei, Csaba Szepesvári, Shalabh Bhatnagar, and Richard S Sutton. 2010 · 2010
Cited alongside, same era.
Bruno Scherrer. 2010 · 2010
Cited alongside, same era.
Gradient temporal-difference learning algorithms
Hamid Reza Maei. 2011 · 2011
Cited alongside, same era.
Insights in reinforcement learning
Hado Philip van Hasselt. 2011 · 2011
Cited alongside, same era.
The arcade learning environment: An evaluation platform for general agents
Marc G Bellemare, Yavar Naddaf, Joel Veness, and Michael Bowling. 2013 · 2013
Cited alongside, same era.
The predictron: End-to-end learning and planning. In Proceedings of the 34th International Conference on Machine Learning
David Silver, Hado van Hasselt, Matteo Hessel, Tom Schaul, Arthur Guez, Tim Harley, Gabriel Dulac-Arnold, David Reichert, Neil Rabinowitz, Andre Barreto, et al · 2017
Later among the works it cites.
Self-correcting models for model-based reinforcement learning. In Proceedings of the 31st AAAI Conference on Artificial Intelligence
Erik Talvitie. 2017 · 2017
Later among the works it cites.
Imagination-augmented agents for deep reinforcement learning
Théophane Weber, Sébastien Racanière, David P Reichert, Lars Buesing, Arthur Guez, Danilo Jimenez Rezende, Adria Puigdomènech Badia, Oriol Vinyals, Nicolas Heess, Yujia Li, et al · 2017
Later among the works it cites.
Sample-efficient reinforcement learning with stochastic ensemble value expansion. In Advances in Neural Information Processing Systems
Jacob Buckman, Danijar Hafner, George Tucker, Eugene Brevdo, and Honglak Lee. 2018 · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Natural temporal difference learning. In Proceedings of the 28th AAAI Conference on Artificial Intelligence
William Dabney and Philip Thomas. 2014 · 2014
Cited alongside, same era.
Policy evaluation with temporal differences: A survey and comparison
Christoph Dann, Gerhard Neumann, and Jan Peters. 2014 · 2014
Cited alongside, same era.
Markov decision processes: discrete stochastic dynamic programming
Martin L Puterman. 2014 · 2014
Cited alongside, same era.
Continuous control with deep reinforcement learning
Timothy P Lillicrap, Jonathan J Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra. 2015 · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Cited alongside, same era.
Online Bellman Residual algorithms with predictive error guarantees. In Proceedings of the 31st Conference on Uncertainty in Artificial Intelligence
Wen Sun and J Andrew Bagnell. 2015 · 2015
Cited alongside, same era.
Continuous deep q-learning with model-based acceleration. In Proceedings of the 33rd International Conference on Machine Learning
Shixiang Gu, Timothy Lillicrap, Ilya Sutskever, and Sergey Levine. 2016 · 2016
Cited alongside, same era.
Deep reinforcement learning in a handful of trials using probabilistic dynamics models. In Advances in Neural Information Processing Systems
Kurtland Chua, Roberto Calandra, Rowan McAllister, and Sergey Levine. 2018 · 2018
Later among the works it cites.
Model-based reinforcement learning via meta-policy optimization
Ignasi Clavera, Jonas Rothfuss, John Schulman, Yasuhiro Fujita, Tamim Asfour, and Pieter Abbeel. 2018 · 2018
Later among the works it cites.
TreeQN and ATreeC: Differentiable tree-structured models for deep reinforcement learning
Gregory Farquhar, Tim Rocktäschel, Maximilian Igl, and Shimon Whiteson. 2018 · 2018
Later among the works it cites.
Model-based value estimation for efficient model-free reinforcement learning
Vladimir Feinberg, Alvin Wan, Ion Stoica, Michael I Jordan, Joseph E Gonzalez, and Sergey Levine. 2018 · 2018
Later among the works it cites.
Addressing function approximation error in actor-critic methods
Scott Fujimoto, Herke van Hoof, and David Meger. 2018 · 2018
Later among the works it cites.
David Ha and Jürgen Schmidhuber. 2018 · 2018
Later among the works it cites.
Learning latent dynamics for planning from pixels
Danijar Hafner, Timothy Lillicrap, Ian Fischer, Ruben Villegas, David Ha, Honglak Lee, and James Davidson. 2018 · 2018
Later among the works it cites.
Model-ensemble trust-region policy optimization
Thanard Kurutach, Ignasi Clavera, Yan Duan, Aviv Tamar, and Pieter Abbeel. 2018 · 2018
Later among the works it cites.
Neural network dynamics for model-based deep reinforcement learning with model-free fine-tuning. In Proceedings of the 2018 International Conference on Robotics and Automation
Anusha Nagabandi, Gregory Kahn, Ronald S Fearing, and Sergey Levine. 2018 · 2018
Later among the works it cites.
Aravind Srinivas, Allan Jabri, Pieter Abbeel, Sergey Levine, and Chelsea Finn. 2018 · 2018
Later among the works it cites.
Reinforcement Learning: An Introduction (2nd Edition)
Richard S Sutton and Andrew G Barto. 2018 · 2018
Later among the works it cites.
Yuval Tassa, Yotam Doron, Alistair Muldal, Tom Erez, Yazhe Li, Diego de Las Casas, David Budden, Abbas Abdolmaleki, Josh Merel, Andrew Lefrancq, et al · 2018
Later among the works it cites.
A Kernel Loss for Solving the Bellman Equation
Yihao Feng, Lihong Li, and Qiang Liu. 2019 · 2019
Closest in time.
ACE: An Actor Ensemble Algorithm for Continuous Control with Tree Search
Shangtong Zhang, Hao Chen, and Hengshuai Yao. 2019 · 2019
Closest in time.