Fetching the paper…
Reading the bibliography…
Model-free reinforcement learning has been successfully applied to a range of challenging problems, and has recently been extended to handle large neural network policies and value functions.
Integrated architectures for learning, planning, and reacting based on approximating dynamic programming
Sutton, Richard S · 1990
Earlier work this paper cites.
Advantage updating
Baird III, Leemon C · 1993
Earlier work this paper cites.
Multi-player residual advantage learning with general function approximation
Harmon, Mance E and Baird III, Leemon C · 1996
Earlier work this paper cites.
Locally weighted learning for control
Atkeson, Christopher G, Moore, Andrew W, and Schaal, Stefan · 1997
Earlier work this paper cites.
Actor-critic algorithms
Konda, Vijay R and Tsitsiklis, John N · 1999
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Sutton, Richard S, McAllester, David A, Singh, Satinder P, Mansour, Yishay, et al · 1999
Earlier work this paper cites.
Iterative linear quadratic regulator design for nonlinear biological movement systems
Li, Weiwei and Todorov, Emanuel · 2004
Earlier work this paper cites.
Policy gradient methods for robotics
Peters, Jan and Schaal, Stefan · 2006
Earlier work this paper cites.
Relative entropy policy search
Peters, Jan, Mülling, Katharina, and Altun, Yasemin · 2010
Earlier work this paper cites.
Pilco: A model-based and data-efficient approach to policy search
Deisenroth, Marc and Rasmussen, Carl E · 2011
Earlier work this paper cites.
Reinforcement learning in feedback control
Hafner, Roland and Riedmiller, Martin · 2011
Earlier work this paper cites.
Reinforcement learning in robotics: A survey
Kober, Jens and Peters, Jan · 2012
Cited alongside, same era.
Synthesis and stabilization of complex behaviors through online trajectory optimization
Tassa, Yuval, Erez, Tom, and Todorov, Emanuel · 2012
Cited alongside, same era.
Mujoco: A physics engine for model-based control
Todorov, Emanuel, Erez, Tom, and Tassa, Yuval · 2012
Cited alongside, same era.
A survey on policy search for robotics
Deisenroth, Marc Peter, Neumann, Gerhard, Peters, Jan, et al · 2013
Cited alongside, same era.
Guided policy search
Levine, Sergey and Koltun, Vladlen · 2013
Cited alongside, same era.
On stochastic optimal control and reinforcement learning by approximate inference
Rawlik, Konrad, Toussaint, Marc, and Vijayakumar, Sethu · 2013
Cited alongside, same era.
Deep reinforcement learning in parameterized action space
Hausknecht, Matthew and Stone, Peter · 2015
Later among the works it cites.
Learning continuous control policies by stochastic value gradients
Heess, Nicolas, Wayne, Gregory, Silver, David, Lillicrap, Tim, Erez, Tom, and Tassa, Yuval · 2015
Later among the works it cites.
End-to-end training of deep visuomotor policies
Levine, Sergey, Finn, Chelsea, Darrell, Trevor, and Abbeel, Pieter · 2015
Later among the works it cites.
Human-level control through deep reinforcement learning
Mnih, Volodymyr, Kavukcuoglu, Koray, Silver, David, Rusu, Andrei A, Veness, Joel, Bellemare, Marc G, Graves, Alex, Riedmiller, Martin, Fidjeland, Andreas K, Ostrovski, Georg, et al · 2015
Later among the works it cites.
Schaul, Tom, Quan, John, Antonoglou, Ioannis, and Silver, David · 2015
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Kingma, Diederik and Ba, Jimmy · 2014
Cited alongside, same era.
Learning neural network policies with guided policy search under unknown dynamics
Levine, Sergey and Abbeel, Pieter · 2014
Cited alongside, same era.
Deterministic policy gradient algorithms
Silver, David, Lever, Guy, Heess, Nicolas, Degris, Thomas, Wierstra, Daan, and Riedmiller, Martin · 2014
Cited alongside, same era.
The importance of experience replay database composition in deep reinforcement learning
de Bruin, Tim, Kober, Jens, Tuyls, Karl, and Babuška, Robert · 2015
Cited alongside, same era.
One-shot learning of manipulation skills with online dynamics adaptation and neural network priors
Fu, Justin, Levine, Sergey, and Abbeel, Pieter · 2015
Cited alongside, same era.
Later among the works it cites.
Trust region policy optimization
Schulman, John, Levine, Sergey, Abbeel, Pieter, Jordan, Michael I., and Moritz, Philipp · 2015
Later among the works it cites.
From pixels to torques: Policy learning with deep dynamical models
Wahlström, Niklas, Schön, Thomas B, and Deisenroth, Marc Peter · 2015
Later among the works it cites.
Dueling network architectures for deep reinforcement learning
Wang, Ziyu, de Freitas, Nando, and Lanctot, Marc · 2015
Later among the works it cites.
Embed to control: A locally linear latent dynamics model for control from raw images
Watter, Manuel, Springenberg, Jost, Boedecker, Joschka, and Riedmiller, Martin · 2015
Later among the works it cites.
Continuous control with deep reinforcement learning
Lillicrap, Timothy P, Hunt, Jonathan J, Pritzel, Alexander, Heess, Nicolas, Erez, Tom, Tassa, Yuval, Silver, David, and Wierstra, Daan · 2016
Closest in time.
High-dimensional continuous control using generalized advantage estimation
Schulman, John, Moritz, Philipp, Levine, Sergey, Jordan, Michael, and Abbeel, Pieter · 2016
Closest in time.