Fetching the paper…
Reading the bibliography…
Sample complexity and safety are major challenges when learning policies with reinforcement learning for real-world tasks, especially when the policies are represented using rich function approximators like deep neural networks.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J. Williams · 1992
Earlier work this paper cites.
Robust and Optimal Control
Kemin Zhou, John C. Doyle, and Keith Glover · 1996
Earlier work this paper cites.
System Identification , pp. 163–173
Lennart Ljung · 1998
Earlier work this paper cites.
A natural policy gradient
Sham Kakade · 2001
Earlier work this paper cites.
Approximately optimal approximate reinforcement learning
Sham Kakade and John Langford · 2002
Earlier work this paper cites.
Design for an optimal probe
Michael O. Duff · 2003
Earlier work this paper cites.
On the Sample Complexity of Reinforcement Learning
Sham Kakade · 2003
Earlier work this paper cites.
Robust control of markov decision processes with uncertain transition matrices
Arnab Nilim and Laurent El Ghaoui · 2005
Earlier work this paper cites.
Using inaccurate models in reinforcement learning
Pieter Abbeel, Morgan Quigley, and Andrew Y. Ng · 2006
Earlier work this paper cites.
Point-based value iteration for continuous pomdps
Josep M. Porta, Nikos A. Vlassis, Matthijs T. J. Spaan, and Pascal Poupart · 2006
Earlier work this paper cites.
An analytic solution to discrete bayesian reinforcement learning
Pascal Poupart, Nikos A. Vlassis, Jesse Hoey, and Kevin Regan · 2006
Earlier work this paper cites.
Bayesian reinforcement learning in continuous pomdps with application to robot navigation
S. Ross, B. Chaib-draa, and J. Pineau · 2008
Earlier work this paper cites.
A survey of robot learning from demonstration
Brenna D. Argall, Sonia Chernova, Manuela Veloso, and Brett Browning · 2009
Earlier work this paper cites.
Transfer learning for reinforcement learning domains: A survey
Matthew E. Taylor and Peter Stone · 2009
Cited alongside, same era.
Real-time reinforcement learning by sequential actor-critics and experience replay
Pawel Wawrzynski · 2009
Cited alongside, same era.
Percentile optimization for markov decision processes with parameter uncertainty
Erick Delage and Shie Mannor · 2010
Cited alongside, same era.
Optimizing walking controllers for uncertain inputs and environments
Jack M. Wang, David J. Fleet, and Aaron Hertzmann · 2010
Cited alongside, same era.
Infinite-horizon model predictive control for periodic tasks with contacts
Tom Erez, Yuval Tassa, and Emanuel Todorov · 2011
Cited alongside, same era.
Learning parameterized skills
Bruno Castro da Silva, George Konidaris, and Andrew G. Barto · 2012
Cited alongside, same era.
Reinforcement learning in robust markov decision processes
Shiau Hong Lim, Huan Xu, and Shie Mannor · 2013
Later among the works it cites.
A comprehensive survey on safe reinforcement learning
Javier García and Fernando Fernández · 2015
Later among the works it cites.
Bayesian reinforcement learning: A survey
Mohammad Ghavamzadeh, Shie Mannor, Joelle Pineau, and Aviv Tamar · 2015
Later among the works it cites.
Continuous control with deep reinforcement learning
T. P. Lillicrap, J. J. Hunt, A. Pritzel, N. Heess, T. Erez, Y. Tassa, D. Silver, and D. Wierstra · 2015
Later among the works it cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih et al · 2015
Later among the works it cites.
Trust region policy optimization
John Schulman, Sergey Levine, Philipp Moritz, Michael Jordan, and Pieter Abbeel · 2015
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Agnostic system identification for model-based reinforcement learning
Stephane Ross and Drew Bagnell · 2012
Cited alongside, same era.
Integrating a partial model into model free reinforcement learning
Aviv Tamar, Dotan Di Castro, and Ron Meir · 2012
Cited alongside, same era.
Mujoco: A physics engine for model-based control
E. Todorov, T. Erez, and Y. Tassa · 2012
Cited alongside, same era.
Bayesian Reinforcement Learning , pp. 359–386
Nikos Vlassis, Mohammad Ghavamzadeh, Shie Mannor, and Pascal Poupart · 2012
Cited alongside, same era.
A survey on policy search for robotics
Marc Peter Deisenroth, Gerhard Neumann, and Jan Peters · 2013
Cited alongside, same era.
Guided policy search
Sergey Levine and Vladlen Koltun · 2013
Cited alongside, same era.
Optimizing the cvar via sampling
Aviv Tamar, Yonatan Glassner, and Shie Mannor · 2015
Later among the works it cites.
High-confidence off-policy evaluation
Philip Thomas, Georgios Theocharous, and Mohammad Ghavamzadeh · 2015
Later among the works it cites.
OpenAI Gym, 2016
Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wojciech Zaremba · 2016
Closest in time.
Benchmarking deep reinforcement learning for continuous control
Yan Duan, Xi Chen, Rein Houthooft, John Schulman, and Pieter Abbeel · 2016
Closest in time.
Terrain-adaptive locomotion skills using deep reinforcement learning
Xue Bin Peng, Glen Berseth, and Michiel van de Panne · 2016
Closest in time.
Mastering the game of go with deep neural networks and tree search
David Silver et al · 2016
Closest in time.