R. J. Williams, “Simple statistical gradient-following algorithms for connectionist reinforcement learning,” Machine learning
1992
Earlier work this paper cites.
C. M. Bishop, “Mixture density networks,” 1994
1994
Earlier work this paper cites.
R. S. Sutton and A. G. Barto, “Reinforcement learning: An introduction,” 1998
1998
Earlier work this paper cites.
R. S. Sutton, D. A. McAllester, S. P. Singh, and Y. Mansour, “Policy gradient methods for reinforcement learning with function approximation,” in Advances in neural information processing systems
2000
Earlier work this paper cites.
V. R. Konda and J. N. Tsitsiklis, “Actor-critic algorithms,” in Advances in neural information processing systems
2000
Earlier work this paper cites.
R. I. Brafman and M. Tennenholtz, “R-max-a general polynomial time algorithm for near-optimal reinforcement learning,” Journal of Machine Learning Research
2002
Earlier work this paper cites.
M. Kuss and C. E. Rasmussen, “Gaussian processes in reinforcement learning,” in Advances in neural information processing systems
2004
Earlier work this paper cites.
Y. Engel, S. Mannor, and R. Meir, “Reinforcement learning with gaussian processes,” in Proceedings of the 22nd international conference on Machine learning
2005
Earlier work this paper cites.
PhD thesis, echnische Universität Darmstadt Darmstadt, Germany, 2006
M. Kuss, Gaussian process models for robust regression, classification, and reinforcement learning · 2006
Earlier work this paper cites.
J. Peters and S. Schaal, “Natural actor-critic,” Neurocomputing
2008
Earlier work this paper cites.
T. Jaksch, R. Ortner, and P. Auer, “Near-optimal regret bounds for reinforcement learning,” Journal of Machine Learning Research
2010
Earlier work this paper cites.
S. Levine, Z. Popovic, and V. Koltun, “Nonlinear inverse reinforcement learning with gaussian processes,” in Advances in Neural Information Processing Systems
2011
Earlier work this paper cites.
T. Degris, M. White, and R. S. Sutton, “Off-policy actor-critic,” arXiv preprint arXiv:1205.4839
Original
2012
Earlier work this paper cites.
E. Todorov, T. Erez, and Y. Tassa, “Mujoco: A physics engine for model-based control.,” in IROS
2012
Earlier work this paper cites.
D. Silver, G. Lever, N. Heess, T. Degris, D. Wierstra, and M. Riedmiller, “Deterministic policy gradient algorithms,” in ICML
2014
Earlier work this paper cites.
R. Grande, T. Walsh, and J. How, “Sample efficient reinforcement learning with gaussian processes,” in International Conference on Machine Learning
2014
Earlier work this paper cites.
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski, et al
2015
Earlier work this paper cites.
T. P. Lillicrap, J. J. Hunt, A. Pritzel, N. Heess, T. Erez, Y. Tassa, D. Silver, and D. Wierstra, “Continuous control with deep reinforcement learning,” arXiv preprint arXiv:1509.02971
Original
2015
Earlier work this paper cites.
J. Schulman, S. Levine, P. Abbeel, M. Jordan, and P. Moritz, “Trust region policy optimization,” in International conference on machine learning
2015
Earlier work this paper cites.