Fetching the paper…
Reading the bibliography…
Many continuous control tasks have bounded action spaces.
Simple Statistical Gradient-Following Algorithms for Connectionist Reinforcement Learning
Williams, R · 1992
Earlier work this paper cites.
The Continuum-Armed Bandit Problem
Agrawal, R · 1995
Earlier work this paper cites.
Policy Gradient Methods for Reinforcement Learning with Function Approximation
Sutton, R. S., Mcallester, D., Singh, S., and Mansour, Y · 1999
Earlier work this paper cites.
Variance Reduction Techniques for Gradient Estimates in Reinforcement Learning
Greensmith, E., Bartlett, P., and Baxter, J · 2004
Earlier work this paper cites.
Off-Policy Actor-Critic
Degris, T., White, M., and Sutton, R. S · 2012
Earlier work this paper cites.
Control of a Free-Falling Cat by Policy-Based Reinforcement Learning
Nakano, D., Maeda, S.-i., and Ishii, S · 2012
Earlier work this paper cites.
MuJoCo: A physics engine for model-based control
Todorov, E., Erez, T., and Tassa, Y · 2012
Earlier work this paper cites.
Lunar Lander : A Continous-Action Case Study for Policy-Gradient Actor-Critic Algorithms
Shariff, R. and Dick, T · 2013
Earlier work this paper cites.
Adam: a Method for Stochastic Optimization
Kingma, D. P. and Ba, J. L · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. a., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., Petersen, S., Beattie, C., Sadik, A., Antonoglou, I., King, H., Kumaran, D., Wierstra, D., Legg, S., and Hassabis, D · 2015
Cited alongside, same era.
Trust Region Policy Optimization
Schulman, J., Levine, S., Moritz, P., Jordan, M., and Abbeel, P · 2015
Cited alongside, same era.
OpenAI Gym
Brockman, G., Cheung, V., Pettersson, L., Schneider, J., Schulman, J., Tang, J., and Zaremba, W · 2016
Cited alongside, same era.
Benchmarking Deep Reinforcement Learning for Continuous Control
Duan, Y., Chen, X., Houthooft, R., Schulman, J., and Abbeel, P · 2016
Cited alongside, same era.
End-to-End Training of Deep Visuomotor Policies
Levine, S., Finn, C., Darrell, T., and Abbeel, P · 2016
Cited alongside, same era.
Asynchronous Methods for Deep Reinforcement Learning
Mnih, V., Badia, A. P., Mirza, M., Graves, A., Lillicrap, T. P., Harley, T., Silver, D., and Kavukcuoglu, K · 2016
Mean Actor Critic
Asadi, K., Allen, C., Roderick, M., Mohamed, A.-r., Konidaris, G., and Littman, M · 2017
Later among the works it cites.
Improving Stochastic Policy Gradients in Continuous Control with Deep Reinforcement Learning using the Beta Distribution
Chou, P.-W., Maturana, D., and Scherer, S · 2017
Later among the works it cites.
OpenAI Baselines
Dhariwal, P., Hesse, C., Klimov, O., Nichol, A., Plappert, M., Radford, A., Schulman, J., Sidor, S., and Wu, Y · 2017
Later among the works it cites.
Emergence of Locomotion Behaviours in Rich Environments
Heess, N., TB, D., Sriram, S., Lemmon, J., Merel, J., Wayne, G., Tassa, Y., Erez, T., Wang, Z., Eslami, S. M. A., Riedmiller, M., and Silver, D · 2017
Later among the works it cites.
Proximal Policy Optimization Algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Later among the works it cites.
Mastering the game of Go without human knowledge
Silver, D., Schrittwieser, J., Simonyan, K., Antonoglou, I., Huang, A., Guez, A., Hubert, T., Baker, L., Lai, M., Bolton, A., Chen, Y., Lillicrap, T., Hui, F., and Sifre, L · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
High-Dimensional Continuous Control Using Generalized Advantage Estimation
Schulman, J., Moritz, P., Levine, S., Jordan, M. I., and Abbeel, P · 2016
Cited alongside, same era.
Mastering the game of Go with deep neural networks and tree search
Silver, D., Huang, A., Maddison, C. J., Guez, A., Sifre, L., Driessche, G. V. D., Schrittwieser, J., Antonoglou, I., Panneershelvam, V., Lanctot, M., Dieleman, S., Grewe, D., Nham, J., Kalchbrenner, N., Sutskever, I., Lillicrap, T., Leach, M., and Kavukcuoglu, K · 2016
Cited alongside, same era.
Off-policy Neural Fitted Actor-Critic
Zimmer, M., Boniface, Y., and Dutech, A · 2016
Cited alongside, same era.
Q-Prop: Sample-Efficient Policy Gradient with an Off-Policy Critic
Gu, S., Lillicrap, T., Ghahramani, Z., Turner, R. E., and Levine, S
Cited in the paper.
Interpolated Policy Gradient : Merging On-Policy and Off-Policy Gradient Estimation for Deep
Gu, S., Lillicrap, T., Ghahramani, Z., Turner, R. E., Schölkopf, B., and Levine, S
Cited in the paper.
Later among the works it cites.
Expected Policy Gradients
Ciosek, K. and Whiteson, S · 2018
Closest in time.
Deep Reinforcement Learning that Matters
Henderson, P., Islam, R., Bachman, P., Pineau, J., Precup, D., and Meger, D · 2018
Closest in time.
DeepMind Control Suite
Tassa, Y., Doron, Y., Muldal, A., Erez, T., Li, Y., De, D., Casas, L., Budden, D., Abdolmaleki, A., Merel, J., Lefrancq, A., Lillicrap, T., and Riedmiller, M · 2018
Closest in time.