Fetching the paper…
Reading the bibliography…
Off-Policy reinforcement learning (RL) is an important class of methods for many problem domains, such as robotics, where the cost of collecting data is high and on-policy methods are consequently intractable.
Optimization of computer simulation models with rare events
Rubinstein, R. Y · 1997
Earlier work this paper cites.
Reinforcement Learning: An Introduction
Sutton, R. S. and Barto, A. G · 1998
Earlier work this paper cites.
Policy Gradient Methods for Reinforcement Learning with Function Approximation
Sutton, R., McAllester, D., Singh, S. P., and Mansour, Y · 1999
Earlier work this paper cites.
Completely derandomized self-adaptation in evolution strategies
Hansen, N. and Ostermeier, A · 2001
Earlier work this paper cites.
Reinforcement learning in feedback control
Hafner, R. and Riedmiller, M · 2011
Earlier work this paper cites.
Playing Atari with Deep Reinforcement Learning
Mnih, V., Kavukcuoglu, K., Silver, D., Graves, A., Antonoglou, I., Wierstra, D., and Riedmiller, M. A · 2013
Earlier work this paper cites.
Continuous control with deep reinforcement learning: Deep Deterministic Policy Gradients (DDPG)
Lillicrap, T. P., Hunt, J. J., Pritzel, A., Heess, N., Erez, T., Tassa, Y., Silver, D., and Wierstra, D · 2015
Earlier work this paper cites.
OpenAI Gym
Brockman, G., Cheung, V., Pettersson, L., Schneider, J., Schulman, J., Tang, J., and Zaremba, W · 2016
Earlier work this paper cites.
Benchmarking Deep Reinforcement Learning for Continuous Control
Duan, Y., Chen, X., Houthooft, R., Schulman, J., and Abbeel, P · 2016
Earlier work this paper cites.
Asynchronous Methods for Deep Reinforcement Learning arXiv : 1602 . 01783v2 [ cs . LG ] 16 Jun 2016
Mnih, V., Mirza, M., Graves, A., Harley, T., Lillicrap, T. P., and Silver, D · 2016
Cited alongside, same era.
Model-free preference-based reinforcement learning
Wirth, C., Fürnkranz, J., and Neumann, G · 2016
Cited alongside, same era.
Andrychowicz, M., Wolski, F., Ray, A., Schneider, J., Fong, R., Welinder, P., McGrew, B., Tobin, J., Abbeel, P., and Zaremba, W · 2017
Cited alongside, same era.
Reinforcement Learning with Deep Energy-Based Policies
Haarnoja, T., Tang, H., Abbeel, P., and Levine, S · 2017
Cited alongside, same era.
Reproducibility of Benchmarked Deep Reinforcement Learning Tasks for Continuous Control
Islam, R., Henderson, P., Gomrokchi, M., and Precup, D · 2017
Cited alongside, same era.
Composable Deep Reinforcement Learning for Robotic Manipulation
Haarnoja, T., Pong, V., Zhou, A., Dalal, M., Abbeel, P., and Levine, S · 2018
Later among the works it cites.
Deep Reinforcement Learning That Matters
Henderson, P., Islam, R., Bachman, P., Pineau, J., Precup, D., and Meger, D · 2018
Later among the works it cites.
Scalable Deep Reinforcement Learning for Vision-Based Robotic Manipulation
Kalashnikov, D., Irpan, A., Pastor, P., Ibarz, J., Herzog, A., Jang, E., Quillen, D., Holly, E., Kalakrishnan, M., Vanhoucke, V., and Levine, S · 2018
Later among the works it cites.
Evolution-guided policy gradient in reinforcement learning
Khadka, S. and Tumer, K · 2018
Later among the works it cites.
Sim-to-Real Reinforcement Learning for Deformable Object Manipulation
Matas, J., James, S., and Davison, A. J · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Evolution Strategies as a Scalable Alternative to Reinforcement Learning
Salimans, T., Ho, J., Chen, X., and Sutskever, I · 2017
Cited alongside, same era.
Proximal Policy Optimization Algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Cited alongside, same era.
Pre-training with Non-expert Human Demonstration for Deep Reinforcement Learning
de la Cruz, G. V., Du, Y., and Taylor, M. E · 2018
Cited alongside, same era.
Addressing Function Approximation Error in Actor-Critic Methods
Fujimoto, S., van Hoof, H., and Meger, D · 2018
Cited alongside, same era.
Relative Entropy Regularized Policy Iteration
Abdolmaleki, A., Springenberg, J. T., Degrave, J., Bohez, S., Tassa, Y., Belov, D., Heess, N., and Riedmiller, M
Cited in the paper.
Maximum a posteriori policy optimisation
Abdolmaleki, A., Springenberg, J. T., Tassa, Y., Munos, R., Heess, N., and Riedmiller, M
Cited in the paper.
Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor
Haarnoja, T., Zhou, A., Abbeel, P., and Levine, S
Cited in the paper.
Silver, D., Lever, G., Heess, N., Degris, T., Wierstra, D., and Riedmiller, M · 2018
Later among the works it cites.
Reinforcement and Imitation Learning for Diverse Visuomotor Skills
Zhu, Y., Wang, Z., Merel, J., Rusu, A., Erez, T., Cabi, S., Tunyasuvunakool, S., Kramár, J., Hadsell, R., de Freitas, N., and Heess, N · 2018
Later among the works it cites.
CEM-RL: Combining evolutionary and gradient-based methods for policy search
Pourchot and Sigaud · 2019
Closest in time.