Fetching the paper…
Reading the bibliography…
Actor-critic methods, a type of model-free reinforcement learning (RL), have achieved state-of-the-art performances in many real-world domains in continuous control.
Dynamic Programming and Optimal Control
Bertsekas, D. P. 2000 · 2000
Earlier work this paper cites.
Convergence Results for Single-Step On-PolicyReinforcement-Learning Algorithms
Singh, S.; Jaakkola, T.; Littman, M. L.; and Szepesvári, C. 2000 · 2000
Earlier work this paper cites.
General duality between optimal control and estimation
Todorov, E. 2008 · 2008
Earlier work this paper cites.
Maximum Entropy Inverse Reinforcement Learning
Ziebart, B. D.; Maas, A. L.; Bagnell, J. A.; and Dey, A. K. 2008 · 2008
Earlier work this paper cites.
Robot Trajectory Optimization Using Approximate Inference
Toussaint, M. 2009 · 2009
Earlier work this paper cites.
Double Q-learning
Hasselt, H. V. 2010 · 2010
Earlier work this paper cites.
Modeling Purposeful Adaptive Behavior with the Principle of Maximum Causal Entropy
Ziebart, B. D. 2010 · 2010
Earlier work this paper cites.
MuJoCo: A physics engine for model-based control
Todorov, E.; Erez, T.; and Tassa, Y. 2012 · 2012
Cited alongside, same era.
Playing Atari with Deep Reinforcement Learning
Mnih, V.; Kavukcuoglu, K.; Silver, D.; Graves, A.; Antonoglou, I.; Wierstra, D.; and Riedmiller, M. A. 2013 · 2013
Cited alongside, same era.
On Stochastic Optimal Control and Reinforcement Learning by Approximate Inference (Extended Abstract)
Rawlik, K.; Toussaint, M.; and Vijayakumar, S. 2013 · 2013
Cited alongside, same era.
Human-level control through deep reinforcement learning
Mnih, V.; Kavukcuoglu, K.; Silver, D.; Rusu, A. A.; Veness, J.; Bellemare, M. G.; Graves, A.; Riedmiller, M.; Fidjeland, A. K.; Ostrovski, G.; Petersen, S.; Beattie, C.; Sadik, A.; Antonoglou, I.; King, H.; Kumaran, D.; Wierstra, D.; Legg, S.; and Hassabis, D. 2015 · 2015
Cited alongside, same era.
Trust Region Policy Optimization
Schulman, J.; Levine, S.; Moritz, P.; Jordan, M. I.; and Abbeel, P. 2015 · 2015
Continuous control with deep reinforcement learning
Lillicrap, T. P.; Hunt, J. J.; Pritzel, A.; Heess, N.; Erez, T.; Tassa, Y.; Silver, D.; and Wierstra, D. 2016 · 2016
Later among the works it cites.
Mastering the game of Go with deep neural networks and tree search
Silver, D.; Huang, A.; Maddison, C. J.; Guez, A.; Sifre, L.; van den Driessche, G.; Schrittwieser, J.; Antonoglou, I.; Panneershelvam, V.; Lanctot, M.; Dieleman, S.; Grewe, D.; Nham, J.; Kalchbrenner, N.; Sutskever, I.; Lillicrap, T.; Leach, M.; Kavukcuoglu, K.; Graepel, T.; and Hassabis, D. 2016 · 2016
Later among the works it cites.
Bridging the Gap Between Value and Policy Based Reinforcement Learning
Nachum, O.; Norouzi, M.; Xu, K.; and Schuurmans, D. 2017 · 2017
Later among the works it cites.
Equivalence Between Policy Gradients and Soft Q-Learning
Schulman, J.; Abbeel, P.; and Chen, X. 2017 · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
OpenAI Gym
Brockman, G.; Cheung, V.; Pettersson, L.; Schneider, J.; Schulman, J.; Tang, J.; and Zaremba, W. 2016 · 2016
Cited alongside, same era.
Deep Reinforcement Learning for Robotic Manipulation
Gu, S.; Holly, E.; Lillicrap, T. P.; and Levine, S. 2016 · 2016
Cited alongside, same era.
Composable Deep Reinforcement Learning for Robotic Manipulation
Haarnoja, T.; Pong, V.; Zhou, A.; Dalal, M.; Abbeel, P.; and Levine, S. 2018a
Cited in the paper.
Soft Actor-Critic Algorithms and Applications
Haarnoja, T.; Zhou, A.; Hartikainen, K.; Tucker, G.; Ha, S.; Tan, J.; Kumar, V.; Zhu, H.; Gupta, A.; Abbeel, P.; and Levine, S. 2018b
Cited in the paper.
Schulman, J.; Wolski, F.; Dhariwal, P.; Radford, A.; and Klimov, O. 2017 · 2017
Later among the works it cites.
Addressing Function Approximation Error in Actor-Critic Methods
Fujimoto, S.; van Hoof, H.; and Meger, D. 2018 · 2018
Later among the works it cites.
Better Exploration with Optimistic Actor-Critic
Ciosek, K.; Vuong, Q.; Loftin, R.; and Hofmann, K. 2019 · 2019
Later among the works it cites.