Fetching the paper…
Reading the bibliography…
Soft Actor-Critic is a state-of-the-art reinforcement learning algorithm for continuous action settings that is not applicable to discrete action settings.
“The Arcade Learning Environment: An Evaluation Platform for General Agents”
M. Bellemare, Y. Naddaf, J. Veness and M. Bowling · 2013
Earlier work this paper cites.
“Auto-Encoding Variational Bayes”
D. Kingma and M. Welling · 2013
Earlier work this paper cites.
“Human-Level Control Through Deep Reinforcement Learning”
V. Mnih, K. Kavukcuoglu, D. Silver, A. Rusu, J. Veness, M. Bellemare, A. Graves, M. Riedmiller, A. Fidjeland, G. Ostrovski, S. Peterson, C. Beattie, A. Sadik, I. Antonoglou, H. King, D. Kumaran, D. Wierstra, S. Legg and D. Hassabis · 2015
Earlier work this paper cites.
“Mastering the Game of Go Without Human Knowledge”
D. Silver, J. Schrittwieser, K. Simonyan, I. Antonoglou, A. Huang, A. Guez, T. Hubert, L. Baker, M. Lai, A. Bolton, Y. Chen, T. Lillicrap, F. Hui, L. Sifre, G. Driessche, T. Graepel and D. Hassabis · 2017
Earlier work this paper cites.
“Scalable Trust-Region Method for Deep Reinforcement Learning Using Kronecker-Factored Approximation”
Y. Wu, E. Mansimov, S. Liao, R. Grosse and J. Ba · 2017
Cited alongside, same era.
“Dopamine: A Research Framework for Deep Reinforcement Learning”
P. Castro, S. Moitra, C. Gelanda, S. Kumar and M. Bellemare · 2018
Cited alongside, same era.
“Addressing Function Approximation Error in Actor-Critic Methods”
S. Fujimoto, H. van Hoof and D. Meger · 2018
Cited alongside, same era.
“Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor”
T. Haarnoja, A. Zhou, P. Abbeel and S. Levine · 2018
Cited alongside, same era.
“Learning Dexterous In-Hand Manipulation”
OpenAI, Marcin Andrychowicz, Bowen Baker, Maciek Chociej, Rafal Józefowicz, Bob McGrew, Jakub. Pachocki, Arthur Petron, Matthias Plappert, Glenn Powell, Alex Ray, Jonas Schneider, Szymon Sidor, Josh Tobin, Peter Welinder, Lilian Weng and Wojciech Zaremba · 2018
Later among the works it cites.
“Soft Actor-Critic Algorithms and Applications”
T. Haarnoja, A. Zhou, K. Hartikainen, G. Tucker, J. S., V. Kumar, H. Zhu, A. Gupta, P. Abbeel and S. Levine · 2019
Closest in time.
“Model Based Reinforcement Learning for Atari”
L. Kaiser, M. Babaeizadech, P. Milos, B. Osinski, R. Campbell, K. Czechowski, D. Erhan, C. Finn, P. Kozakowski, S. Levine, A. Mohiuddin, R. Sepassi, G. Tucker and H. Michalewski · 2019
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…