Fetching the paper…
Reading the bibliography…
We present an algorithm for learning an approximate action-value soft Q-function in the relative entropy regularised reinforcement learning setting, for which an optimal improved policy can be recovered in closed form.
Reinforcement learning in continuous time: Advantage updating
L. C. Baird · 1994
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
R. S. Sutton, D. A. McAllester, S. P. Singh, and Y. Mansour · 2000
Earlier work this paper cites.
Natural actor-critic
J. Peters and S. Schaal · 2008
Earlier work this paper cites.
Relative entropy policy search
J. Peters, K. Mülling, and Y. Altun · 2010
Earlier work this paper cites.
Learning continuous control policies by stochastic value gradients
N. Heess, G. Wayne, D. Silver, T. Lillicrap, T. Erez, and Y. Tassa · 2015
Earlier work this paper cites.
Variational inference with normalizing flows
D. J. Rezende and S. Mohamed · 2015
Cited alongside, same era.
Trust region policy optimization
J. Schulman, S. Levine, P. Abbeel, M. Jordan, and P. Moritz · 2015
Cited alongside, same era.
Density estimation using real nvp
L. Dinh, J. Sohl-Dickstein, and S. Bengio · 2016
Cited alongside, same era.
Continuous deep q-learning with model-based acceleration
S. Gu, T. Lillicrap, I. Sutskever, and S. Levine · 2016
Cited alongside, same era.
Weight normalization: A simple reparameterization to accelerate training of deep neural networks
T. Salimans and D. P. Kingma · 2016
Cited alongside, same era.
Latent space policies for hierarchical reinforcement learning
T. Haarnoja, K. Hartikainen, P. Abbeel, and S. Levine
Cited in the paper.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
T. Haarnoja, A. Zhou, P. Abbeel, and S. Levine
Cited in the paper.
Reinforcement learning with deep energy-based policies
T. Haarnoja, H. Tang, P. Abbeel, and S. Levine · 2017
Later among the works it cites.
Maximum a posteriori policy optimisation
A. Abdolmaleki, J. T. Springenberg, Y. Tassa, R. Munos, N. Heess, and M. Riedmiller · 2018
Later among the works it cites.
Quinoa: a Q-function you infer normalized over actions
J. Degrave, A. Abdolmaleki, J. Springenberg, N. Heess, and M. Riedmiller · 2018
Later among the works it cites.
C.-W. Huang, D. Krueger, A. Lacoste, and A. Courville · 2018
Later among the works it cites.
Y. Tassa, Y. Doron, A. Muldal, T. Erez, Y. Li, D. d. L. Casas, D. Budden, A. Abdolmaleki, J. Merel, A. Lefrancq, et al · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Later among the works it cites.