Fetching the paper…
Reading the bibliography…
Deep Reinforcement Learning (DRL) algorithms for continuous action spaces are known to be brittle toward hyperparameters as well as \cut{being}sample inefficient.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Williams, R. J · 1992
Earlier work this paper cites.
Reinforcement learning - an introduction
Sutton, R. S. and Barto, A. G · 1998
Earlier work this paper cites.
Maximum entropy inverse reinforcement learning
Ziebart, B. D., Maas, A. L., Bagnell, J. A., and Dey, A. K · 2008
Earlier work this paper cites.
Modeling purposeful adaptive behavior with the principle of maximum causal entropy
Ziebart, B. D · 2010
Earlier work this paper cites.
Auto-encoding variational bayes
Kingma, D. P. and Welling, M · 2013
Earlier work this paper cites.
Stochastic backpropagation and approximate inference in deep generative models
Rezende, D. J., Mohamed, S., and Wierstra, D · 2014
Earlier work this paper cites.
Learning continuous control policies by stochastic value gradients
Heess, N., Wayne, G., Silver, D., Lillicrap, T., Erez, T., and Tassa, Y · 2015
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Ioffe, S. and Szegedy, C · 2015
Cited alongside, same era.
Gradient estimation using stochastic computation graphs
John Schulman, Nicolas Heess, T. W. P. A · 2015
Cited alongside, same era.
Variational inference with normalizing flows
Rezende, D. J. and Mohamed, S · 2015
Cited alongside, same era.
Gradient estimation using stochastic computation graphs
Schulman, J., Heess, N., Weber, T., and Abbeel, P · 2015
Cited alongside, same era.
Density estimation using real nvp
Dinh, L., Sohl-Dickstein, J., and Bengio, S · 2016
Cited alongside, same era.
Improving stochastic policy gradients in continuous control with deep reinforcement learning using the beta distribution
Chou, P.-W., Maturana, D., and Scherer, S · 2017
Later among the works it cites.
Value prediction network
Oh, J., Singh, S., and Lee, H · 2017
Later among the works it cites.
Ffjord: Free-form continuous dynamics for scalable reversible generative models
Grathwohl, W., Chen, R. T., Betterncourt, J., Sutskever, I., and Duvenaud, D · 2018
Later among the works it cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Haarnoja, T., Zhou, A., Abbeel, P., and Levine, S · 2018
Later among the works it cites.
Huang, C.-W., Krueger, D., Lacoste, A., and Courville, A · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Deep reinforcement learning with double q-learning
Van Hasselt, H., Guez, A., and Silver, D · 2016
Cited alongside, same era.
Arjovsky, M., Chintala, S., and Bottou, L · 2017
Cited alongside, same era.
Later among the works it cites.
Glow: Generative flow with invertible 1x1 convolutions
Kingma, D. P. and Dhariwal, P · 2018
Later among the works it cites.
Boosting trust region policy optimization by normalizing flows policy
Tang, Y. and Agrawal, S · 2018
Later among the works it cites.