Fetching the paper…
Reading the bibliography…
We propose to improve trust region policy search with normalizing flows policy.
Numerical optimization
Stephen Wright and Jorge Nocedal · 1999
Earlier work this paper cites.
Approximately optimal approximate reinforcement learning
Sham Kakade and John Langford · 2002
Earlier work this paper cites.
A natural policy gradient
Sham M Kakade · 2002
Earlier work this paper cites.
General duality between optimal control and estimation
Emanuel Todorov · 2008
Earlier work this paper cites.
Nice: Non-linear independent components estimation
Laurent Dinh, David Krueger, and Yoshua Bengio · 2014
Earlier work this paper cites.
Optimizing neural networks with kronecker-factored approximate curvature
James Martens and Roger Grosse · 2015
Earlier work this paper cites.
Variational inference with normalizing flows
Danilo Jimenez Rezende and Shakir Mohamed · 2015
Earlier work this paper cites.
Trust region policy optimization
John Schulman, Sergey Levine, Pieter Abbeel, Michael Jordan, and Philipp Moritz · 2015
Earlier work this paper cites.
Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wojciech Zaremba · 2016
Cited alongside, same era.
Density estimation using real nvp
Laurent Dinh, Jascha Sohl-Dickstein, and Samy Bengio · 2016
Cited alongside, same era.
Benchmarking deep reinforcement learning for continuous control
Yan Duan, Xi Chen, Rein Houthooft, John Schulman, and Pieter Abbeel · 2016
Cited alongside, same era.
Stein variational gradient descent: A general purpose bayesian inference algorithm
Qiang Liu and Dilin Wang · 2016
Cited alongside, same era.
Asynchronous methods for deep reinforcement learning
Volodymyr Mnih, Adria Puigdomenech Badia, Mehdi Mirza, Alex Graves, Timothy Lillicrap, Tim Harley, David Silver, and Koray Kavukcuoglu · 2016
Cited alongside, same era.
Reinforcement learning with deep energy-based policies
Tuomas Haarnoja, Haoran Tang, Pieter Abbeel, and Sergey Levine · 2017
Later among the works it cites.
Deep reinforcement learning that matters
Peter Henderson, Riashat Islam, Philip Bachman, Joelle Pineau, Doina Precup, and David Meger · 2017
Later among the works it cites.
Multiplicative normalizing flows for variational bayesian neural networks
Christos Louizos and Max Welling · 2017
Later among the works it cites.
Glow: Generative flow with invertible 1x1 convolutions
Diederik P Kingma and Prafulla Dhariwal · 2018
Closest in time.
Reinforcement learning and control as probabilistic inference: Tutorial and review
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yuhuai Wu, Elman Mansimov, Roger B Gross, Shun Liao, and Jummy Ba · 2016
Cited alongside, same era.
Improving stochastic policy gradients in continuous control with deep reinforcement learning using the beta distribution
Po-Wei Chou, Daniel Maturana, and Sebastian Scherer · 2017
Cited alongside, same era.
Openai baselines
Prafulla Dhariwal, Christopher Hesse, Oleg Klimov, Alex Nichol, Matthias Plappert, Alec Radford, John Schulman, Szymon Sidor, and Yuhuai Wu · 2017
Cited alongside, same era.
Latent space policies for hierarchical reinforcement learning
Tuomas Haarnoja, Kristian Hartikainen, Pieter Abbeel, and Sergey Levine
Cited in the paper.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine
Cited in the paper.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov
Cited in the paper.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov
Cited in the paper.
Sergey Levine · 2018
Closest in time.
Simple random search provides a competitive approach to reinforcement learning
Horia Mania, Aurelia Guy, and Benjamin Recht · 2018
Closest in time.
Implicit policy for reinforcement learning
Yunhao Tang and Shipra Agrawal · 2018
Closest in time.