Fetching the paper…
Reading the bibliography…
The ability to discover approximately optimal policies in domains with sparse rewards is crucial to applying reinforcement learning (RL) in many real-world scenarios.
A markovian decision process
R. Bellman · 1957
Earlier work this paper cites.
Function optimization using connectionist reinforcement learning algorithms
R. J. Williams and J. Peng · 1991
Earlier work this paper cites.
Estimating the number of clusters in a data set via the gap statistic
R. Tibshirani, G. Walther, and T. Hastie · 2001
Earlier work this paper cites.
Novelty search and the problem with objectives
J. Lehman and K. O. Stanley · 2011
Earlier work this paper cites.
Evolution through the search for novelty
J. Lehman · 2012
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
E. Todorov, T. Erez, and Y. Tassa · 2012
Earlier work this paper cites.
Generative adversarial nets
I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio · 2014
Earlier work this paper cites.
Markov decision processes: discrete stochastic dynamic programming
M. L. Puterman · 2014
Earlier work this paper cites.
Auto-encoding variational bayes
D. P. Kingma and M. Welling · 2014
Earlier work this paper cites.
Variational inference with normalizing flows
D. J. Rezende and S. Mohamed · 2015
Earlier work this paper cites.
Trust region policy optimization
J. Schulman, S. Levine, P. Moritz, M. I. Jordan, and P. Abbeel · 2015
Earlier work this paper cites.
Markov chain monte carlo and variational inference: Bridging the gap
T. Salimans, D. Kingma, and M. Welling · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2015
Cited alongside, same era.
Pybullet, a python module for physics simulation for games, robotics and machine learning
E. Coumans and Y. Bai · 2016
Cited alongside, same era.
Continuous control with deep reinforcement learning
T. P. Lillicrap, J. J. Hunt, A. Pritzel, N. Heess, T. Erez, Y. Tassa, D. Silver, and D. Wierstra · 2016
Cited alongside, same era.
Asynchronous methods for deep reinforcement learning
V. Mnih, A. P. Badia, M. Mirza, A. Graves, T. P. Lillicrap, T. Harley, D. Silver, and K. Kavukcuoglu · 2016
Cited alongside, same era.
Dropout as a bayesian approximation: Representing model uncertainty in deep learning
Y. Gal and Z. Ghahramani · 2016
Cited alongside, same era.
Deep exploration via bootstrapped DQN
Reinforcement learning with deep energy-based policies
T. Haarnoja, H. Tang, P. Abbeel, and S. Levine · 2017
Later among the works it cites.
Bayesian policy gradients via alpha divergence dropout inference
P. Henderson, T. Doan, R. Islam, and D. Meger · 2017
Later among the works it cites.
Proximal policy optimization algorithms
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov · 2017
Later among the works it cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
T. Haarnoja, A. Zhou, P. Abbeel, and S. Levine · 2018
Later among the works it cites.
Addressing function approximation error in actor-critic methods
S. Fujimoto, H. van Hoof, and D. Meger · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
I. Osband, C. Blundell, A. Pritzel, and B. V. Roy · 2016
Cited alongside, same era.
Improving variational inference with inverse autoregressive flow
D. P. Kingma, T. Salimans, and M. Welling · 2016
Cited alongside, same era.
Density estimation using real nvp
L. Dinh, J. Sohl-Dickstein, and S. Bengio · 2016
Cited alongside, same era.
Iterative refinement of the approximate posterior for directed belief networks
D. Hjelm, R. R. Salakhutdinov, K. Cho, N. Jojic, V. Calhoun, and J. Chung · 2016
Cited alongside, same era.
Benchmarking deep reinforcement learning for continuous control
Y. Duan, X. Chen, R. Houthooft, J. Schulman, and P. Abbeel · 2016
Cited alongside, same era.
Improving exploration in evolution strategies for deep reinforcement learning via a population of novelty-seeking agents
E. Conti, V. Madhavan, F. P. Such, J. Lehman, K. O. Stanley, and J. Clune · 2017
Cited alongside, same era.
T. Doan, B. Mazoure, and C. Lyle · 2018
Later among the works it cites.
Boosting trust region policy optimization by normalizing flows policy
Y. Tang and S. Agrawal · 2018
Later among the works it cites.
Latent space policies for hierarchical reinforcement learning
T. Haarnoja, K. Hartikainen, P. Abbeel, and S. Levine · 2018
Later among the works it cites.
Neural autoregressive flows
C. Huang, D. Krueger, A. Lacoste, and A. C. Courville · 2018
Later among the works it cites.
Go-explore: a new approach for hard-exploration problems
A. Ecoffet, J. Huizinga, J. Lehman, K. O. Stanley, and J. Clune · 2019
Closest in time.
The machine learning reproducibility checklist, v.1.2 , 2019
J. Pineau · 2019
Closest in time.