Fetching the paper…
Reading the bibliography…
We introduce a new approach for comparing reinforcement learning policies, using Wasserstein distances (WDs) in a newly defined latent behavioral space.
On information and sufficiency
S. Kullback and R. A. Leibler · 1951
Earlier work this paper cites.
General Topology: Elements of Mathematics
N. Bourbaki · 1966
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
R. J. Williams · 1992
Earlier work this paper cites.
Introduction to Reinforcement Learning
R. S. Sutton, A. G. Barto, et al · 1998
Earlier work this paper cites.
Approximately optimal approximate reinforcement learning
S. Kakade and J. Langford · 2002
Earlier work this paper cites.
Universal kernels
C. A. Micchelli, Y. Xu, and H. Zhang · 2006
Earlier work this paper cites.
Exploiting open-endedness to solve problems through the search for novelty
J. Lehman and K. O. Stanley · 2008
Earlier work this paper cites.
Random features for large-scale kernel machines
A. Rahimi and B. Recht · 2008
Earlier work this paper cites.
Optimal transport: Old and new
C. Villani · 2008
Earlier work this paper cites.
Learning behavior styles with inverse reinforcement learning
S. J. Lee and Z. Popovic · 2010
Earlier work this paper cites.
Evolution through the Search for Novelty
J. Lehman · 2012
Earlier work this paper cites.
Auto-encoding variational bayes
D. P. Kingma and M. Welling · 2013
Earlier work this paper cites.
Playing atari with deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. Graves, I. Antonoglou, D. Wierstra, and M. A. Riedmiller · 2013
Earlier work this paper cites.
Trust region policy optimization
J. Schulman, S. Levine, P. Abbeel, M. I. Jordan, and P. Moritz · 2015
Cited alongside, same era.
G. Brockman, V. Cheung, L. Pettersson, J. Schneider, J. Schulman, J. Tang, and W. Zaremba · 2016
Cited alongside, same era.
Stochastic optimization for large-scale optimal transport
A. Genevay, M. Cuturi, G. Peyré, and F. Bach · 2016
Cited alongside, same era.
Learning behavior characterizations for novelty search
E. Meyerson, J. Lehman, and R. Miikkulainen · 2016
Cited alongside, same era.
Quality diversity: A new frontier for evolutionary computation
J. K. Pugh, L. B. Soros, and K. O. Stanley · 2016
Cited alongside, same era.
Wasserstein generative adversarial networks
M. Arjovsky, S. Chintala, and L. Bottou · 2017
Implicit quantile networks for distributional reinforcement learning
W. Dabney, G. Ostrovski, D. Silver, and R. Munos · 2018
Later among the works it cites.
Noisy networks for exploration
M. Fortunato, M. G. Azar, B. Piot, J. Menick, M. Hessel, I. Osband, A. Graves, V. Mnih, R. Munos, D. Hassabis, O. Pietquin, C. Blundell, and S. Legg · 2018
Later among the works it cites.
Simple random search of static linear policies is competitive for reinforcement learning
H. Mania, A. Guy, and B. Recht · 2018
Later among the works it cites.
An analysis of categorical distributional reinforcement learning
M. Rowland, M. Bellemare, W. Dabney, R. Munos, and Y. W. Teh · 2018
Later among the works it cites.
Y. Tassa, Y. Doron, A. Muldal, T. Erez, Y. Li, D. d. L. Casas, D. Budden, A. Abdolmaleki, J. Merel, A. Lefrancq, et al · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
A distributional perspective on reinforcement learning
M. G. Bellemare, W. Dabney, and R. Munos · 2017
Cited alongside, same era.
Openai baselines
P. Dhariwal, C. Hesse, O. Klimov, A. Nichol, M. Plappert, A. Radford, J. Schulman, S. Sidor, Y. Wu, and P. Zhokhov · 2017
Cited alongside, same era.
On wasserstein reinforcement learning and the fokker-planck equation
P. H. Richemond and B. Maginnis · 2017
Cited alongside, same era.
Evolution strategies as a scalable alternative to reinforcement learning
T. Salimans, J. Ho, X. Chen, S. Sidor, and I. Sutskever · 2017
Cited alongside, same era.
Proximal policy optimization algorithms
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov · 2017
Cited alongside, same era.
Improving exploration in evolution strategies for deep reinforcement learning via a population of novelty-seeking agents
E. Conti, V. Madhavan, F. P. Such, J. Lehman, K. O. Stanley, and J. Clune · 2018
Cited alongside, same era.
Policy optimization as wasserstein gradient flows
R. Zhang, C. Chen, C. Li, and L. Carin · 2018
Later among the works it cites.
Distributional reinforcement learning with linear function approximation
M. G. Bellemare, N. L. Roux, P. S. Castro, and S. Moitra · 2019
Closest in time.
Provably robust blackbox optimization for reinforcement learning
K. Choromanski, A. Pacchiano, J. Parker-Holder, Y. Tang, D. Jain, Y. Yang, A. Iscen, J. Hsu, and V. Sindhwani · 2019
Closest in time.
Evolvability es: Scalable and direct optimization of evolvability
A. Gajewski, J. Clune, K. O. Stanley, and J. Lehman · 2019
Closest in time.
Wasserstein fair classification
R. Jiang, A. Pacchiano, T. Stepleton, H. Jiang, and S. Chiappa · 2019
Closest in time.
Nonlinear distributional gradient temporal-difference learning
C. Qu, S. Mannor, and H. Xu · 2019
Closest in time.
Statistics and samples in distributional reinforcement learning
M. Rowland, R. Dadashi, S. Kumar, R. Munos, M. G. Bellemare, and W. Dabney · 2019
Closest in time.