Fetching the paper…
Reading the bibliography…
Deep Reinforcement Learning (RL) has emerged as a powerful paradigm for training neural policies to solve complex control tasks.
Auto-encoding variational bayes
Diederik P Kingma and Max Welling · 2013
Earlier work this paper cites.
Robots that can adapt like animals
Antoine Cully, Jeff Clune, Danesh Tarapore, and Jean-Baptiste Mouret · 2015
Earlier work this paper cites.
Illuminating search spaces by mapping elites
Jean-Baptiste Mouret and Jeff Clune · 2015
Earlier work this paper cites.
Variational intrinsic control
Karol Gregor, Danilo Jimenez Rezende, and Daan Wierstra · 2016
Earlier work this paper cites.
Quality diversity: A new frontier for evolutionary computation
Justin K Pugh, Lisa B Soros, and Kenneth O Stanley · 2016
Earlier work this paper cites.
Towards principled methods for training generative adversarial networks
Martin Arjovsky and Léon Bottou · 2017
Earlier work this paper cites.
Quality and diversity optimization: A unifying modular framework
Antoine Cully and Yiannis Demiris · 2017
Earlier work this paper cites.
Robust adversarial reinforcement learning
Lerrel Pinto, James Davidson, Rahul Sukthankar, and Abhinav Gupta · 2017
Earlier work this paper cites.
Evolution strategies as a scalable alternative to reinforcement learning
Tim Salimans, Jonathan Ho, Xi Chen, Szymon Sidor, and Ilya Sutskever · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Earlier work this paper cites.
Mastering chess and shogi by self-play with a general reinforcement learning algorithm
David Silver, Thomas Hubert, Julian Schrittwieser, Ioannis Antonoglou, Matthew Lai, Arthur Guez, Marc Lanctot, Laurent Sifre, Dharshan Kumaran, Thore Graepel, et al · 2017
Earlier work this paper cites.
Domain randomization for transferring deep neural networks from simulation to the real world
Josh Tobin, Rachel Fong, Alex Ray, Jonas Schneider, Wojciech Zaremba, and Pieter Abbeel · 2017
Earlier work this paper cites.
Using centroidal voronoi tessellations to scale up the multidimensional archive of phenotypic elites algorithm
Vassilis Vassiliades, Konstantinos Chatzilygeroudis, and Jean-Baptiste Mouret · 2017
Earlier work this paper cites.
JAX: composable transformations of Python+NumPy programs, 2018
James Bradbury, Roy Frostig, Peter Hawkins, Matthew James Johnson, Chris Leary, Dougal Maclaurin, George Necula, Adam Paszke, Jake VanderPlas, Skye Wanderman-Milne, and Qiao Zhang · 2018
Earlier work this paper cites.
Improving exploration in evolution strategies for deep reinforcement learning via a population of novelty-seeking agents
Edoardo Conti, Vashisht Madhavan, Felipe Petroski Such, Joel Lehman, Kenneth Stanley, and Jeff Clune · 2018
Earlier work this paper cites.
Meta learning shared hierarchies
Kevin Frans, Jonathan Ho, Xi Chen, Pieter Abbeel, and John Schulman · 2018
Earlier work this paper cites.
Addressing function approximation error in actor-critic methods
Scott Fujimoto, Herke Hoof, and David Meger · 2018
Earlier work this paper cites.
Learning an embedding space for transferable robot skills
Karol Hausman, Jost Tobias Springenberg, Ziyu Wang, Nicolas Heess, and Martin Riedmiller · 2018
Cited alongside, same era.
Discovering the elite hypervolume by leveraging interspecies correlation
Vassiiis Vassiliades and Jean-Baptiste Mouret · 2018
Cited alongside, same era.
Continuously discovering novel strategies via reward-switching policy optimization
Zihan Zhou, Wei Fu, Bingliang Zhang, and Yi Wu · 2018
Cited alongside, same era.
Autonomous skill discovery with quality-diversity and unsupervised descriptors
Antoine Cully · 2019
Cited alongside, same era.
Diversity is all you need: Learning skills without a reward function
Benjamin Eysenbach, Abhishek Gupta, Julian Ibarz, and Sergey Levine · 2019
Cited alongside, same era.
Diversity-inducing policy gradient: Using maximum mean discrepancy to find a set of diverse policies
Quality-diversity optimization: a novel branch of stochastic optimization
Konstantinos Chatzilygeroudis, Antoine Cully, Vassilis Vassiliades, and Jean-Baptiste Mouret · 2021
Later among the works it cites.
Wasserstein distance maximizing intrinsic control
Ishan Durugkar, Steven Hansen, Stephen Spencer, and Volodymyr Mnih · 2021
Later among the works it cites.
Brax–a differentiable physics engine for large scale rigid body simulation
C Daniel Freeman, Erik Frey, Anton Raichuk, Sertan Girgin, Igor Mordatch, and Olivier Bachem · 2021
Later among the works it cites.
Learning a subspace of policies for online adaptation in reinforcement learning
Jean-Baptiste Gaya, Laure Soulier, and Ludovic Denoyer · 2021
Later among the works it cites.
Entropic desired dynamics for intrinsic control
Steven Hansen, Guillaume Desjardins, Kate Baumli, David Warde-Farley, Nicolas Heess, Simon Osindero, and Volodymyr Mnih · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Muhammad A Masood and Finale Doshi-Velez · 2019
Cited alongside, same era.
Dynamics-aware unsupervised discovery of skills
Archit Sharma, Shixiang Gu, Sergey Levine, Vikash Kumar, and Karol Hausman · 2019
Cited alongside, same era.
Learning novel policies for tasks
Yunbo Zhang, Wenhao Yu, and Greg Turk · 2019
Cited alongside, same era.
Learning dexterous in-hand manipulation
OpenAI: Marcin Andrychowicz, Bowen Baker, Maciek Chociej, Rafal Jozefowicz, Bob McGrew, Jakub Pachocki, Arthur Petron, Matthias Plappert, Glenn Powell, Alex Ray, et al · 2020
Cited alongside, same era.
Explore, discover and learn: Unsupervised discovery of state-covering skills
Víctor Campos, Alexander Trott, Caiming Xiong, Richard Socher, Xavier Giró-i Nieto, and Jordi Torres · 2020
Cited alongside, same era.
Scaling map-elites to deep neuroevolution
Cédric Colas, Vashisht Madhavan, Joost Huizinga, and Jeff Clune · 2020
Cited alongside, same era.
Covariance matrix adaptation for the rapid illumination of behavior space
Matthew C Fontaine, Julian Togelius, Stefanos Nikolaidis, and Amy K Hoover · 2020
Cited alongside, same era.
Later among the works it cites.
Direct then diffuse: Incremental unsupervised skill discovery for state covering and goal reaching
Pierre-Alexandre Kamienny, Jean Tarbouriech, Sylvain Lamprier, Alessandro Lazaric, and Ludovic Denoyer · 2021
Later among the works it cites.
Policy gradient assisted map-elites
Olle Nilsson and Antoine Cully · 2021
Later among the works it cites.
Discovering a set of policies for the worst case reward
Tom Zahavy, Andre Barreto, Daniel J Mankowitz, Shaobo Hou, Brendan O’Donoghue, Iurii Kemaev, and Satinder Baveja Singh · 2021
Later among the works it cites.
Magnetic control of tokamak plasmas through deep reinforcement learning
Jonas Degrave, Federico Felici, Jonas Buchli, Michael Neunert, Brendan Tracey, Francesco Carpanese, Timo Ewalds, Roland Hafner, Abbas Abdolmaleki, Diego de Las Casas, et al · 2022
Closest in time.
Diversifying behaviors for learning in asymmetric multiagent systems
Gaurav Dixit, Everardo Gonzalez, and Kagan Tumer · 2022
Closest in time.
Cic: Contrastive intrinsic control for unsupervised skill discovery
Michael Laskin, Hao Liu, Xue Bin Peng, Denis Yarats, Aravind Rajeswaran, and Pieter Abbeel · 2022
Closest in time.
Accelerated quality-diversity for robotics through massive parallelism
Bryan Lim, Maxime Allard, Luca Grillotti, and Antoine Cully · 2022
Closest in time.
Discovering diverse solutions in deep reinforcement learning by maximizing state–action-based mutual information
Takayuki Osa, Voot Tangkaratt, and Masashi Sugiyama · 2022
Closest in time.
Diversity policy gradient for sample efficient quality-diversity optimization
Thomas Pierrot, Valentin Macé, Félix Chalumeau, Arthur Flajolet, Geoffrey Cideron, Karim Beguir, Antoine Cully, Olivier Sigaud, and Nicolas Perrin-Gilbert · 2022
Closest in time.
Approximating gradients for differentiable quality diversity in reinforcement learning
Bryon Tjanaka, Matthew C Fontaine, Julian Togelius, and Stefanos Nikolaidis · 2022
Closest in time.
Qdax: A library for quality-diversity and population-based algorithms with hardware acceleration, 2023
Felix Chalumeau, Bryan Lim, Raphael Boige, Maxime Allard, Luca Grillotti, Manon Flageat, Valentin Macé, Arthur Flajolet, Thomas Pierrot, and Antoine Cully · 2023
Closest in time.