Fetching the paper…
Reading the bibliography…
Finding different solutions to the same problem is a key aspect of intelligence associated with creativity and adaptation to novel situations.
Applied imagination
Alex F Osborn · 1953
Earlier work this paper cites.
Convex analysis
Ralph Tyrell Rockafellar · 1970
Earlier work this paper cites.
Markov decision processes: discrete stochastic dynamic programming
Martin L Puterman · 1984
Earlier work this paper cites.
Constrained Markov decision processes , volume 7
Eitan Altman · 1999
Earlier work this paper cites.
Apprenticeship learning via inverse reinforcement learning
Pieter Abbeel and Andrew Y Ng · 2004
Earlier work this paper cites.
An actor-critic algorithm for constrained markov decision processes
Vivek S Borkar · 2005
Earlier work this paper cites.
Transfer in variable-reward hierarchical reinforcement learning
Neville Mehta, Sriraam Natarajan, Prasad Tadepalli, and Alan Fern · 2008
Earlier work this paper cites.
Evolving a diversity of virtual creatures through novelty search and local competition
Joel Lehman and Kenneth O Stanley · 2011
Earlier work this paper cites.
Maximizing population diversity in single-objective optimization
Tamara Ulrich and Lothar Thiele · 2011
Earlier work this paper cites.
An online actor–critic algorithm with function approximation for constrained markov decision processes
Shalabh Bhatnagar and K Lakshmanan · 2012
Earlier work this paper cites.
Robots that can adapt like animals
Antoine Cully, Jeff Clune, Danesh Tarapore, and Jean-Baptiste Mouret · 2015
Earlier work this paper cites.
Illuminating search spaces by mapping elites
Jean-Baptiste Mouret and Jeff Clune · 2015
Earlier work this paper cites.
Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wojciech Zaremba · 2016
Earlier work this paper cites.
Generative adversarial imitation learning
Jonathan Ho and Stefano Ermon · 2016
Earlier work this paper cites.
Quality diversity: A new frontier for evolutionary computation
Justin K Pugh, Lisa B Soros, and Kenneth O Stanley · 2016
Earlier work this paper cites.
Chapter 2 - structure, synthesis, and application of nanoparticles
Ashok K. Singh · 2016
Earlier work this paper cites.
How do different encodings influence the performance of the map-elites algorithm?
Danesh Tarapore, Jeff Clune, Antoine Cully, and Jean-Baptiste Mouret · 2016
Earlier work this paper cites.
Scaling up map-elites using centroidal voronoi tessellations
Vassilis Vassiliades, Konstantinos Chatzilygeroudis, and Jean-Baptiste Mouret · 2016
Earlier work this paper cites.
Continuously discovering novel strategies via reward-switching policy optimization
Zihan Zhou, Wei Fu, Bingliang Zhang, and Yi Wu · 2016
Earlier work this paper cites.
On frank-wolfe and equilibrium computation
Jacob D Abernethy and Jun-Kun Wang · 2017
Earlier work this paper cites.
Successor features for transfer in reinforcement learning
Andre Barreto, Will Dabney, Remi Munos, Jonathan J Hunt, Tom Schaul, Hado P van Hasselt, and David Silver · 2017
Earlier work this paper cites.
Quality and diversity optimization: A unifying modular framework
Antoine Cully and Yiannis Demiris · 2017
Earlier work this paper cites.
Variational intrinsic control
Karol Gregor, Danilo Jimenez Rezende, and Daan Wierstra · 2017
Earlier work this paper cites.
Stein variational policy gradient
Yang Liu, Prajit Ramachandran, Qiang Liu, and Jian Peng · 2017
Cited alongside, same era.
Curiosity-driven exploration by self-supervised prediction
Deepak Pathak, Pulkit Agrawal, Alexei A Efros, and Trevor Darrell · 2017
Cited alongside, same era.
Transfer in deep reinforcement learning using successor features and generalised policy improvement
Andre Barreto, Diana Borsa, John Quan, Tom Schaul, David Silver, Matteo Hessel, Daniel Mankowitz, Augustin Zidek, and Remi Munos · 2018
Cited alongside, same era.
Impala: Scalable distributed deep-rl with importance weighted actor-learner architectures
Lasse Espeholt, Hubert Soyer, Remi Munos, Karen Simonyan, Vlad Mnih, Tom Ward, Yotam Doron, Vlad Firoiu, Tim Harley, Iain Dunning, et al · 2018
Cited alongside, same era.
Reinforcement learning for improving agent design
David Ha · 2018
Cited alongside, same era.
Off-policy actor-critic with shared experience replay
Simon Schmitt, Matteo Hessel, and Karen Simonyan · 2020
Later among the works it cites.
Dynamics-aware unsupervised discovery of skills
Archit Sharma, Shixiang Gu, Sergey Levine, Vikash Kumar, and Karol Hausman · 2020
Later among the works it cites.
Responsive safety in reinforcement learning by pid lagrangian methods
Adam Stooke, Joshua Achiam, and Pieter Abbeel · 2020
Later among the works it cites.
Novel policy seeking with constrained optimization
Hao Sun, Zhenghao Peng, Bo Dai, Jian Guo, Dahua Lin, and Bolei Zhou · 2020
Later among the works it cites.
Constrained mdps and the reward hypothesis
Csaba Szepesvári · 2020
Later among the works it cites.
Apprenticeship learning via frank-wolfe
Tom Zahavy, Alon Cohen, Haim Kaplan, and Yishay Mansour · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Diversity-driven exploration strategy for deep reinforcement learning
Zhang-Wei Hong, Tzu-Yun Shann, Shih-Yang Su, Yi-Hsiang Chang, Tsu-Jui Fu, and Chun-Yi Lee · 2018
Cited alongside, same era.
Yuval Tassa, Yotam Doron, Alistair Muldal, Tom Erez, Yazhe Li, Diego de Las Casas, David Budden, Abbas Abdolmaleki, Josh Merel, Andrew Lefrancq, et al · 2018
Cited alongside, same era.
The option keyboard: Combining skills in reinforcement learning
André Barreto, Diana Borsa, Shaobo Hou, Gheorghe Comanici, Eser Aygün, Philippe Hamel, Daniel Toyama, Shibl Mourad, David Silver, Doina Precup, et al · 2019
Cited alongside, same era.
Autonomous skill discovery with quality-diversity and unsupervised descriptors
Antoine Cully · 2019
Cited alongside, same era.
Challenges of real-world reinforcement learning
Gabriel Dulac-Arnold, Daniel Mankowitz, and Todd Hester · 2019
Cited alongside, same era.
Diversity is all you need: Learning skills without a reward function
Benjamin Eysenbach, Abhishek Gupta, Julian Ibarz, and Sergey Levine · 2019
Cited alongside, same era.
Provably efficient maximum entropy exploration
Elad Hazan, Sham Kakade, Karan Singh, and Abby Van Soest · 2019
Cited alongside, same era.
Later among the works it cites.
Variational policy gradient method for reinforcement learning with general utilities
Junyu Zhang, Alec Koppel, Amrit Singh Bedi, Csaba Szepesvari, and Mengdi Wang · 2020
Later among the works it cites.
Constructing a good behavior basis for transfer using generalized policy updates
Safa Alver and Doina Precup · 2021
Later among the works it cites.
Relative variational intrinsic control
Kate Baumli, David Warde-Farley, Steven Hansen, and Volodymyr Mnih · 2021
Later among the works it cites.
Inverse reinforcement learning in contextual mdps
Stav Belogolovsky, Philip Korsunsky, Shie Mannor, Chen Tessler, and Tom Zahavy · 2021
Later among the works it cites.
Balancing constraints and rewards with meta-gradient d4pg
Dan A. Calian, Daniel J Mankowitz, Tom Zahavy, Zhongwen Xu, Junhyuk Oh, Nir Levine, and Timothy Mann · 2021
Later among the works it cites.
The information geometry of unsupervised reinforcement learning
Benjamin Eysenbach, Ruslan Salakhutdinov, and Sergey Levine · 2021
Later among the works it cites.
Adversarially guided actor-critic
Yannis Flet-Berliac, Johan Ferret, Olivier Pietquin, Philippe Preux, and Matthieu Geist · 2021
Later among the works it cites.
Concave utility reinforcement learning: the mean-field game viewpoint
Matthieu Geist, Julien Pérolat, Mathieu Laurière, Romuald Elie, Sarah Perrin, Olivier Bachem, Rémi Munos, and Olivier Pietquin · 2021
Later among the works it cites.
Podracer architectures for scalable reinforcement learning
Matteo Hessel, Manuel Kroiss, Aidan Clark, Iurii Kemaev, John Quan, Thomas Keck, Fabio Viola, and Hado van Hasselt · 2021
Later among the works it cites.
Trajectory diversity for zero-shot coordination
Andrei Lupu, Brandon Cui, Hengyuan Hu, and Jakob Foerster · 2021
Later among the works it cites.
Policy gradient assisted map-elites
Olle Nilsson and Antoine Cully · 2021
Later among the works it cites.
Online apprenticeship learning
Lior Shani, Tom Zahavy, and Shie Mannor · 2021
Later among the works it cites.
Generalisation in lifelong reinforcement learning through logical composition
Geraud Nangue Tasse, Steven James, and Benjamin Rosman · 2021
Later among the works it cites.
Challenging common assumptions in convex reinforcement learning
Mirco Mutti, Riccardo De Santi, Piersilvio De Bartolomeis, and Marcello Restelli · 2022
Closest in time.
Diversity policy gradient for sample efficient quality-diversity optimization
Thomas Pierrot, Valentin Macé, Felix Chalumeau, Arthur Flajolet, Geoffrey Cideron, Karim Beguir, Antoine Cully, Olivier Sigaud, and Nicolas Perrin-Gilbert · 2022
Closest in time.
Approximating gradients for differentiable quality diversity in reinforcement learning
Bryon Tjanaka, Matthew C Fontaine, Julian Togelius, and Stefanos Nikolaidis · 2022
Closest in time.