Fetching the paper…
Reading the bibliography…
We present Reward-Switching Policy Optimization (RSPO), a paradigm to discover diverse strategies in complex RL environments by iteratively finding novel policies that are both locally optimal and sufficiently different from existing ones.
Markov decision processes: Discrete stochastic dynamic programming
M. Puterman · 1994
Earlier work this paper cites.
Genetic algorithms with dynamic niche sharing for multimodal function optimization
B. Miller and Michael J. Shaw · 1996
Earlier work this paper cites.
The cross-entropy method for continuous multi-extremal optimization
Dirk P. Kroese, S. Porotsky, and R. Rubinstein · 2006
Earlier work this paper cites.
Finding multiple solutions for multimodal optimization problems using a multi-objective evolutionary approach
K. Deb and Amit Saha · 2010
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
E. Todorov, T. Erez, and Y. Tassa · 2012
Earlier work this paper cites.
Empirical evaluation of gated recurrent neural networks on sequence modeling
Junyoung Chung, Caglar Gulcehre, KyungHyun Cho, and Yoshua Bengio · 2014
Earlier work this paper cites.
Robots that can adapt like animals
Antoine Cully, Jeff Clune, Danesh Tarapore, and Jean-Baptiste Mouret · 2015
Earlier work this paper cites.
Trust region policy optimization
John Schulman, Sergey Levine, P. Abbeel, Michael I. Jordan, and P. Moritz · 2015
Earlier work this paper cites.
Deep reinforcement learning for dialogue generation
Jiwei Li, Will Monroe, Alan Ritter, Dan Jurafsky, Michel Galley, and Jianfeng Gao · 2016
Earlier work this paper cites.
Quality diversity: A new frontier for evolutionary computation
Justin K. Pugh, L. B. Soros, and K. Stanley · 2016
Earlier work this paper cites.
A unified game-theoretic approach to multiagent reinforcement learning
Marc Lanctot, V. Zambaldi, A. Gruslys, A. Lazaridou, K. Tuyls, J. Pérolat, D. Silver, and T. Graepel · 2017
Earlier work this paper cites.
Curiosity-driven exploration by self-supervised prediction
Deepak Pathak, Pulkit Agrawal, Alexei A Efros, and Trevor Darrell · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
John Schulman, F. Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Earlier work this paper cites.
Starcraft ii: A new challenge for reinforcement learning
Oriol Vinyals, Timo Ewalds, Sergey Bartunov, Petko Georgiev, Alexander Sasha Vezhnevets, Michelle Yeo, Alireza Makhzani, Heinrich Küttler, John Agapiou, Julian Schrittwieser, et al · 2017
Earlier work this paper cites.
Addressing function approximation error in actor-critic methods
Scott Fujimoto, Herke Hoof, and David Meger · 2018
Earlier work this paper cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine · 2018
Earlier work this paper cites.
Diversity-driven exploration strategy for deep reinforcement learning
Zhang-Wei Hong, Tzu-Yun Shann, Shih-Yang Su, Y. Chang, and Chun-Yi Lee · 2018
Earlier work this paper cites.
Emergence of grounded compositional language in multi-agent populations
Igor Mordatch and P. Abbeel · 2018
Cited alongside, same era.
Exploration by random network distillation
Yuri Burda, Harrison Edwards, A. Storkey, and Oleg Klimov · 2019
Cited alongside, same era.
Diversity is all you need: Learning skills without a reward function
Benjamin Eysenbach, A. Gupta, J. Ibarz, and Sergey Levine · 2019
Cited alongside, same era.
Human-level performance in 3d multiplayer games with population-based reinforcement learning
Max Jaderberg, Wojciech M Czarnecki, Iain Dunning, Luke Marris, Guy Lever, Antonio Garcia Castaneda, Charles Beattie, Neil C Rabinowitz, Ari S Morcos, Avraham Ruderman, et al · 2019
Cited alongside, same era.
Emergent coordination through competition
Siqi Liu, G. Lever, J. Merel, S. Tunyasuvunakool, N. Heess, and T. Graepel · 2019
Cited alongside, same era.
Maven: Multi-agent variational exploration
Efficient continuous pareto exploration in multi-task learning
Pingchuan Ma, Tao Du, and Wojciech Matusik · 2020
Later among the works it cites.
Why do local methods solve nonconvex problems?
Tengyu Ma · 2020
Later among the works it cites.
Navigating the landscape of multiplayer games
Shayegan Omidshafiei, Karl Tuyls, Wojciech M Czarnecki, Francisco C Santos, Mark Rowland, Jerome Connor, Daniel Hennes, Paul Muller, Julien Pérolat, Bart De Vylder, et al · 2020
Later among the works it cites.
Reward prediction error as an exploration objective in deep rl
Riley Simmons-Edler, Ben Eisner, Daniel Yang, Anthony Bisulco, Eric Mitchell, Sebastian Seung, and Daniel Lee · 2020
Later among the works it cites.
Novel policy seeking with constrained optimization
Hao Sun, Zhenghao Peng, Bo Dai, Jian Guo, Dahua Lin, and Bolei Zhou · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Anuj Mahajan, Tabish Rashid, Mikayel Samvelyan, and S. Whiteson · 2019
Cited alongside, same era.
Diversity-inducing policy gradient: Using maximum mean discrepancy to find a set of diverse policies
M. A. Masood and Finale Doshi-Velez · 2019
Cited alongside, same era.
The Starcraft multi-agent challenge
Tabish Rashid, Philip HS Torr, Gregory Farquhar, Chia-Man Hung, Tim GJ Rudner, Nantas Nardelli, Shimon Whiteson, Christian Schroeder de Witt, Jakob Foerster, and Mikayel Samvelyan · 2019
Cited alongside, same era.
Qxplore: Q-learning exploration by maximizing temporal difference error
Riley Simmons-Edler, Ben Eisner, Daniel Yang, Anthony Bisulco, Eric Mitchell, Sebastian Seung, and Daniel Lee · 2019
Cited alongside, same era.
Grandmaster level in StarCraft II using multi-agent reinforcement learning
Oriol Vinyals, Igor Babuschkin, Wojciech M Czarnecki, Michaël Mathieu, Andrew Dudzik, Junyoung Chung, David H Choi, Richard Powell, Timo Ewalds, Petko Georgiev, et al · 2019
Cited alongside, same era.
Poet: open-ended coevolution of environments and their optimized solutions
Rui Wang, Joel Lehman, Jeff Clune, and Kenneth O Stanley · 2019
Cited alongside, same era.
Emergent tool use from multi-agent autocurricula
Bowen Baker, Ingmar Kanitscheider, Todor Markov, Yi Wu, Glenn Powell, Bob McGrew, and Igor Mordatch · 2020
Cited alongside, same era.
Tonghan Wang*, Jianhao Wang*, Yi Wu, and Chongjie Zhang · 2020
Later among the works it cites.
The AI economist: Improving equality and productivity with AI-driven tax policies
Stephan Zheng, Alexander Trott, Sunil Srinivasa, Nikhil Naik, Melvin Gruesbeck, David C Parkes, and Richard Socher · 2020
Later among the works it cites.
Unifying behavioral and response diversity for open-ended learning in zero-sum games
Xiangyu Liu, Hangtian Jia, Ying Wen, Yaodong Yang, Yujing Hu, Yingfeng Chen, Changjie Fan, and Zhipeng Hu · 2021
Later among the works it cites.
Trajectory diversity for zero-shot coordination
Andrei Lupu, Brandon Cui, Hengyuan Hu, and Jakob Foerster · 2021
Later among the works it cites.
Modelling behavioural diversity for learning in open-ended games
Nicolas Perez Nieves, Yaodong Yang, Oliver Slumbers, David Mguni, and Jun Wang · 2021
Later among the works it cites.
Policy gradient assisted map-elites
Olle Nilsson and Antoine Cully · 2021
Later among the works it cites.
Diversity oriented deep reinforcement learning for targeted molecule generation
Tiago Pereira, Maryam Abbasi, Bernardete Ribeiro, and Joel P. Arrais · 2021
Later among the works it cites.
Discovering diverse multi-agent strategic behavior via reward randomization
Zhenggang Tang, C. Yu, Boyuan Chen, Huazhe Xu, Xiaolong Wang, Fei Fang, S. Du, Yu Wang, and Yi Wu · 2021
Later among the works it cites.
The surprising effectiveness of mappo in cooperative, multi-agent games
Chao Yu, Akash Velu, Eugene Vinitsky, Yu Wang, Alexandre Bayen, and Yi Wu · 2021
Later among the works it cites.
Discovering diverse nearly optimal policies with successor features
Tom Zahavy, Brendan O’Donoghue, Andre Barreto, Volodymyr Mnih, Sebastian Flennerhag, and Satinder Singh · 2021
Later among the works it cites.
On learning intrinsic rewards for policy gradient methods
Zeyu Zheng, Junhyuk Oh, and Satinder Singh · 2022
Closest in time.