Fetching the paper…
Reading the bibliography…
Most reinforcement learning algorithms seek a single optimal strategy that solves a given task.
The starcraft multi-agent challenge
Samvelyan, M.; Rashid, T.; De Witt, C. S.; Farquhar, G.; Nardelli, N.; Rudner, T. G.; Hung, C.-M.; Torr, P. H.; Foerster, J.; and Whiteson, S. 2019 · 1902
Earlier work this paper cites.
Stochastic latent actor-critic: Deep reinforcement learning with a latent variable model
Lee, A. X.; Nagabandi, A.; Abbeel, P.; and Levine, S. 2019 · 1907
Earlier work this paper cites.
Maven: Multi-agent variational exploration
Mahajan, A.; Rashid, T.; Samvelyan, M.; and Whiteson, S. 2019 · 1910
Earlier work this paper cites.
Dota 2 with large scale deep reinforcement learning
Berner, C.; Brockman, G.; Chan, B.; Cheung, V.; Debiak, P.; Dennison, C.; Farhi, D.; Fischer, Q.; Hashme, S.; Hesse, C.; et al. 2019 · 1912
Earlier work this paper cites.
Skill Discovery of Coordination in Multi-agent Reinforcement Learning
He, S.; Shao, J.; and Ji, X. 2020 · 2006
Earlier work this paper cites.
Maximum entropy inverse reinforcement learning
Ziebart, B. D.; Maas, A. L.; Bagnell, J. A.; and Dey, A. K. 2008 · 2008
Earlier work this paper cites.
Overcoming the bootstrap problem in evolutionary robotics using behavioral diversity
Mouret, J.-B.; and Doncieux, S. 2009 · 2009
Earlier work this paper cites.
Variational methods for reinforcement learning
Furmston, T.; and Barber, D. 2010 · 2010
Earlier work this paper cites.
The arcade learning environment: An evaluation platform for general agents
Bellemare, M. G.; Naddaf, Y.; Veness, J.; and Bowling, M. 2013 · 2013
Earlier work this paper cites.
Variational information maximisation for intrinsically motivated reinforcement learning
Mohamed, S.; and Rezende, D. J. 2015 · 2015
Earlier work this paper cites.
Deep reinforcement learning for dialogue generation
Li, J.; Monroe, W.; Ritter, A.; Galley, M.; Gao, J.; and Jurafsky, D. 2016 · 2016
Earlier work this paper cites.
A unified game-theoretic approach to multiagent reinforcement learning
Lanctot, M.; Zambaldi, V.; Gruslys, A.; Lazaridou, A.; Tuyls, K.; Pérolat, J.; Silver, D.; and Graepel, T. 2017 · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
Schulman, J.; Wolski, F.; Dhariwal, P.; Radford, A.; and Klimov, O. 2017 · 2017
Cited alongside, same era.
Diversity is all you need: Learning skills without a reward function
Eysenbach, B.; Gupta, A.; Ibarz, J.; and Levine, S. 2018 · 2018
Cited alongside, same era.
Learning an embedding space for transferable robot skills
Hausman, K.; Springenberg, J. T.; Wang, Z.; Heess, N.; and Riedmiller, M. 2018 · 2018
Cited alongside, same era.
Deep variational reinforcement learning for POMDPs
Igl, M.; Zintgraf, L.; Le, T. A.; Wood, F.; and Whiteson, S. 2018 · 2018
Cited alongside, same era.
Reinforcement learning and control as probabilistic inference: Tutorial and review
Levine, S. 2018 · 2018
Cited alongside, same era.
One solution is not all you need: Few-shot extrapolation via structured maxent rl
Kumar, S.; Kumar, A.; Levine, S.; and Finn, C. 2020 · 2020
Later among the works it cites.
Effective diversity in population based reinforcement learning
Parker-Holder, J.; Pacchiano, A.; Choromanski, K. M.; and Roberts, S. J. 2020 · 2020
Later among the works it cites.
Increasing Diversity with Deep Reinforcement Learning for Chatbots
Pavel, C.; Budulan, S.; and Rebedea, T. 2020 · 2020
Later among the works it cites.
Adaptable Agent Populations via a Generative Model of Policies
Derek, K.; and Isola, P. 2021 · 2021
Later among the works it cites.
The Information Geometry of Unsupervised Reinforcement Learning
Eysenbach, B.; Salakhutdinov, R.; and Levine, S. 2021 · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Emergence of grounded compositional language in multi-agent populations
Mordatch, I.; and Abbeel, P. 2018 · 2018
Cited alongside, same era.
S-RL Toolbox: Environments, Datasets and Evaluation Metrics for State Representation Learning
Raffin, A.; Hill, A.; Traoré, R.; Lesort, T.; Díaz-Rodríguez, N.; and Filliat, D. 2018 · 2018
Cited alongside, same era.
Qmix: Monotonic value function factorisation for deep multi-agent reinforcement learning
Rashid, T.; Samvelyan, M.; Schroeder, C.; Farquhar, G.; Foerster, J.; and Whiteson, S. 2018 · 2018
Cited alongside, same era.
Reinforcement learning: An introduction
Sutton, R. S.; and Barto, A. G. 2018 · 2018
Cited alongside, same era.
Generating multiple diverse responses for short-text conversation
Gao, J.; Bi, W.; Liu, X.; Li, J.; and Shi, S. 2019 · 2019
Cited alongside, same era.
Learning to coordinate manipulation skills via skill behavior diversification
Lee, Y.; Yang, J.; and Lim, J. J. 2019 · 2019
Cited alongside, same era.
Grandmaster level in StarCraft II using multi-agent reinforcement learning
Vinyals, O.; Babuschkin, I.; Czarnecki, W. M.; Mathieu, M.; Dudzik, A.; Chung, J.; Choi, D. H.; Powell, R.; Ewalds, T.; Georgiev, P.; et al. 2019 · 2019
Cited alongside, same era.
Huang, S.; Chen, W.; Zhang, L.; Li, Z.; Zhu, F.; Ye, D.; Chen, T.; and Zhu, J. 2021 · 2021
Later among the works it cites.
Discovering diverse solutions in deep reinforcement learning
Osa, T.; Tangkaratt, V.; and Sugiyama, M. 2021 · 2021
Later among the works it cites.
Discovering Diverse Multi-Agent Strategic Behavior via Reward Randomization
Tang, Z.; Yu, C.; Chen, B.; Xu, H.; Wang, X.; Fang, F.; Du, S.; Wang, Y.; and Wu, Y. 2021 · 2021
Later among the works it cites.
Discovering diverse nearly optimal policies with successor features
Zahavy, T.; O’Donoghue, B.; Barreto, A.; Flennerhag, S.; Mnih, V.; and Singh, S. 2021 · 2021
Later among the works it cites.
A Mixture-of-Expert Approach to RL-based Dialogue Management
Chow, Y.; Tulepbergenov, A.; Nachum, O.; Ryu, M.; Ghavamzadeh, M.; and Boutilier, C. 2022 · 2022
Closest in time.
Diverse dialogue generation by fusing mutual persona-aware and self-transferrer
Xu, F.; Xu, G.; Wang, Y.; Wang, R.; Ding, Q.; Liu, P.; and Zhu, Z. 2022 · 2022
Closest in time.
Continuously Discovering Novel Strategies via Reward-Switching Policy Optimization
Zhou, Z.; Fu, W.; Zhang, B.; and Wu, Y. 2022 · 2022
Closest in time.