Fetching the paper…
Reading the bibliography…
We propose a simple, general and effective technique, Reward Randomization for discovering diverse strategic policies in complex multi-agent games.
Xxii. programming a computer for playing chess
Claude E Shannon · 1950
Earlier work this paper cites.
Iterative solution of games by fictitious play
George W Brown · 1951
Earlier work this paper cites.
Non-cooperative games
John Nash · 1951
Earlier work this paper cites.
An iterative method of solving a game
Julia Robinson · 1951
Earlier work this paper cites.
Spieltheoretische behandlung eines oligopolmodells mit nachfrageträgheit: Teil i: Bestimmung des dynamischen preisgleichgewichts
Reinhard Selten · 1965
Earlier work this paper cites.
Reexamination of the perfectness concept for equilibrium points in extensive games
R Selten · 1975
Earlier work this paper cites.
Refinements of the Nash equilibrium concept
Roger B Myerson · 1978
Earlier work this paper cites.
A discourse on inequality
Jean-Jacques Rousseau · 1984
Earlier work this paper cites.
Equilibrium selection in signaling games
Jeffrey S Banks and Joel Sobel · 1987
Earlier work this paper cites.
Coalition-proof Nash equilibria i. concepts
B Douglas Bernheim, Bezalel Peleg, and Michael D Whinston · 1987
Earlier work this paper cites.
Learning, local interaction, and coordination
Glenn Ellison · 1993
Earlier work this paper cites.
Learning, mutation, and long run equilibria in games
Michihiro Kandori, George J Mailath, and Rafael Rob · 1993
Earlier work this paper cites.
Quantal response equilibria for normal form games
Richard D McKelvey and Thomas R Palfrey · 1995
Earlier work this paper cites.
Potential games
Dov Monderer and Lloyd S Shapley · 1996
Earlier work this paper cites.
Policy invariance under reward transformations: Theory and application to reward shaping
Andrew Y Ng, Daishi Harada, and Stuart Russell · 1999
Earlier work this paper cites.
Algorithms for inverse reinforcement learning
Andrew Y Ng, Stuart J Russell, et al · 2000
Earlier work this paper cites.
Nash convergence of gradient dynamics in general-sum games
Satinder P Singh, Michael J Kearns, and Yishay Mansour · 2000
Earlier work this paper cites.
Friend-or-foe q-learning in general-sum games
Michael L Littman · 2001
Earlier work this paper cites.
Deep blue
Murray Campbell, A Joseph Hoane Jr, and Feng-hsiung Hsu · 2002
Earlier work this paper cites.
On adaptive emergence of trust behavior in the game of stag hunt
Christina Fang, Steven Orla Kimbrough, Stefano Pace, Annapurna Valluri, and Zhiqiang Zheng · 2002
Earlier work this paper cites.
Planning in the presence of cost functions controlled by an adversary
H Brendan McMahan, Geoffrey J Gordon, and Avrim Blum · 2003
Earlier work this paper cites.
Learning to play Bayesian games
Eddie Dekel, Drew Fudenberg, and David K Levine · 2004
Earlier work this paper cites.
The stag hunt and the evolution of social structure
Brian Skyrms · 2004
Earlier work this paper cites.
Social reward shaping in the prisoner’s dilemma
Monica Babes, Enrique Munoz de Cote, and Michael L Littman · 2008
Earlier work this paper cites.
Maximum entropy inverse reinforcement learning
Brian D Ziebart, Andrew L Maas, J Andrew Bagnell, and Anind K Dey · 2008
Earlier work this paper cites.
A dynamic model of social network formation
Brian Skyrms and Robin Pemantle · 2009
Earlier work this paper cites.
Individual and cultural learning in stag hunt games with multiple actions
Russell Golman and Scott E Page · 2010
Earlier work this paper cites.
Refinement of strong stackelberg equilibria in security games
Bo An, Milind Tambe, Fernando Ordonez, Eric Shieh, and Christopher Kiekintveld · 2011
Cited alongside, same era.
Theoretical considerations of potential-based reward shaping for multi-agent systems
Sam Devlin and Daniel Kudenko · 2011
Cited alongside, same era.
Protecting moving targets with multiple mobile resources
Fei Fang, Albert Xin Jiang, and Milind Tambe · 2013
Cited alongside, same era.
Robots that can adapt like animals
Antoine Cully, Jeff Clune, Danesh Tarapore, and Jean-Baptiste Mouret · 2015
Cited alongside, same era.
Escaping from saddle points—online stochastic gradient for tensor decomposition
Rong Ge, Furong Huang, Chi Jin, and Yang Yuan · 2015
Cited alongside, same era.
No-regret learning in Bayesian games
Jason Hartline, Vasilis Syrgkanis, and Eva Tardos · 2015
Cited alongside, same era.
Building generalizable agents with a realistic and rich 3D environment
Yi Wu, Yuxin Wu, Georgia Gkioxari, and Yuandong Tian · 2018
Later among the works it cites.
Emergent tool use from multi-agent autocurricula, 2019
Bowen Baker, Ingmar Kanitscheider, Todor Markov, Yi Wu, Glenn Powell, Bob McGrew, and Igor Mordatch · 2019
Later among the works it cites.
Open-ended learning in symmetric zero-sum games
David Balduzzi, Marta Garnelo, Yoram Bachrach, Wojciech M Czarnecki, Julien Perolat, Max Jaderberg, and Thore Graepel · 2019
Later among the works it cites.
Exploration by random network distillation
Yuri Burda, Harrison Edwards, Amos Storkey, and Oleg Klimov · 2019
Later among the works it cites.
Diversity is all you need: Learning skills without a reward function
Benjamin Eysenbach, Abhishek Gupta, Julian Ibarz, and Sergey Levine · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Deep reinforcement learning from self-play in imperfect-information games
Johannes Heinrich and David Silver · 2016
Cited alongside, same era.
Coordinate to cooperate or compete: abstract goals and joint intentions in social interaction
Max Kleiman-Weiner, Mark K Ho, Joseph L Austerweil, Michael L Littman, and Joshua B Tenenbaum · 2016
Cited alongside, same era.
End-to-end training of deep visuomotor policies
Sergey Levine, Chelsea Finn, Trevor Darrell, and Pieter Abbeel · 2016
Cited alongside, same era.
Learning to navigate in complex environments
Piotr Mirowski, Razvan Pascanu, Fabio Viola, Hubert Soyer, Andrew J Ballard, Andrea Banino, Misha Denil, Ross Goroshin, Laurent Sifre, Koray Kavukcuoglu, et al · 2016
Cited alongside, same era.
Training agent for first-person shooter game with actor-critic curriculum learning
Yuxin Wu and Yuandong Tian · 2016
Cited alongside, same era.
Intrinsically motivated goal exploration processes with automatic curriculum learning
Sébastien Forestier, Rémy Portelas, Yoan Mollard, and Pierre-Yves Oudeyer · 2017
Cited alongside, same era.
Deep fictitious play for finding Markovian Nash equilibrium in multi-agent games
Jiequn Han and Ruimeng Hu · 2019
Later among the works it cites.
Human-level performance in 3D multiplayer games with population-based reinforcement learning
Max Jaderberg, Wojciech M Czarnecki, Iain Dunning, Luke Marris, Guy Lever, Antonio Garcia Castaneda, Charles Beattie, Neil C Rabinowitz, Ari S Morcos, Avraham Ruderman, et al · 2019
Later among the works it cites.
Deep fictitious play for games with continuous action spaces
Nitin Kamra, Umang Gupta, Kai Wang, Fei Fang, Yan Liu, and Milind Tambe · 2019
Later among the works it cites.
Robust multi-agent reinforcement learning via minimax deep deterministic policy gradient
Shihui Li, Yi Wu, Xinyue Cui, Honghua Dong, Fei Fang, and Stuart Russell · 2019
Later among the works it cites.
Maven: Multi-agent variational exploration
Anuj Mahajan, Tabish Rashid, Mikayel Samvelyan, and Shimon Whiteson · 2019
Later among the works it cites.
Dota 2 with large scale deep reinforcement learning, 2019
OpenAI, :, Christopher Berner, Greg Brockman, Brooke Chan, Vicki Cheung, Przemysław Dębiak, Christy Dennison, David Farhi, Quirin Fischer, Shariq Hashme, Chris Hesse, Rafal Józefowicz, Scott Gray, Catherine Olsson, Jakub Pachocki, Michael Petrov, Henrique Pondé de Oliveira Pinto, Jonathan Raiman, Tim Salimans, Jeremy Schlatter, Jonas Schneider, Szymon Sidor, Ilya Sutskever, Jie Tang, Filip Wolski, and Susan Zhang · 2019
Later among the works it cites.
Macheng Shen and Jonathan P How · 2019
Later among the works it cites.
Grandmaster level in StarCraft II using multi-agent reinforcement learning
Oriol Vinyals, Igor Babuschkin, Wojciech M Czarnecki, Michaël Mathieu, Andrew Dudzik, Junyoung Chung, David H Choi, Richard Powell, Timo Ewalds, Petko Georgiev, et al · 2019
Later among the works it cites.
Deep reinforcement learning for green security games with real-time information
Yufei Wang, Zheyuan Ryan Shi, Lantao Yu, Yi Wu, Rohit Singh, Lucas Joppa, and Fei Fang · 2019
Later among the works it cites.
Learning to interactively learn and assist
Mark Woodward, Chelsea Finn, and Karol Hausman · 2019
Later among the works it cites.
Meta-world: A benchmark and evaluation for multi-task and meta reinforcement learning
Tianhe Yu, Deirdre Quillen, Zhanpeng He, Ryan Julian, Karol Hausman, Chelsea Finn, and Sergey Levine · 2019
Later among the works it cites.
Provable self-play algorithms for competitive reinforcement learning
Yu Bai and Chi Jin · 2020
Later among the works it cites.
Emergent tool use from multi-agent autocurricula
Bowen Baker, Ingmar Kanitscheider, Todor Markov, Yi Wu, Glenn Powell, Bob McGrew, and Igor Mordatch · 2020
Later among the works it cites.
Other-play for zero-shot coordination
Hengyuan Hu, Adam Lerer, Alex Peysakhovich, and Jakob Foerster · 2020
Later among the works it cites.
Towards practical multi-object manipulation using relational reinforcement learning
Richard Li, Allan Jabri, Trevor Darrell, and Pulkit Agrawal · 2020
Later among the works it cites.
Evolutionary population curriculum for scaling multi-agent reinforcement learning
Qian Long, Zihan Zhou, Abhinav Gupta, Fei Fang, Yi Wu, and Xiaolong Wang · 2020
Later among the works it cites.
Social diversity and social preferences in mixed-motive reinforcement learning
Kevin R McKee, Ian Gemp, Brian McWilliams, Edgar A Duéñez-Guzmán, Edward Hughes, and Joel Z Leibo · 2020
Later among the works it cites.
Julien Perolat, Remi Munos, Jean-Baptiste Lespiau, Shayegan Omidshafiei, Mark Rowland, Pedro Ortega, Neil Burch, Thomas Anthony, David Balduzzi, Bart De Vylder, et al · 2020
Later among the works it cites.
Dynamics-aware unsupervised discovery of skills
Archit Sharma, Shixiang Gu, Sergey Levine, Vikash Kumar, and Karol Hausman · 2020
Later among the works it cites.
Agar.io, 2020
Wikipedia · 2020
Later among the works it cites.
Prosocial learning agents solve generalized stag hunts better than selfish ones
Alexander Peysakhovich and Adam Lerer · 2044
Closest in time.